NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3407 most downloaded on PyPI
Optimum Library is an extension of the Hugging Face Transformers library, providing a framework to integrate third-party libraries from Hardware Partners and interface with their specific functionality.
Last release 1 months ago
04 Aug 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 57 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
5 years old
84 releases · first in 2021
The parameter from_transformers of ORTModel.from_pretrained will be deprecated in favor of export.
Additional architectures are supported in the ONNX export: PoolFormer, Pegasus, Audio Spectrogram Transformer, Hubert, SEW, Speech2Text, UniSpeech, UniSpeech-SAT, Wav2Vec2, Wav2Vec2-Conformer, WavLM, Data2Vec Audio, MPNet, stable diffusion VAE encoder, vision encoder decoder, Nystromformer, Splinter, GPT NeoX.
optimum.exporters.onnx by @michaelbenayoun in https://github.com/huggingface/optimum/pull/622A few additional architectures are supported in BetterTransformer: RoCBERT, RoFormer, Marian
With ORTModelForMaskedLM, ORTModelForVision2Seq, ORTModelForAudioClassification, ORTModelForCTC, ORTModelForAudioXVector, ORTModelForAudioFrameClassification, ORTStableDiffusionPipeline.
Reference: https://huggingface.co/docs/optimum/main/en/onnxruntime/package_reference/modeling_ort and https://huggingface.co/docs/optimum/main/en/onnxruntime/usage_guides/models#export-and-inference-of-stable-diffusion-models
In the ONNX export, it is possible to pass the options --fp16 --device cuda to export using float16 when a GPU is available, directly with the native torch.onnx.export.
Example: optimum-cli export onnx --model gpt2 --fp16 --device cuda gpt2_onnx/
torch.float16 type by @fxmarty in https://github.com/huggingface/optimum/pull/749TFLite export is now supported, with static shapes:
optimum-cli export tflite --help
optimum-cli export tflite --model bert-base-uncased --sequence_length 128 bert_tflite/
exporters.tflite initial support by @michaelbenayoun in https://github.com/huggingface/optimum/pull/716The ONNX export optionally supports the ONNX Runtime optimizations directly in the export, passing the --optimize O1, up to --optimize O4 option:
optimum-cli export onnx --help
optimum-cli export onnx --model t5-small --optimize O3 t5small_onnx/
ONNX Runtime quantization is supported directly in command line, using optimum-cli onnxruntime quantize:
optimum-cli onnxruntime quantize --help
optimum-cli onnxruntime quantize --onnx_model distilbert_onnx --avx512
ONNX Runtime optimization is supported directly in command line, using optimum-cli onnxruntime optimize:
optimum-cli onnxruntime optimize --help
optimum-cli onnxruntime optimize --onnx_model distilbert_onnx -O3
Up no now, for decoders, two ONNX were used:
This release introduces the support in the ONNX export and in ORTModelForCausalLM of a single ONNX handling both steps of the decoding. This allows to reduce memory usage, as weights are not duplicated between two separate models during inference.
Using a single ONNX for decoders can be used by passing use_merged=True to ORTModelForCausalLM.from_pretrained, loading directly from a PyTorch model:
from optimum.onnxruntime import ORTModelForCausalLM
model = ORTModelForCausalLM.from_pretrained("gpt2", export=True, use_merged=True)
Alternatively, using a single ONNX for decoders is the default behavior in the ONNX export, that can later be used for example with ORTModelForCausalLM, the command optimum-cli export onnx --model gpt2 gpt2_onnx/ will produce:
└── gpt2_onnx
├── config.json
├── decoder_model_merged.onnx
├── decoder_model.onnx
├── decoder_with_past_model.onnx
├── merges.txt
├── special_tokens_map.json
├── tokenizer_config.json
├── tokenizer.json
└── vocab.json
The decoder_model.onnx and decoder_with_past_model.onnx are kept separate for backward compatibility, but during inference using solely decoder_model_merged.onnx is enough.
ORTModelForCausalLM by @JingyaHuang in https://github.com/huggingface/optimum/pull/647ORTModel accept numpy arrays as inputs, in addition to PyTorch tensors. This is only the case for models that use a single ONNX.
--monolith.--task causal-lm instead of --task causal-lm-with-past.block_sparse attention type being written in pure numpy in Transformers, and hence not exportable to ONNX: https://github.com/huggingface/optimum/pull/778from_transformers of ORTModel.from_pretrained will be deprecated in favor of export.use_cache=True to ORTModel and no ONNX with cache is available by @fxmarty in https://github.com/huggingface/optimum/pull/650from optimum.onnxruntime import QuantizationConfig by @fxmarty in https://github.com/huggingface/optimum/pull/715ORTTrainer by @JingyaHuang in https://github.com/huggingface/optimum/pull/709onnxruntime/modeling_ort.py refactor, part 1 by @michaelbenayoun in https://github.com/huggingface/optimum/pull/698ORTTrainer inference with ONNX Runtime backend by @JingyaHuang in https://github.com/huggingface/optimum/pull/737BetterTransformer.transform() by @fxmarty in https://github.com/huggingface/optimum/pull/750exporters.onnx output names and dynamic axes fix by @michaelbenayoun in https://github.com/huggingface/optimum/pull/731BT] Add stable layer-norm Wav2vec2 by @younesbelkada in https://github.com/huggingface/optimum/pull/803Full Changelog: https://github.com/huggingface/optimum/compare/v1.6.0...v1.7.0
One column per quarter.
Fix past key/value reuse in decoders following transformers 4.26.0 release and renaming: https://github.com/huggingface/optimum/commit/b9211d6826b9270
Full Changelog: https://github.com/huggingface/optimum/compare/v1.6.3...v1.6.4
Fixes ORTTrainer for the inference with the ONNX Runtime backend.
Fixes ORTTrainer for the inference with the ONNX Runtime backend.
Support generation config in ORTModel by @fxmarty in https://github.com/huggingface/optimum/pull/651
The export of speech-to-text architecture as a single ONNX file (that handles both the encoding and decoding) fails do to a regression with the latest transformers version: https://github.com/huggingface/optimum/issues/721
Full Changelog: https://github.com/huggingface/optimum/compare/v1.6.1...v1.6.2
Revert breaking removal of EncoderOnnxConfig, DecoderOnnxConfig, _DecoderWithLMhead by @fxmarty in https://github.com/huggingface/optimum/pull/643
Full Changelog: https://github.com/huggingface/optimum/compare/v1.6.0...v1.6.1
Deprecate PyTorch 1.12. for BetterTransformer by @fxmarty in https://github.com/huggingface/optimum/pull/513
The Optimum command line interface is introduced, and is now the official entrypoint for the ONNX export. Example commands:
optimum-cli --help
optimum-cli export onnx --help
optimum-cli export onnx --model bert-base-uncased --task sequence-classification bert_onnx/
Optimum now supports the ONNX export of stable diffusion models from the diffusers library:
optimum-cli export onnx --model runwayml/stable-diffusion-v1-5 sd_v15_onnx/
BetterTransformer integration includes new models in this release: CLIP, RemBERT, mBART, ViLT, FSMT
The complete list of supported models is available in the documentation.
Bettertransformer support for FSMT by @Sumanth077 in https://github.com/huggingface/optimum/pull/494BetterTransformer support for ViLT architecture by @ka00ri in https://github.com/huggingface/optimum/pull/508MBart support for BetterTransformer by @ravenouse in https://github.com/huggingface/optimum/pull/516The ONNX export now supports Swin, MobileNet-v1, MobileNet-v2.
ONNX] add mobilenet support by @younesbelkada in https://github.com/huggingface/optimum/pull/633Encoder-decoder or decoder-only models normally making use of the generate() method in transformers can now be exported in several files using the --for-ort argument:
optimum-cli export onnx --model t5-small --task seq2seq-lm-with-past --for-ort t5_small_onnx
yielding:
.
└── t5_small_onnx
├── config.json
├── decoder_model.onnx
├── decoder_with_past_model.onnx
├── encoder_model.onnx
├── special_tokens_map.json
├── spiece.model
├── tokenizer_config.json
└── tokenizer.json
Passing --for-ort, exported models are expected to be loadable directly into ORTModel.
--for-ort from optimum.exporters.onnx in ORTDecoder by @fxmarty in https://github.com/huggingface/optimum/pull/554The ONNX export from PyTorch normally creates external data in case the exported model is larger than 2 GB. This release introduces a better support for the export and use of large models, writting all external data into a .onnx_data file if necessary.
Various improvements to allow for a better user experience in the ONNX Runtime integration:
ORTModel, ORTModelDecoder and ORTModelForConditionalGeneration can now load any ONNX model files regardless of their names, allowing to load optimized and quantized models without having to specify a file name argument.
ORTModel.from_pretrained() with from_transformers=True now downloads and loads the model in a temporary directory instead of the cache, which was not a right place to store it.
ORTQuantizer.save_pretrained() now saves the model configuration and the preprocessor, making the exported directory usable end-to-end.
ORTOptimizer.save_pretrained() now saves the preprocessor, making the exported directory usable end-to-end.
ONNX Runtime integration API improvement by @michaelbenayoun in https://github.com/huggingface/optimum/pull/515
The shape of the example input to provide for the export to ONNX can be overridden in case the validity of the ONNX model is sensitive to the shape used during the export.
Read more: optimum-cli export onnx --help
use_cache=True for ORTModelForCausalLMReusing past key values for models using ORTModelForCausalLM (e.g. gpt2) is now possible using use_cache=True, avoiding to recompute them at each iteration of the decoding:
from transformers import AutoTokenizer
from optimum.onnxruntime import ORTModelForCausalLM
import torch
tokenizer = AutoTokenizer.from_pretrained("gpt2")
model = ORTModelForCausalLM.from_pretrained("gpt2", from_transformers=True, use_cache=True)
inputs = tokenizer("My name is Arthur and I live in", return_tensors="pt")
gen_tokens = model.generate(**inputs)
tokenizer.batch_decode(gen_tokens)
ORTModelForCustomTasks now supports IO Binding when using CUDAExecutionProvider.
Along with --for-ort, when passing --task causal-lm-with-past , --task seq2seq-with-past or --task speech2seq-lm-with-past during the ONNX export exports two models: one not using the previously computed keys/values, and one using them.
An experimental support is introduced to merge the two models in one. Example:
optimum-cli export onnx --model t5-small --task seq2seq-lm-with-past --for-ort t5_onnx/
import onnx
from optimum.onnx import merge_decoders
decoder = onnx.load("t5_onnx/decoder_model.onnx")
decoder_with_past = onnx.load("t5_onnx/decoder_with_past_model.onnx")
merged_model = merge_decoders(decoder, decoder_with_past)
onnx.save(merged_model, "t5_onnx/decoder_merged_model.onnx")
norm_first by @younesbelkada in https://github.com/huggingface/optimum/pull/510encoder_last_hidden_state as an output for encoder-decoder models by @fxmarty in https://github.com/huggingface/optimum/pull/601use_io_binding default value for different execution providers by @JingyaHuang in https://github.com/huggingface/optimum/pull/604Full Changelog: https://github.com/huggingface/optimum/compare/v1.5.2...v1.6.0
The following contributors have made significant changes to the library over the last release:
Constraint temporarily numpy<1.24.0
Constraint temporarily numpy<1.24.0 (#614)
Deprecate PyTorch 1.12. for BetterTransformer with better error message
Deprecate PyTorch 1.12. for BetterTransformer with better error message (#513)
Convert your model into its PyTorch `BetterTransformer` format using a one liner with the new BetterTransformer integration for faster inference on CP
Convert your model into its PyTorch BetterTransformer format using a one liner with the new BetterTransformer integration for faster inference on CPU and GPU!
from optimum.bettertransformer import BetterTransformer
model = BetterTransformer.transform(model)
Check the full list of supported models in the documentaiton, and check out the Google Colab demo.
BetterTransformer integration (#423)ORT models (except for ORTModelForCustomTasks) now support IOBinding to avoid data copying overheads between the host and device. Significant inference speedup during the decoding process on GPU.
By default, use_io_binding is set to True when using CUDA. You can turn off the IOBinding in case of any memory issue:
from optimum.onnxruntime import ORTModelForSeq2SeqLM
model = ORTModelForSeq2SeqLM.from_pretrained("optimum/t5-small", use_io_binding=False)
optimum.exporters is a new module that handles the export of PyTorch and TensorFlow models to several backends. Only ONNX is supported for now, and more than 50 architectures can already be exported, among which BERT, GPT-Neo, Bloom, T5, ViT, Whisper, CLIP.
The export can be done via the CLI:
python -m optimum.exporters.onnx --model openai/whisper-tiny.en whisper_onnx/
For more information, check the documentation.
optimum.exporters creation (#403)optimum.exporters.optimum.onnxruntime, IO binding is also supported.Note: For the now the export from optimum.exporters will not be usable by ORTModelForSpeechSeq2Seq. To be able to run inference, export Whisper directly using ORTModelForSpeechSeq2Seq. This will be solved in the next release.
optimum.onnxruntime and optimum.exporters (#420)transformers 4.23.1 (#434)ORTModel can load models from subfolders in a similar fashion as in transformers (#443)ORTOptimizer has been refactored, and a factory class has been added to create common OptimizationConfigs (#457)Add inference with ORTModel to ORTTrainer and ORTSeq2SeqTrainer #189
ORTModel to ORTTrainer and ORTSeq2SeqTrainer #189InferenceSession options and provider to ORTModel #271ORTOptimizertorch.fx transformations #348torch.fx transformations now use the marking methods mark_as_transformed, mark_as_restored, get_transformed_nodes #385BaseConfig for transformers 4.22.0 release #386ORTTrainer for transformers 4.22.1 release #388provider_options to ORTModel #401ORTModel, as transformers does for pipelines #427Refactorization of ORTQuantizer (#270) and ORTOptimizer
ORTQuantizer (#270) and ORTOptimizer (#294)ORTModelForCustomTasks allowing ONNX Runtime inference support for custom tasks (#303)ORTModelForMultipleChoice allowing ONNX Runtime inference for models with multiple choice classification head (#358)FuseBiasInLinear a transformation that fuses the weight and the bias of linear modules (#253)past_key_values during ONNX Runtime inference of Seq2Seq models (#241)The optimum.fx.optimization module (#232) provides a set of torch.fx graph transformations, along with classes and functions to write your own transfo
The optimum.fx.optimization module (#232) provides a set of torch.fx graph transformations, along with classes and functions to write your own transformations and compose them.
Transformation and ReversibleTransformation represent non-reversible and reversible transformations, and it is possible to write such transformations by inheriting from those classescompose utility function enables transformation compositionMergeLinears: merges linear layers that have the same inputChangeTrueDivToMulByInverse: changes a division by a static value to a multiplication of its inverseORTModelForSeq2SeqLM (#199) allows ONNX export and ONNX Runtime inference for Seq2Seq models.
Below is an example that downloads a T5 model from the Hugging Face Hub, exports it through the ONNX format and saves it :
from optimum.onnxruntime import ORTModelForSeq2SeqLM
# Load model from hub and export it through the ONNX format
model = ORTModelForSeq2SeqLM.from_pretrained("t5-small", from_transformers=True)
# Save the exported model in the given directory
model.save_pretrained(output_dir)
ORTModelForImageClassification (#226) allows ONNX Runtime inference for models with an image classification head.
Below is an example that downloads a ViT model from the Hugging Face Hub, exports it through the ONNX format and saves it :
from optimum.onnxruntime import ORTModelForImageClassification
# Load model from hub and export it through the ONNX format
model = ORTModelForImageClassification.from_pretrained("google/vit-base-patch16-224", from_transformers=True)
# Save the exported model in the given directory
model.save_pretrained(output_dir)
Adds support for converting model weights from fp32 to fp16 by adding a new optimization parameter (fp16) to OptimizationConfig (#273).
Additional pipelines tasks are now supported, here is a list of the supported tasks along with the default model for each:
Below is an example that downloads a T5 small model from the Hub and loads it with transformers pipeline for translation :
from transformers import AutoTokenizer, pipeline
from optimum.onnxruntime import ORTModelForSeq2SeqLM
tokenizer = AutoTokenizer.from_pretrained("optimum/t5-small")
model = ORTModelForSeq2SeqLM.from_pretrained("optimum/t5-small")
onnx_translation = pipeline("translation_en_to_fr", model=model, tokenizer=tokenizer)
text = "What a beautiful day !"
pred = onnx_translation(text)
# [{'translation_text': "C'est une belle journée !"}]
The ORTModelForXXX execution provider default value is now set to CPUExecutionProvider (#203). Before, if no execution provider was provided, it was set to CUDAExecutionProvider if a gpu was detected, or to CPUExecutionProvider otherwise.
Remove intel sub-package, migrating to `optimum-intel`
optimum-intel (#212)ORTModel optimized and quantized models (#214)Extend QuantizationPreprocessor to dynamic quantization
QuantizationPreprocessor to dynamic quantization (https://github.com/huggingface/optimum/pull/196)huggingface_hub version and protobuf fix (https://github.com/huggingface/optimum/pull/205)Add support to Python version 3.7
Add support to Python version 3.7 (https://github.com/huggingface/optimum/pull/176)
`ORTModelForXXX` classes such as `ORTModelForSequenceClassification` were integrated with the Hugging Face Hub in order to easily export models throug
ORTModelForXXX classes such as ORTModelForSequenceClassification were integrated with the Hugging Face Hub in order to easily export models through the ONNX format, load ONNX models, as well as easily save the resulting model and push it to the 🤗 Hub by using respectively the save_pretrained and push_to_hub methods. An already optimized and / or quantized ONNX model can also be loaded using the ORTModelForXXX classes using the from_pretrained method.
Below is an example that downloads a DistilBERT model from the Hub, exports it through the ONNX format and saves it :
from optimum.onnxruntime import ORTModelForSequenceClassification
# Load model from hub and export it through the ONNX format
model = ORTModelForSequenceClassification.from_pretrained(
"distilbert-base-uncased-finetuned-sst-2-english",
from_transformers=True
)
# Save the exported model
model.save_pretrained("a_local_path_for_convert_onnx_model")
Built-in support for transformers pipelines was added. This allows us to leverage the same API used from Transformers, with the power of accelerated runtimes such as ONNX Runtime.
The currently supported tasks with the default model for each are the following :
Below is an example that downloads a RoBERTa model from the Hub, exports it through the ONNX format and loads it with transformers pipeline for question-answering.
from transformers import AutoTokenizer, pipeline
from optimum.onnxruntime import ORTModelForQuestionAnswering
# load vanilla transformers and convert to onnx
model = ORTModelForQuestionAnswering.from_pretrained("deepset/roberta-base-squad2",from_transformers=True)
tokenizer = AutoTokenizer.from_pretrained("deepset/roberta-base-squad2")
# test the model with using transformers pipeline, with handle_impossible_answer for squad_v2
optimum_qa = pipeline(task, model=model, tokenizer=tokenizer, handle_impossible_answer=True)
prediction = optimum_qa(
question="What's my name?", context="My name is Philipp and I live in Nuremberg."
)
print(prediction)
# {'score': 0.9041663408279419, 'start': 11, 'end': 18, 'answer': 'Philipp'}
ORTTrainer, previously not enabled when inference was performed with ONNX Runtime in #152Installation details added for Optimum-Habana which provides optimized transformers integration for Intel's Habana Gaudi Processor (HPU).
ORTModel.IncludeFullyConnectedNodes class to find the nodes composing the fully connected layers in order to (only) target the latter for quantization to limit the accuracy drop.QuantizationPreprocessor so that the intersection of the two sets representing the nodes to quantize and the nodes to exclude from quantization to be an empty set.Seq2SeqORTTrainer to ORTSeq2SeqTrainer for clarity and to keep consistency.ORTOptimizer support for ELECTRA models.ORTConfig which contains optimization and quantization config.The ORTTrainer and Seq2SeqORTTrainer are two newly experimental classes.
The ORTTrainer and Seq2SeqORTTrainer are two newly experimental classes.
ORTTrainer and Seq2SeqORTTrainer were created to have a similar user-facing API as the Trainer and Seq2SeqTrainer of the Transformers library.ORTTrainer allows the usage of the ONNX Runtime backend to train a given PyTorch model in order to accelerate training. ONNX Runtime will run the forward and backward passes using an optimized automatically-exported ONNX computation graph, while the rest of the training loop is executed by native PyTorch.ORTTrainer allows the usage of ONNX Runtime inferencing during both the evaluation and the prediction step.Seq2SeqORTTrainer, ONNX Runtime inferencing is incompatible with --predict_with_generate, as the generate method is not supported yet.The ORTQuantizer and ORTOptimizer classes underwent a massive refactoring that should allow a simpler and more flexible user-facing API.
ORTQuantizer method partial_fit. This is especially useful when using memory-hungry calibration methods such as Entropy and Percentile methods.OptimizationConfig, QuantizationConfig and CalibrationConfig were added in order to better segment the different ONNX Runtime related parameters instead of having one unique configuration ORTConfig.QuantizationPreprocessor class was added in order to find the nodes to include and / or exclude from quantization, by finding the nodes following a given pattern (such as the nodes forming LayerNorm for example). This is particularly useful in the context of static quantization, where the quantization of modules such as LayerNorm or GELU are responsible of important drop in accuracy.An ORTConfig class was introduced, allowing the user to define the desired export, optimization and quantization strategies.
ORTConfig class was introduced, allowing the user to define the desired export, optimization and quantization strategies.ORTOptimizer class takes care of the model's ONNX export as well as the graph optimization provided by ONNX Runtime. In order to create an instance of ORTOptimizer, the user needs to provide an ORTConfig object, defining the export and graph-level transformations informations. Then optimization can be perfomed by calling the ORTOptimizer.fit method.ORTQuantizer class. In order to create an instance of ORTQuantizer, the user needs to provide an ORTConfig object, defining the export and quantization informations, such as the quantization approach to use or the activations and weights data types. Then quantization can be applied by calling the ORTQuantizer.fit method.We have also added a new class called IncOptimizer which will take care of combining the pruning and the quantization processes.
Nothing published for this version
With this release, we enable Intel Neural Compressor v1.8 magnitude pruning for a variety of NLP tasks with the introduction of IncTrainer which handl
With this release, we enable Intel Neural Compressor v1.8 magnitude pruning for a variety of NLP tasks with the introduction of IncTrainer which handles the pruning process.
With this release, we enable Intel Neural Compressor v1.7 PyTorch dynamic, post-training and aware-training quantization for a variety of NLP tasks. T
With this release, we enable Intel Neural Compressor v1.7 PyTorch dynamic, post-training and aware-training quantization for a variety of NLP tasks. This support includes the overall process, from quantization application to the loading of the resulting quantized model. The latter being enabled by the introduction of the IncQuantizedModel class.
Nothing published for this version
Initial release for early access to Optimum library featuring Intel's LPOT quantization and pruning support.
Initial release for early access to Optimum library featuring Intel's LPOT quantization and pruning support.
Your coding agent can read these notes before it upgrades. Set up the MCP server →