NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
crates.io · #2232 most downloaded on crates.io
A safe Rust wrapper for ONNX Runtime 1.28 - Optimize and accelerate machine learning inference & training
Last release 2 months ago
28 Jul 2026
Release timing varies
gaps range from 3 weeks to 7 months
Unknown
no stable releases
29 versions withdrawn
withdrawn after publishing
4 years old
43 releases · first in 2022
One column per quarter.
rc.13 skips ahead 4 ONNX Runtime versions to v1.28, bringing new operator support, bug & security fixes, and performance improvements.
ort useful, please consider sponsoring us on Open Collective 💖🤔 Need help upgrading? Ask questions in GitHub Discussions or in the pyke.io Discord server!
EP structs like ep::CUDA are now compile-time gated behind their respective Cargo feature flags. The EP feature flags only modifying runtime behavior was a common pain point, and this change brings ort in line with, like, every other Rust crate in existence.
⚠️ ort will now throw an error at link time if download-binaries is enabled and no binaries contain all of the requested EPs. Previously, if you enabled a nonsensical combination like cuda + coreml, ort would silently fall back to a CPU-only build. This caused great confusion as the only indicator of what went wrong was in log messages that don't show by default. A new lax-feature-matching Cargo feature will fallback to the closest fit instead of erroring.
rc.13 skips ahead 4 ONNX Runtime versions to v1.28, bringing new operator support, bug & security fixes, and performance improvements.
rc.13 only ships CUDA 13 binaries, as ONNX Runtime has deprecated CUDA 12.
Custom operators have been reworked to have a more Rusty interface:
struct MyOperator;
impl Operator for MyOperator {
+ type Kernel<'attr> = ort::operator::BoxedKernel<'attr>;
fn name(&self) -> &str {
"MyOperator"
}
- fn inputs(&self) -> Vec<OperatorInput> {
- vec![OperatorInput::required(TensorElementType::Float32), OperatorInput::required(TensorElementType::Float32)]
+ fn inputs(&self) -> impl IntoIterator<Item = OperatorInput> {
+ [OperatorInput::required(TensorElementType::Float32), OperatorInput::required(TensorElementType::Float32)]
}
- fn outputs(&self) -> Vec<OperatorOutput> {
- vec![OperatorOutput::required(TensorElementType::Float32)]
+ fn outputs(&self) -> impl IntoIterator<Item = OperatorOutput> {
+ [OperatorOutput::required(TensorElementType::Float32)]
}
- fn create_kernel(&self, _: &KernelAttributes) -> ort::Result<Box<dyn Kernel>> {
- Ok(Box::new(|ctx: &KernelContext| {
+ fn create_kernel<'attr>(&self, _: &KernelContext<'attr>) -> ort::Result<Self::Kernel<'attr>> {
+ Ok(Box::new(|ctx: &ComputeContext| {
...
}))
}
}ort now uses the wonderful #[diagnostic] attributes to drastically improve the clarity of commonly encountered errors.
Test coverage was expanded and ort was ran against Miri, and a few unsoundness bugs were fixed.
ort-web): Allow creating tensors from images.f16 type on nightly Rustort-candle): Update candle to 0.11ort-tract): Update tract to 0.23ort-sys can't create a system-wide cache dir.DynTensor to be Cloned.0 dimensions in tensors.try_extract_scalar to work on tensors with shape [1].load-dynamic fails.std for ndarray when std is enabled.copy-dylibs to work independently of download-binaries.build.rs on macOS when Xcode updates.💖 If you find ort useful, please consider sponsoring us on Open Collective 💖
ort useful, please consider sponsoring us on Open Collective 💖🤔 Need help upgrading? Ask questions in GitHub Discussions or in the pyke.io Discord server!
This release was made possible by Rime.ai!
Authentic AI voice models for enterprise.
🚨 If you used
ortwithdefault-features = false, enable theapi-24feature to use the latest features.
The big highlight of this release is multiversioning: ort can now use any minor version of ONNX Runtime from v1.17 to v1.24. New features are gated behind api-* feature flags, like api-20 or api-24. These flags will set the minimum version of ONNX Runtime required by ort.
More info 👉 https://ort.pyke.io/setup/multiversion
With ONNX Runtime 1.22 or later, ort will now automatically use an NPU if one is available for maximum efficiency & power savings! Setting your own execution providers will override this.
This is thanks to the super cool new SessionBuilder::with_auto_device API! There's also SessionBuilder::with_devices for finer control.
ort now ships builds for both CUDA 12 & CUDA 13! It should automatically detect which CUDA you're using, but if it gets it wrong, you can override it by setting the ORT_CUDA_VERSION environment variable to 12 or 13.
SessionBuilder error recoveryYou can now recover from errors when building a session by calling .recover() on the error type to get the SessionBuilder back.
Prebuilt binaries are now attested via GitHub Actions, so you can verify that they are untampered builds of ONNX Runtime coming straight from pyke.io.
To verify, download your binary package of choice and use the gh CLI to verify:
➜ gh attestation verify --owner pykeio ./x86_64-pc-windows-msvc+cu13.tar.lzma2
Loaded digest sha256:e96616510082108be228ad6ea026246a31650b7d446b330c6b9671fcb9ae6267 for file://./x86_64-pc-windows-msvc+cu13.tar.lzma2
Loaded 1 attestation from GitHub API
The following policy criteria will be enforced:
- OIDC Issuer must match:................... https://token.actions.githubusercontent.com
- Source Repository Owner URI must match:... https://github.com/pykeio
- Predicate type must match:................ https://slsa.dev/provenance/v1
- Subject Alternative Name must match regex: (?i)^https://github.com/pykeio/
✓ Verification succeeded!
sha256:e96616510082108be228ad6ea026246a31650b7d446b330c6b9671fcb9ae6267 was attested by:
REPO PREDICATE_TYPE WORKFLOW
pykeio/ort-artifacts https://slsa.dev/provenance/v1 .github/workflows/build-runner.yml@refs/heads/main
(Also note that the SHA-256 hash lines up with the one defined in dist.txt.)
ORT_LIB_LOCATION environment variable has been renamed to ORT_LIB_PATH.
_LOCATION.ort::tensor is now in ort::value, because why have a tensor module if the Tensor<T> type actually comes from the value module?IoBinding and Adapter were moved from their own modules into ort::session. All sub-modules of ort::session besides builder were collapsed into ort::session.ort::operator were collapsed into ort::operator.with_denormal_as_zero -> with_flush_to_zerowith_device_allocator_for_initializers -> with_device_allocated_initializersTensor::clone.💖 If you find ort useful, please consider sponsoring us on Open Collective 💖
ort useful, please consider sponsoring us on Open Collective 💖🤔 Need help upgrading? Ask questions in GitHub Discussions or in the pyke.io Discord server!
I'm sorry it took so long to get to this point, but the next big release of ort should be, finally, 2.0.0 🎉. I know I said that about one of the old alpha releases (if you can even remember those), but I mean it this time! Also, I would really like to not have to do another major release right after, so if you have any concerns about any APIs, please speak now or forever hold your peace!
A huge thank you to all the individuals who have contributed to the Collective over the years: Marius, Urban Pistek, Phu Tran, Haagen, Yunho Cho, Laco Skokan, Noah, Matouš Kučera, mush42, Thomas, Bartek, Kevin Lacker, & Okabintaro. You guys have made these past rc releases possible.
If you are a business using ort, please consider sponsoring me. Egress bandwidth from pyke.io has quadrupled in the last 4 months, and 90% of that comes from just a handful of businesses. I'm lucky enough that I don't have to pay for egress right now, but I don't expect that arrangement to last forever. pyke & ort have been funded entirely from my own personal savings for years, and (as I'm sure you're well aware 😂) everything is getting more expensive, so that definitely isn't sustainable.
Seeing companies that raise tens of millions in funding build large parts of their business on ort, ask for support, and then not give anything back just... seems kind of unfair, no?
ort-webort-web allows you to use the fully-featured ONNX Runtime on the Web! This time, it's hack-free and thus here to stay (it won't be removed, and then added back, and then removed again like last time!)
See the crate docs for info on how to port your application to ort-web; there is a little bit of work involved. For a very barebones sample application, see ort-web-sample.
Documentation for ort-web, like the rest of ort, will improve by the time 2.0.0 comes around. If you ever have any questions, you can always reach out via GitHub Discussions or Discord!
5d85209 Add WebNN & WASM execution providers for ort-web.#430 (💖 @jhonboy121) Support statically linking to iOS frameworks.#433 (💖 @rMazeiks) Implement more traits for GraphOptimizationLevel.6727c98 Make PrepackedWeights Send + Sync.15bd15c Make the TLS backend configurable with new tls-* Cargo features.f3cd995 Allow overriding the cache dir with the ORT_CACHE_DIR environment variable.8b3a1ed Load the dylib immediately when using ort::init_from.
#484 (💖 @michael-p) Update ndarray to v0.17.
ndarray dependency to v0.17, too.0084d08 New ort::lifetime tracing target tracks when objects are allocated/freed to aid in debugging leaks.2ee17aa Fix a memory leak in IoBinding.317be20 Don't store Environment as a static.
mutex lock failed: Invalid argument crash on macOS when exiting the process.466025c Fix unexpected CPU usage when copying GPU tensors.ecca246 Fix UB when extracting empty tensors.22f71ba Gate the ArrayExtensions trait behind the std feature, fixing #![no_std] builds.af63cea Fix an illegal memory access on no_std builds.#444 (💖 @pembem22) Fix Android link.1585268 Don't allow sessions to be created with non-CPU allocators#485 (💖 @mayocream) Fix load order when using cuda::preload_dylibs.c5b68a1 Fix AsyncInferenceFut drop behavior.ort in CI, please cache the ~/.cache/ort.pyke.io directory between runs.ort's dependency tree has shrunk a little bit, so it should build a little faster!b68c928 Overhaul build.rs
pkg-config support now requires the pkg-config feature.d269461 Make Metadata methods return Option<T> instead of Result<T>.47e5667 Gate preload_dylib and cuda::preload_dylibs behind a new preload-dylibs feature flag instead of load-dynamic.3b408b1 Shorten execution_providers to ep and XXXExecutionProvider to XXX.
38573e0 Simplify ThreadManager trait.x86_64-apple-darwin) has been dropped following upstream changes to ONNX Runtime & Rust.
--client_package_build, meaning default options will optimize for low-resource edge inference rather than high throughput.
x86-64-v3, aka Intel Haswell/Broadwell and AMD Zen (any Ryzen) or later.ort-tracttract to 0.22.2d40e05 ort-tract no longer claims it is ort-candle in ort::info().ort-candlecandle to 0.9.💖 If you find ort useful, please consider sponsoring us on Open Collective 💖
ort useful, please consider sponsoring us on Open Collective 💖🤔 Need help upgrading? Ask questions in GitHub Discussions or in the pyke.io Discord server!
You can now create a TensorRef directly from an ArrayView. Previously, tensors could only be created via Tensor::from_array (which, in many cases, performed a copy if borrowed data was provided). The new TensorRef::from_array_view (and the complementary TensorRefMut::from_array_view_mut) method(s) allows for the zero-copy creation of tensors directly from an ArrayView.
Tensor::from_array now only accepts owned data, so you should either refactor your code to use TensorRefs or pass ownership of the array to the Tensor.
⚠️
ndarrays must be in standard/contiguous memory layout to be converted to aTensorRef(Mut); see.as_standard_layout().
rc.10 now allows you to manually copy tensors between devices using Tensor::to!
// Create our tensor in CUDA memory
let cuda_allocator = Allocator::new(
&session,
MemoryInfo::new(AllocationDevice::CUDA, 0, AllocatorType::Device, MemoryType::Default)?
)?;
let cuda_tensor = Tensor::<f32>::new(&cuda_allocator, [1_usize, 3, 224, 224])?;
// Copy it back to CPU
let cpu_tensor = cuda_tensor.to(AllocationDevice::CPU, 0)?;There's also Tensor::to_async, which replicates the functionality of PyTorch's non_blocking=True. Additionally, Tensors now implement Clone.
ort is no longer just a wrapper for ONNX Runtime; it's a one-stop shop for inferencing ONNX models in Rust thanks to the addition of the alternative backend API.
Alternative backends wrap other inference engines behind ONNX Runtime's API, which can simply be dropped in and used in ort - all it takes is one line of code:
fn main() {
ort::set_api(ort_tract::api()); // <- magic!
let session = Session::builder()?
...
}2 alternative backends are shipping alongside rc.10 - ort-tract, powered by tract, and ort-candle, powered by candle, with more to come in the future.
Outside of the Rust ecosystem, these alternative backends can also be compiled as standalone libraries that can be directly dropped in to applications as a replacement for libonnxruntime. 🦀🦠
Models can be created entirely programmatically, or edited from an existing ONNX model via the new Model Editor API.
See src/editor/tests.rs for an example of how an ONNX model can be created programmatically. You can combine the Model Editor API with SessionBuilder::with_optimized_model_path to export the model outside Rust.
Many execution providers internally convert ONNX graphs to a framework-specific graph representation, like CoreML networks/TensorRT engines. This process can take a long time, especially for larger and more complex models. Since these generated artifacts aren't persisted between runs, they have to be created every time a session is loaded.
The new Compiler API allows you to compile an optimized, EP-ready graph ahead-of-time, so subsequent loads are lighting fast! ⚡
ModelCompiler::new(
Session::builder()?
.with_execution_providers([
TensorRTExecutionProvider::default().build()
])?
)?
.with_model_from_file("model.onnx")?
.compile_to_file("compiled_trt_model.onnx")?;#![no_std]🚨 BREAKING: If you previously usedortwithdefault-features = false...That will now disable
ort'sstdfeature, which means you don't get to use APIs that interact with the operating system, likeSessionBuilder::commit_from_file- APIs you probably need!To minimize breakage, manually enable the
stdfeature:[dependencies] ort = { version = "=2.0.0-rc.10", default-features = false, features = [ "std", ... ] }
ort no longer depends on std (but does still depend on alloc) - default-features = false will enable #![no_std] for ort.
🚨 BREAKING: Boolean options for ArmNN, CANN, CoreML, CPU, CUDA, MIGraphX, NNAPI, OpenVINO, & ROCm...If you previously used an option setter on one of these EPs that took no parameters (i.e. a boolean option that was
falseby default), note that these functions now do take a boolean parameter to align with Rust idiom.Migrating is as simple as passing
trueto these functions. Affected functions include:
ArmNNExecutionProvider::with_arena_allocatorCANNExecutionProvider::with_dump_graphsCPUExecutionProvider::with_arena_allocatorCUDAExecutionProvider::with_cuda_graphCUDAExecutionProvider::with_skip_layer_norm_strict_modeCUDAExecutionProvider::with_prefer_nhwcMIGraphXExecutionProvider::with_fp16MIGraphXExecutionProvider::with_int8NNAPIExecutionProvider::with_fp16NNAPIExecutionProvider::with_nchwNNAPIExecutionProvider::with_disable_cpuNNAPIExecutionProvider::with_cpu_onlyOpenVINOExecutionProvider::with_opencl_throttlingOpenVINOExecutionProvider::with_dynamic_shapesOpenVINOExecutionProvider::with_npu_fast_compileROCmExecutionProvider::with_exhaustive_conv_search
🚨 BREAKING: Renamed enum options for CANN, CUDA, QNN...The following EP option enums have been renamed to reduce verbosity:
CANNExecutionProviderPrecisionMode->CANNPrecisionModeCANNExecutionProviderImplementationMode->CANNImplementationModeCUDAExecutionProviderAttentionBackend->CUDAAttentionBackendCUDAExecutionProviderCuDNNConvAlgoSearch->CuDNNConvAlgorithmSearchQNNExecutionProviderPerformanceMode->QNNPerformanceModeQNNExecutionProviderProfilingLevel->QNNProfilingLevelQNNExecutionProviderContextPriority->QNNContextPriority
🚨 BREAKING: Updated CoreML options...
CoreMLExecutionProviderhas been updated to use a new registration API, unlocking more options. To migrate old options:
.with_cpu_only()->.with_compute_units(CoreMLComputeUnits::CPUOnly).with_ane_only()->.with_compute_units(CoreMLComputeUnits::CPUAndNeuralEngine).with_subgraphs()->.with_subgraphs(true)
rc.10 adds support for 3 execution providers:
ort.All binaries are now statically linked! This means the cuda and tensorrt features no longer use onnxruntime.dll/libonnxruntime.so. The EPs themselves do still require separate DLLs - like libonnxruntime_providers_cuda - but this change should make it significantly easier to set up and use ort with CUDA/TRT.
🚨 BREAKING: Migrating your custom operators...
- All methods under
Operatornow take&self.- The operator's kernel is no longer an associated type -
create_kernelis instead expected to return aBox<dyn Kernel>(which can now be created directly from a function!)impl Operator for MyCustomOp { - type Kernel = MyCustomOpKernel; - fn name() -> &'static str { + fn name(&self) -> &str { "MyCustomOp" } - fn inputs() -> Vec<OperatorInput> { + fn inputs(&self) -> Vec<OperatorInput> { vec![OperatorInput::required(TensorElementType::Float32)] } - fn outputs() -> Vec<OperatorOutput> { + fn outputs(&self) -> Vec<OperatorOutput> { vec![OperatorOutput::required(TensorElementType::Float32)] } - fn create_kernel(_: &KernelAttributes) -> ort::Result<Self::Kernel> { - Ok(MyCustomOpKernel) - } + fn create_kernel(&self, _: &KernelAttributes) -> ort::Result<Box<dyn Kernel>> { + Ok(Box::new(|ctx: &KernelContext| { + ... + })) + } }To add an operator to an
OperatorDomain, you now pass the operator by value instead of as a type parameter:let mut domain = OperatorDomain::new("io.pyke")?; -domain = domain.add::<MyCustomOp>()?; +domain = domain.add(MyCustomOp)?;
Custom operators have been internally revamped to reduce code size & compilation time, and allow operators to be Sized.
tracing dependency is now optional (but enabled by default).
tracing with default-features = false, enable the tracing feature.WARN but can be controlled at runtime via the ORT_LOG environment variable by setting it to one of verbose, info, warning, error, or fatal.parcel.pyke.io to cdn.pyke.io, so make sure to update firewall exclusions.build.rs hack for Apple platforms is no longer required. (9b31680)ureq dependency (used by download-binaries/fetch-models) has been ugpraded to v3.0.
ort with the fetch-models feature will use rustls as the TLS provider.ort-sys with the download-binaries feature will use native-tls since that pulls less dependencies (it previously used rustls). No prerequisites are required when building on Windows & macOS, but other platforms now require OpenSSL to be installed.Complex64 & Complex128, 4-bit integers, and 8 bit floats!
DynTensor::new to allocate a tensor and DynTensor::data_ptr to access its data.e136869)
Session::run can now be zero-alloc (on the Rust side)!Session::run now takes &mut self.
SessionOutputs::remove to get an owned session output.ort::inputs! no longer outputs a Result, so remove the trailing ? from any invocations of the macro.extract_tensor to extract a tensor to an ndarray has been renamed to extract_array, with extract_raw_tensor now taking the place of extract_tensor.
DynValue::try_extract_tensor(_mut) -> DynValue::try_extract_array(_mut)Tensor::extract_tensor(_mut) -> Tensor::extract_array(_mut)DynValue::try_extract_raw_tensor(_mut) -> DynValue::try_extract_tensor(_mut)Tensor::extract_raw_tensor(_mut) -> Tensor::extract_tensor(_mut)Session::run_async now always takes &RunOptions; Session::run_async_with_options has been removed.ValueType::tensor_dimensions) has been replaced with "shape" (so ValueType::tensor_shape) for consistency.ort::tensor::Shape, instead of a Vec<i64> directly.
ValueType::Tensor.dimension_symbols is its own struct, SymbolicDimensions.::from()/.into().SessionBuilder::with_execution_providers now takes AsRef<[EP]> instead of any iterable type.SessionBuilder::with_external_initializer_file_in_memory requires a Path for the path parameter instead of a regular &str.Tensor::new. (7a95f98)
Tensor::new will be manually zeroed on the Rust side.IoBinding::synchronize_* now takes &self so synchronize_outputs can actually be used as intended (e8d873a)XNNPACKExecutionProvider::is_available always returning false (5ad997c)AllocationDevice & MemoryInfo (3ca14c2)3e7e8fe)
5661450)ort-sys crate now specifies links, hopefully preventing linking conflicts (d2dc7c8)AllocationDevice (46c3376)ort-sys no longer tries to download binaries when building with --offline (d7d4493)4b6b163)tracing level instead of being knocked down a level (d8bcfd7)TensorRTExecutionProvider::with_context_memory_sharing (#327)
with_build_heuristics & with_sparisty (b6ddfd8)commit_from_url or ort-sys (eb51646/#323)Note that this does come with some breaking changes:
A previous ort release 'flattened' all exports, such that everything was exported at the crate root - ort::{TensorElementType, Session, Value}. This was done at a time when ort didn't export much, but now it exports a lot, so this was leading to some big, ugly use blocks.
rc.9 now has most exports behind their respective modules - Session is now imported as ort::session::Session, Tensor as ort::value::Tensor, etc. rust-analyzer and some quick searches on docs.rs can help you find the right paths to import.
extract optimization (1dbad54)Previously, calling any of the extract_tensor_* methods would have to call back to ONNX Runtime to determine the value's ValueType to ensure it was OK to extract. This involved a lot of FFI calls and a few allocations which could have a notable performance impact in hot loops.
Since a value's type never changes after it is created, the ValueType is now created when the Value is constructed (i.e. via Tensor::from_array or returned from a session). This makes extract_tensor_* a lot cheaper!
Note that this does come with some breaking changes:
&[i64] for their dimensions instead of Vec<i64>.Value::dtype() and Tensor::memory_info() now return &ValueType and &MemoryInfo respectively, instead of their non-borrowed counterparts.ValueType::Tensor now has an extra field for symbolic dimensions, dimension_symbols, so you might have to update matches on ValueType.2.0.0-rc.9 introduces a new trait: ThreadManager. This allows you to define custom thread create & join functions for session & environment thread pools! See the thread_manager.rs test for an example of how to create your own ThreadManager and apply it to a session, or an environment's GlobalThreadPoolOptions (previously EnvironmentGlobalThreadPoolOptions).
Additionally, sessions may now opt out of the environment's global thread pool if one is configured.
ort now provides ShapeInferenceContext, an interface for custom operators to provide a hint to ONNX Runtime about the shape of the operator's output tensors based on its inputs, which may open the doors to memory optimizations.
See the updated custom_operators.rs example to see how it works.
SessionOutputs has been slightly refactored to reduce memory usage and slightly increase performance. Most notably, it no longer derefs to a &BTreeMap.
The new SessionOutputs interface closely mirrors BTreeMap's API, so most applications require no changes unless you were explicitly dereferencing to a &BTreeMap.
ONNX Runtime v1.20.0 introduces a new Adapter format for supporting LoRA-like weight adapters, and now ort has it too!
An Adapter essentially functions as a map of tensors, loaded from disk or memory and copied to a device (typically whichever device the session resides on). When you add an Adapter to RunOptions, those tensors are automatically added as inputs (except faster, because they don't need to be copied anywhere!)
With some modification to your ONNX graph, you can add LoRA layers using optional inputs which Adapter can then override. (Hopefully ONNX Runtime will provide some documentation on how this can be done soon, but until then, it's ready to use in ort!)
let model = Session::builder()?.commit_from_file("tests/data/lora_model.onnx")?;
let lora = Adapter::from_file("tests/data/adapter.orl", None)?;
let mut run_options = RunOptions::new()?;
run_options.add_adapter(&lora)?;
let outputs = model.run_with_options(ort::inputs![Tensor::<f32>::from_array(([4, 4], vec![1.0; 16]))?]?, &run_options)?;PrepackedWeights allows multiple sessions to share the same weights across multiple sessions. If you create multiple Sessions from one model file, they can all share the same memory!
Currently, ONNX Runtime only supports prepacked weights for the CPU execution provider.
You can now override dynamic dimensions in a graph using SessionBuilder::with_dimension_override, allowing ONNX Runtime to do more optimizations.
Not all workloads need full performance all the time! If you're using ort to perform background tasks, you can now set a session's workload type to prioritize either efficiency (by lowering scheduling priority or utilizing more efficient CPU cores on some architectures), or performance (the default).
let session = Session::builder()?.commit_from_file("tests/data/upsample.onnx")?;
session.set_workload_type(WorkloadType::Efficient)?;ortsys! macro.
ort::api() return &ort_sys::OrtApi instead of NonNull<ort_sys::OrtApi>.AsPointer trait.
ptr() method now have an AsPointer implementation instead.RunOptions.ORT_CXX_STDLIB environment variable (mirroring CXXSTDLIB) to allow changing the C++ standard library ort links to.ValueRef & ValueRefMut leaking value memory.MemoryInfo's DeviceType instead of its allocation device to determine whether Tensors can be extracted.ORT_PREFER_DYNAMIC_LINK to work even when cuda or tensorrt are enabled.Sequence<T>.If you have any questions about this release, we're here to help:
Thank you to Thomas, Johannes Laier, Yunho Cho, Phu Tran, Bartek, Noah, Matouš Kučera, Kevin Lacker, and Okabintaro, whose support made this release possible. If you'd like to support ort as well, consider contributing on Open Collective 💖
🩷💜🩷💜
Nothing published for this version
The following functions have been updated to return T instead of ort::Result<T> :
The following functions have been updated to return T instead of ort::Result<T>:
MemoryInfo::memory_typeMemoryInfo::allocator_typeMemoryInfo::allocation_deviceMemoryInfo::device_idValue::<T>::memory_infoValue::<T>::dtypeValueType now implements Display.Sync for Value<T>.ep_context_embed_mode.Send for Allocator.Session::overridable_initializers to get a list of overridable initializers in the graph.ValueRef or ValueRefMut to a Value in certain cases.SessionBuilder::with_config_entry for adding custom session config options.ORT_PREFER_DYNAMIC_LINK, to override whether or not ort should prefer static or dynamic libs when ORT_LIB_LOCATION is specified.IoBinding.::ptr() to every C-backed struct to expose ort_sys pointers.Clone for MemoryInfo.with_arena_allocator is now with_use_arena.)IoBinding so it can be stored in a struct alongside a session.Sequence::extract_sequence now returns Value<T> instead of ValueRef<T>.Environment and ExecutionProvider Send + Sync.If you have any questions about this release, we're here to help:
Thank you to Brad Neuman, web3nomad, and Julien Cretin for contributing to this release!
Thank you to Thomas, Johannes Laier, Yunho Cho, Phu Tran, Bartek, Noah, Matouš Kučera, Kevin Lacker, and Okabintaro, whose support made this release possible. If you'd like to support ort as well, consider supporting us on Open Collective 💖
💜🩷💜🩷
ort::Error is no longer an enum, but rather an opaque struct with a message and a new ErrorCode field.
ort::Error refactorort::Error is no longer an enum, but rather an opaque struct with a message and a new ErrorCode field.
ort::Error still implements std::error::Error, so this change shouldn't be too breaking; however, if you were previously matching on ort::Errors, you'll have to refactor your code to instead match on the error's code (acquired with the Error::code() function).
AllocationDevice refactorThe AllocationDevice type has also been converted from an enum to a struct. Common devices like CUDA or DirectML are accessible via associated constants like AllocationDevice::CUDA & AllocationDevice::DIRECTML.
ModelMetadata::custom_keys() to get a Vec of all custom keys.SessionBuilder options affecting compute & graph optimizations.Allocator API. You can now allocate & free buffers acquired from a session or operator kernel context.ValueType::Optional.KernelContext::par_for, allowing operator kernels to use ONNX Runtime's thread pool without needing an extra dependency on a crate like rayon.Tensors from &CowArrays.tracing's attributes feature - a --no-default-features build of ort now only builds 9 crates!operator-libraries feature - you can still use SessionBuilder::with_operator_library, it's just no longer gated behind the feature!If you have any questions about this release, we're here to help:
Love ort? Consider supporting us on Open Collective 💖
❤️💚💙💛
Pre-built static libraries (i.e. not cuda or tensorrt ) are now linked with /MD instead of /MT on Windows ; i.e. MSVC CRT is no longer statically link
cuda or tensorrt) are now linked with /MD instead of /MT on Windows; i.e. MSVC CRT is no longer statically linked. This should resolve linking issues in some cases (particularly crates using other FFI libraries), but may cause issues for others. I have personally tested this in 2 internal pyke projects that depend on ort & many FFI libraries and haven't encountered any issues, but your mileage may vary.ort now depends on ndarray 0.16.wasm32-unknown-unknown support has been removed.
wasm32-unknown-unknown working in the first place was basically a miracle. Hacking ONNX Runtime to work outside of Emscripten took a lot of effort, but recent changes to Emscripten and ONNX Runtime have made this exponentially more difficult. Given I am not adequately versed on ONNX Runtime's internals, the nigh-impossibility of debugging weird errors, ort.ort in WASM, I suggest you use and/or support the development of alternative WASM-supporting ONNX inference crates like tract or WONNX.commit_from_url will be redownloaded.Trainer API, just like HF's TrainerCallbacks! This allows you to write custom logging/LR scheduling callbacks. See the updated train-clm-simple example for usage details.If you have any questions about this release, we're here to help:
Love ort? Consider supporting us on Open Collective 💖
❤️💚💙💛
This release addresses important linking issues with rc3, particularly regarding CUDA on Linux.
This release addresses important linking issues with rc3, particularly regarding CUDA on Linux.
cuDNN 9 is no longer required for CUDA 12 builds (but is still the default); set the ORT_CUDNN_VERSION environment variable to 8 to use cuDNN 8 with CUDA 12.
If you have any questions about this release, we're here to help:
Love ort? Consider supporting us on Open Collective 💖
❤️💚💙💛
ort now supports a (currently limited subset of) ONNX Runtime's Training API. You can use the on-device Training API for fine-tuning, online learning,
ort now supports a (currently limited subset of) ONNX Runtime's Training API. You can use the on-device Training API for fine-tuning, online learning, or even full pretraining, on any CPU or GPU.
The train-clm example pretrains a language model from scratch. There's also a 'simple' API and related example, which offers a basically one-line training solution akin to 🤗 Transformers' Trainer API:
trainer.train(
TrainingArguments::new(dataloader)
.with_lr(7e-5)
.with_max_steps(5000)
.with_ckpt_strategy(CheckpointStrategy::Steps(500))
)?;You can learn more about training with ONNX Runtime here. Please try it out and let us know how we can improve the training experience!
ort now ships with ONNX Runtime v1.18.
The CUDA 12 build requires cuDNN 9.x, so if you're using CUDA 12, you need to update cuDNN. The CUDA 11 build still requires cuDNN 8.x.
IoBindingIoBinding's previously rather unsound API has been reworked and actually documented.
Sometimes, you don't need to calculate all of the outputs of a session. Other times, you need to pre-allocate a session's outputs to save on slow device copies or expensive re-allocations. Now, you can do both of these things without IoBinding through a new API: OutputSelector.
let options = RunOptions::new()?.with_outputs(
OutputSelector::no_default()
.with("output")
.preallocate("output", Tensor::<f32>::new(&Allocator::default(), [1, 3, 224, 224])?)
);
let outputs = model.run_with_options(inputs!["input" => input.view()]?, &options)?;In this example, each call to run_with_options that uses the same options struct will use the same allocation in memory, saving the cost of re-allocating the output; and any outputs that aren't the output aren't even calculated.
String tensors are now Tensor<String> instead of DynTensor. They also no longer require an allocator to be provided to create or extract them. Additionally, Maps can also have string keys, and no longer require allocators.
Since value specialization, IntoTensorElementType was used to describe only primitive (i.e. f32, i64) elements. This has since been changed to PrimitiveTensorElementType, which is a subtrait of IntoTensorElementType. If you have type bounds that depended on IntoTensorElementType, you probably want to update them to use PrimitiveTensorElementType instead.
Operator kernels now support i64, string, Vec<f32>, Vec<i64>, and TensorRef attributes, among most other previously missing C API features.
Additionally, the API for adding an operator to a domain has been changed slightly; it is now .add::<Operator>() instead of .add(Operator).
ValueRef & ValueRefMut.EnvironmentBuilder::with_telemetry.
ExecutionProviderDispatch::error_on_failure will immediately error out session creation if the registration of an EP fails.RunOptions is now taken by reference instead of via an Arc.Session::run_async_with_options.libonnxruntime in library builds where crate-type=rlib/staticlib.i686-pc-windows-msvc.pkg-config.If you have any questions about this release, we're here to help:
Thank you to Florian Kasischke, cagnolone, Ryo Yamashita, and Julien Cretin for contributing to this release!
Thank you to Johannes Laier, Noah, Yunho Cho, Okabintaro, and Matouš Kučera, whose support made this release possible. If you'd like to support ort as well, consider supporting us on Open Collective 💖
❤️💚💙💛
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →