NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
npm · #4595 most downloaded on npm
ONNXRuntime JavaScript API library
Last release 20 days ago
14 Sep 2026
Release timing varies
gaps range from 9 days to 3 months
Nearly every release is documented
notes for 34 of 35 stable releases
Nothing withdrawn
no release was ever pulled
5 years old
133 releases · first in 2021
One column per quarter.
Nothing published for this version
ONNX Runtime 1.30.0 expands generative AI inference, improves CPU and GPU performance, adds Go bindings, and strengthens runtime reliability. These no
ONNX Runtime 1.30.0 expands generative AI inference, improves CPU and GPU performance, adds Go bindings, and strengthens runtime reliability. These notes cover changes since ONNX Runtime 1.29.1.
-Donnxruntime_USE_FP4_QMOE=OFF (#32096, #32163).block_size=32. Set -Donnxruntime_USE_FPA_INTB_GEMM_FULL=ON when building from source to retain the full kernel set, including BF16, zero-point, bias, larger-block-size, and native Hopper variants (#32324).Gemm and MatMul execution is gated on hardware acceleration. CPU-assigned FP16 nodes without a matching kernel now fall back to FP32 (#32301, #32197).Split, Scan, GatherND, ScatterND, SpaceToDepth/DepthToSpace, Crop, Conv, Normalizer, and pooling (#29461, #31668, #32034, #32039, #32076, #32157, #32160, #32161, #32345, #32349).BifurcationDetector inputs, generation subgraph shapes, and QEmbed segment inputs. BeamSearch buffer expansion now uses dynamic shape storage (#31648, #31701, #32009, #32078, #32144).TreeEnsemble node references and bounded subtree comparison, rejected non-finite CPU RoiAlign coordinates, and required ImageScaler bias to match the channel count (#32031, #32043, #32011, #32002).MatMulFpQ4 shape inputs, and checked MLAS blockwise quantization/dequantization index ranges (#31682, #32032, #32007).MatMulNBits, RemovePadding, RotaryEmbedding, SparseAttention, Whisper beam search, NMS, QDQ, and GatherElements (#31643, #31994, #31995, #31996, #31998, #32014, #32029, #32030).CudaAsyncBuffer staging storage alive across CUDA graph replay (#31968, #32121).js-yaml, joi, fast-uri, and the Next.js end-to-end fixture (#32397, #32486, #32488, #32505, #32508).KernelContext::GetPreallocatedOutput (#29726, #32089).EngramGate and NGramHashMapping, and expanded kernel coverage for Qwen-3.5 operators (#32268, #32106).is_causal attribute to PagedAttention (#32515, #32225).VarlenCausalConvWithState for continuous batching and compact variable-length causal-convolution state updates (#32168, #32290).GatedDeltaNet operator and BFloat16 support for CUDA GatedDeltaNet (#32282, #32307).ORT_QMOE_FP4_DEEPGEMM=1, default off). This path is disabled on Windows (#32122, #32485).MatMul shapes and refined fpA-intB GEMV support checks (#31478, #32338).TopK, ArgMax, and ArgMin performance for wide last axes, and accelerated low-lane INT64 CumSum (#32404, #32092, #32238).Slice fast path for contiguous subregions and removed pinned-buffer use from Split and Concat fast paths (#28902, #32410).ReduceMean kernels and fixed ScatterElements reduction dispatch by element type and signed-zero handling in Abs (#32326, #29879, #31477).Gather support and optimized MatMulNBits wide tiles with subgroup shuffle (#31714, #31703).Split when all output segments are vec4-aligned, and selected pooling paths based on occupancy (#29820, #32251, #32313).onnxruntime_perf_test -i options to WebGPU (#32074, #31971, #32316).GPUDevice instances, corrected MatMul pipeline-cache keys and the 1D-dispatch shader fast path, and changed copy_tensors misuse to report errors instead of terminating the process (#32259, #32048, #32343, #32315).SkipLayerNormalization support and corrected output-rank validation and fallback data-type support checks (#32377, #31708, #32067, #32293).kernel_shape and output_padding lengths during DirectML kernel setup (#31999).QuantizeLinear rounding, prevented CPU TensorScatter index overflow, serialized ScatterND string updates, and widened Compress loop counters (#32452, #32012, #32033, #32008).LpNormalization inputs, zero-element BiasGelu/FastGelu, extreme Trilu diagonals, and empty reduction axes (#32020, #31698, #32013, #32156).Slice starts rank (#31670, #31678, #32018, #32044).Gemm (#32038, #32143, #32426, #32435).NodeAttrHelper string-default lifetimes (#32051, #32138, #32019).Env, and clarified how external-initializer paths interact with EP context paths (#32502, #32503, #32442).RunAsync arguments until completion (#32041, #32015).Conv3DNaive shader (#32469, #32357).Thanks to our 57 human contributors for this release!
@4n4ny4, @apsonawane, @arnej27959, @baijumeswani, @bmehta001, @chilo-ms, @crvineeth97, @daijh, @danfiedler-msft, @danielsongmicrosoft, @dannyota, @DKAIN-py, @edgchen1, @ericcraw, @eserscor, @fanchenkong1, @hanbitmyths, @hariharans29, @hdharpure9922, @Honry, @jambayk, @javier-intel, @jchen10, @jiafatom, @jnagi-intel, @justinchuby, @kadyrbekovhamit-cyber, @kunal-vaishnavi, @Lapis0x0, @LOGO127, @Manogna-Sree, @martin-klacer-arm, @mei1127, @miaobin, @mirounga, @miyanyan, @MohamedElashri, @mustjab, @Nikhi00718, @Noperi0r, @Novestars, @pkubaj, @preetha-intel, @qjia7, @rvandermeulen, @sanaa-hamel-microsoft, @skottmckay, @sushraja-msft, @swetha097, @sylvesterkaczmarek, @tianleiwu, @titaiwangms, @toothache, @xadupre, @xhcao, @xiaofeihan1, @Zestion
Full Changelog: rel-1.29.1...rel-1.30.0
Release highlights were prepared with AI assistance.
Nothing published for this version
Fixed a path traversal vulnerability in TensorRT and NvTensorRTRTX engine refitting by making external-data path validation unconditional ( #29396 ).
ORT_DISABLE_TELEMETRY=1 before initialization disables non-Windows telemetry for the process (#27379, #29872).onnxruntime/python/tools/tensorrt dashboard tooling was removed. This does not affect the TensorRT Execution Provider APIs (#29395).k attribute against the number of experts and fixed a CPU TensorScatter security issue (#29907, #29916).Range, and CropAndResize (#29254, #29255, #29265, #29579, #29595, #29605, #29871, #31636, #31671, #31675, #31676, #31684).OrtApi::GetValue and validated DML constant tensor byte sizes (#29157, #31665).adm-zip for onnxruntime-node (#29827, #29926, #31192).ORT_INTRA_OP_NUM_THREADS and ORT_INTER_OP_NUM_THREADS. Explicit thread settings still take precedence, and 0 preserves machine-sized defaults (#29688).EpContext nodes, and wired maximum-shape inference into workspace estimation (#29607, #29799, #31613).MRotaryEmbedding contrib operator for Qwen mRoPE variants (#29261, #31728).onnxruntime_perf_test through --data_shape, plus verbose graph-transformer tracing and broader inference-session error-path coverage (#29555, #29558, #29569, #29571).PagedAttention with quantized KV cache, XQA decode, MLA, QK-Norm, and head-sink support (#29912).Attention CUDA kernel and enabled cuDNN SDPA for contrib Attention (#29715, #29717).attention_bias support to the GroupQueryAttention unfused path and state_window support to LinearAttention and CausalConvWithState for MTP (#29525, #31157).MatMulBlockQuantizedFp4Weight and MatMulBlockQuantizedFp8Weight, plus block-scaled tensor-core/GEMV decode paths, packed FP4 decode, M-tiling, and folded W8A8 activation QDQ (#29818, #29850, #29896, #31155, #31481).LinearAttentionGate, GatedRMSNorm, and GatedAdd contrib operators (#31158, #31835).AllReduce, AllGather, and AllToAll (#31571).GatherBlockQuantized (#31693).GatherBlockQuantized and LpNormalization, reused the shared WASM loader for Blob-backed external data, and fixed per-axis QDQ and MatMulNBits edge cases (#29475, #29801, #31151, #31152, #31197).M from the input tensor at compute time (#31189).Cos and int32 support to CPU Trilu, and fixed int8 QLinearSoftmax saturation and AvgPool ceil_mode/count_include_pad behavior (#28975, #29476, #29629, #29728).TfIdfVectorizer weight indexing and skipped MinLength logits-processor construction when eos_token_id is negative (#29604, #31649).GetOverridableInitializerNames() (#29349, #29589, #29616).Split axes now produce an error instead of being accepted (#31149).ceil_mode, allowed DFT to ignore excess input data, and fixed a webpack/Terser release-build crash (#29627, #29680, #31652).CUDA package architecture selections are now aligned across plugin EP, Python, C API, TensorRT, and Node.js pipelines. Windows arm64 is only available in CUDA plugin EP (#31992):
| OS | CUDA | CUDA architectures (all in -real form) |
|---|---|---|
| Linux x64 | 12.8 | 60;70;75;80;86;89;90a;120a |
| Linux x64 | 13.x | 75;80;86;89;90a;120a |
| Linux aarch64 | 13.x | 89;90a;120a;121a |
| Windows x64 | 12.8 | 61;75;86;89;120a |
| Windows x64 | 13.x | 75;80;86;89;120a |
| Windows arm64 | 13.x | 120a;121a |
Reduced CUDA compilation time and memory usage by splitting generated SM80 MoE, fpA_intB, and MatMulNBits translation units and adding two-level workspace estimation (#29614, #29699, #29811, #31834, #31837).
Fixed CUDA 13 plugin and packaging builds on Windows, Linux, and Windows ARM64, including MSVC/TMA compatibility and CI memory limits (#31608, #31609, #31615, #31616, #31617, #31622, #31729, #31748).
Fixed MLAS AVX2 builds on toolchains without AVX-VNNI assembler support, GCC 15 -Werror builds, and an MSVC C1001 issue in the W2 AVX-512-VNNI dispatch path (#28767, #29679, #29885).
Improved Windows compatibility by skipping DXGI discovery when Win32k system calls are unavailable and delay-loading shell32 (#29755, #30889).
Fixed Dawn parallel-build races, and GPU discovery in build/test environments (#29858, #29866).
Thanks to our 63 contributors for this release!
@adrastogi, @ahsan-ca, @AngelGalindo7, @ankitm3k, @apsonawane, @blazingphoenix7, @bmehta001, @chilo-ms, @claude, @daijh, @ducviet00, @edgchen1, @elwhyjay, @eserscor, @GopalakrishnanN, @guptaishaan, @hariharans29, @Honry, @huningxin, @jchen10, @jiafatom, @jiangzhuo, @Jiawei-Shao, @JonathanC-ARM, @justinchuby, @kjg0724, @kunal-vaishnavi, @kylo5aby, @Laan33, @martin-klacer-arm, @mastryukov1990, @mcollinswisc, @miaobin, @mingmingtasd, @mirounga, @mustjab, @n1harika, @namgyu-youn, @neilmsft, @nenad1002, @nicholascelestin, @OscarFree, @prathikr, @qjia7, @quic-muchhsu, @Sammy-Dabbas, @sanaa-hamel-microsoft, @shiyi9801, @skottmckay, @tairenpiao, @TedThemistokleous, @the0cp, @tianleiwu, @titaiwangms, @velonica0, @wangw-1991, @wuisabel-gif, @xadupre, @xhcao, @xiaofeihan1, @xiaoyu-work, @yen-shi, @zlma7001
Full Changelog: v1.28.0...v1.29.0
Nothing published for this version
Nothing published for this version
CUDA 12 packages are deprecated, please move to CUDA 13 ASAP.
n.b. This release is targeting ONNX 1.21. ONNX 1.22 will be supported in ORT 1.28. n.b. This changelog was generated via LLM. Only the contributor list has been verified. As always, only trust the commit history.
SoftmaxCrossEntropyLoss via label bounds validation (#28004)OneHot input validation and output-size computation (#28014)Expand and capped constant-folding output sizes (#28055)Tile kernel (#28070)MaxpoolWithMask::Compute (#28223)BitShift UB for shift amounts greater than or equal to bit width (#28272)seqlens_k vs cos_cache) (#28277)WordConvEmbedding to prevent OOB reads (#28279)CropBase scale handling (#28399)torch.load() calls to weights_only=True (#28421)OrtEp::OnSessionInitializationEnd() callback (#28319)kOrtEpDevice_EpMetadataKey_OSDriverVersion example and docs (#28282)quantize_static (#28221)ActivationRestrictedAsymmetric quantization option (#28237)block_size attribute support to QDQ quantization (#28522)FusedAdam optimizer in ORT Training (#28233)ConvTranspose-22 support (#27710)Split (split attribute path) and scalar Gather indices (#28270, #28278)FusedConv, Identity, Ceil, Tile, Cast(bool), Sin, Cos, and GatherND support (#28289, #28293, #28595, #28596, #28598)next, postcss, tmp, qs, body-parser, and other npm packages) (#27705, #27894, #28304, #28547, #28644, #28683, [#28694]present_key/present_value outputs in GQA and Gemma4 support (#28242)Session and InferenceSession (#27802)sympy an optional runtime dependency (#28141)py.typed marker to the onnxruntime package (#28438)UserLoggingFunction is used (#28314)OrtModelEditorApi ownership transfer (#28123)allowzero=1 handling for chained zero-size tensors (#28455)Resize nearest-mode rounding bug for negative halfway values (#28345)Thanks to our 106 contributors for this release!
@adrastogi, @adrianlizarraga, @AIFrameworksIntegration, @AlekseiNikiforovIBM, @angelser, @angelserMS, @ankitm3k, @anzzraju1997-glitch, @apsonawane, @arajendra, @ayappanec, @bachelor-dou, @badranX, @baijumeswani, @BoarQing, @bopeng1234, @bsosnader, @cbourjau, @chilo-ms, @chunghow-qti, @chwarr, @Craigacp, @daijh, @derdeljan-msft, @dparikh79, @edgchen1, @elwhyjay, @ericcraw, @eserscor, @feich-ms, @fs-eire, @gaugarg-nv, @gblong1, @GopalakrishnanN, @gramalingam, @guschmue, @hariharans29, @HectorSVC, @intbf, @ishwar-raut1, @Jaswanth51, @jatinwadhwa921, @javier-intel, @jchen10, @jiafatom, @Jiawei-Shao, @jnagi-intel, @jocelyn-stericker, @JonathanC-ARM, @justinchuby, @jwludzik, @kevinch-nv, @Kotomi-Du, @kpkbandi, @KV2773, @Laan33, @lhrios, @maxwbuckley, @MayureshV1, @mdvoretc-intel, @mingyueliuh, @mklimenk, @mustjab, @n1harika, @nazanin-beheshti, @orlmon01, @prathikr, @preetha-intel, @psakhamoori, @qiurui144, @qjia7, @qti-ashwshan, @qti-hungjuiw, @qti-yuduo, @rajatmonga, @RajeevSekar, @Rishi-Dave, @rvandermeulen, @RyanMetcalfeInt8, @SamuelLess, @sanaa-hamel-microsoft, @sfatimar, @sgbihu, @shiyi9801, @simonbyrne, @skottmckay, @susbhere, @sushraja-msft, @tairenpiao, @TejalKhade28, @theHamsta, @tianleiwu, @titaiwangms, @umangb-09, @velonica0, @vraspar, @vthaniel, @wenqinI, @xadupre, @xenova, @xhan65, @xiaofeihan1, @xieofxie, @yuslepukhin, @ZackyLake, @zejianzhang1982, @zz002
Full Changelog: https://github.com/microsoft/onnxruntime/compare/v0.1.4...v1.27.0
SVM and TreeEnsemble bounds/security fixes ( #27950 , #27951 , #27952 , #27989 ).
n.b. The following was generated via LLM from Git history. Only the contributor list has been verified.
onnxruntime-<os>-<arch>-gpu_cuda13-<version>.<ext>.ort model loads (#28164).setattr configuration with an allowlist (#28083).roundingType -> outputShapeRounding (#28172).@tianleiwu, @yuslepukhin, @edgchen1, @vraspar, @hariharans29, @skottmckay, @eserscor, @xadupre, @sanaa-hamel-microsoft, @elwhyjay, @Rishi-Dave, @titaiwangms, @adrianlizarraga, @jatinwadhwa921, @jchen10, @Jiawei-Shao, @maxwbuckley, @preetha-intel, @qjia7, @qti-hungjuiw, @RajeevSekar, @umangb-09, @adrastogi, @akote123, @amd-genmingz, @ankitm3k, @apsonawane, @bachelor-dou, @baijumeswani, @bopeng1234, @chilo-ms, @chwarr, @Craigacp, @dccarmo, @derdeljan-msft, @ericcraw, @fdwr, @fs-eire, @gaugarg-nv, @gblong1, @GopalakrishnanN, @Honry, @intbf, @ishwar-raut1, @Jaswanth51, @javier-intel, @JonathanC-ARM, @julia-thorn, @justinchuby, @jwludzik, @Kevin-Taha, @Kotomi-Du, @MayureshV1, @mdvoretc-intel, @miaobin, @milpuz01, @mingyueliuh, @mklimenk, @n1harika, @prathikr, @psakhamoori, @qti-yuduo, @quic-calvnguy, @RyanMetcalfeInt8, @sfatimar, @sgbihu, @ShirasawaSama, @ssam18, @susbhere, @sushraja-msft, @TejalKhade28, @theHamsta, @TomCrypto, @TsofnatMaman, @velonica0, @vthaniel, @wenqinI, @xhan65, @xhcao
n.b. This changelog is LLM generated. Only the contributor listing has been verified.
n.b. This changelog is LLM generated. Only the contributor listing has been verified.
SetRawDataInTensorProto in NVIDIA TensorRT RTX tests (#28065)Thanks to our 7 contributors for this release: @guschmue, @sanaa-hamel-microsoft, @apsonawane, @eserscor, @ishwar-raut1, @qjia7, @theHamsta
Full Changelog: https://github.com/microsoft/onnxruntime/compare/v1.25.0...v1.25.1
This is a patch release for ONNX Runtime 1.24, containing bug fixes, security improvements, performance enhancements, and execution provider updates.
This is a patch release for ONNX Runtime 1.24, containing bug fixes, security improvements, performance enhancements, and execution provider updates.
OrtEnv.DisableDllImportResolver to prevent fatal error on resolver conflict. (#27535)s_kernel_registry_vitisaiep.reset() in deinitialize_vitisai_ep(). (#27295)OrtEpDevice instances for plugin and provider bridge EPs. (#27522)-Warray-bounds build error in MLAS on clang 17+. (#27499)kMaxValueLength to 8192. (#27521)Full Changelog: v1.24.2...v1.24.3
@tianleiwu, @fs-eire, @adrianlizarraga, @yuslepukhin, @0-don, @anujj, @chaya2350, @chilo-ms, @dabhattimsft, @edgchen1, @eserscor, @hariharans29, @JonathanC-ARM, @lukas-folle-snkeos, @patryk-kaiser-ARM, @praneshgo, @skottmckay, @theHamsta, @vektah, @vishalpandya1990, @vthaniel, @xieofxie, @zz002
Security: Fixed an out-of-bounds read vulnerability in ArrayFeatureExtractor.
This is a patch release for ONNX Runtime 1.24, containing several bug fixes, security improvements, and execution provider updates.
SparseTensorProtoToDenseTensorProto to improve robustness. (#27323)ArrayFeatureExtractor. (#27275)BaseTester to support plugin EPs with both compiled nodes and registered kernels. (#27176)Full Changelog: v1.24.1...v1.24.2
@tianleiwu, @hariharans29, @edgchen1, @xiaofeihan1, @adrianlizarraga, @angelser, @angelserMS, @ankitm3k, @baijumeswani, @bmehta001, @ericcraw, @eserscor, @fs-eire, @guschmue, @mc-nv, @qjia7, @qti-monumeen, @titaiwangms, @yuslepukhin
DoS vulnerability in FuseReluClip
A major infrastructure enhancement enabling plugin-based EPs with dynamic loading:
OrtKernelInfo APIs for kernel-based plugin EPs (#26803)OrtApi::CreateEnvWithOptions() and OrtEpApi::GetEnvConfigEntries() (#26971)KernelInfo (#26589)Arm is formally deprecating the Arm NN Execution Provider (EP) in ONNX Runtime. The Arm NN EP is still experimental and depends on technology that is no longer actively maintained. Keeping it available now only adds complexity and potential confusion for users.
What to expect:
--enable_arm_neon_nchwc to enable this feature (#25580 #26838 #26691 #26171). This feature may be turned ON by default in a future release based on community feedback.SiLU activation perf improvement (#26753)add_external_initializers_from_files (#26012)FuseReluClip (#26878)KernelContext_GetAllocator (#26883)Thanks to our 170 contributors for this release!
@fs-eire, @tianleiwu, @edgchen1, @qjia7, @yuslepukhin, @hariharans29, @Honry, @qti-yuduo, @adrianlizarraga, @snnn, @eserscor, @vraspar, @xiaofeihan1, @guschmue, @daijh, @quic-muchhsu, @qti-jkilpatrick, @tirupath-qti, @Jiawei-Shao, @qti-hungjuiw, @quic-ashwshan, @titaiwangms, @qti-mattsinc, @chilo-ms, @jchen10, @xhcao, @skottmckay, @quic-calvnguy, @JonathanC-ARM, @Rohanjames1997, @sushraja-msft, @jambayk, @adrastogi, @xenova, @quic-tirupath, @justinchuby, @HectorSVC, @kunal-vaishnavi, @wenqinI, @prathikr, @baijumeswani, @preetha-intel, @jatinwadhwa921, @umangb-09, @qti-ashwshan, @carzh, @bachelor-dou, @ranjitshs, @gedoensmax, @xadupre, @nenad1002, @TedThemistokleous, @keshavv27, @zpye, @jnagi-intel, @jiafatom, @mingyueliuh, @Colm-in-Arm, @borg323, @chunghow-qti, @Craigacp, @BODAPATIMAHESH, @AlekseiNikiforovIBM, @hans00, @thevishalagarwal, @MaanavD, @qti-kromero, @damdoo01-arm, @BoarQing, @naomiOvad, @yuhuchua-qti, @hadiFute, @vishalpandya1990, @rivkastroh, @minfhong-qti, @kuanyul-qti, @xieofxie, @ankitm3k, @RyanMetcalfeInt8, @MayureshV1, @bopeng1234, @vthaniel, @mdvoretc-intel, @ericcraw, @javier-intel, @saurabhkale17, @sfatimar, @Kotomi-Du, @intbf, @n1harika, @TejalKhade28, @gupta-pallavi, @cbourjau, @nieubank, @r-devulap, @wszqkzqk, @sanketkaleoss, @amancini-N, @fanchenkong1, @meakbiyik, @hisham-hchowdhu, @shaoboyan091, @Stonesjtu, @qwu16, @wangw-1991, @bonktree, @naetherm, @nikhilfujitsu, @Panxuefeng-loongson, @selenayang888, @moyo1997, @chwarr, @patryk-kaiser-ARM, @fdwr, @SavaLione, @shiyi9801, @mcost45, @aciddelgado, @prudhvi-qti, @Jonahcb, @lifang-zhang, @zhaoxul-qti, @gaugarg-nv, @cocotdf, @WangFengtu1996, @orlmon01, @weidu-tpvision, @theHamsta, @kevinch-nv, @XXXXRT666, @movedancer, @melkap01-Arm, @KingSora, @urpetkov-amd, @junchao-loongson, @jixiongdeng, @wcy123, @GrigoryEvko, @anujj, @peishenyan, @quic-ankus, @jchen351, @yihonglyu, @satyajandhyala, @co63oc, @mschofie, @quic-ashigarg, @asoldano, @nproshun, @jiangzhaoming, @seungtaek94, @liqunfu, @jaholme, @hanbitmyths, @quic-boyuc, @rM-planet, @qti-vaiskv, @AndreyOrb, @pkubaj, @xhan65, @Jaswanth51, @quic-hungjuiw, @jywu-msft, @mklimenk, @derdeljan-msft, @ianfhunter, @NingW101, @feich-ms, @Akupadhye, @wschin
Full Changelog: v1.23.2...rel-1.24.1
Nothing published for this version
Nothing published for this version
Nothing published for this version
This release introduces Execution Provider (EP) Plugin API, which is a new infrastructure for building plugin-based EPs. (#24887 , #25137, #25124, #25
This release introduces Execution Provider (EP) Plugin API, which is a new infrastructure for building plugin-based EPs. (#24887 , #25137, #25124, #25147, #25127, #25159, #25191, #2524)
This release introduces the ability to dynamically download and install execution providers. This feature is exclusively available in the WinML build and requires Windows 11 version 25H2 or later. To leverage this new capability, C/C++/C# users should use the builds distributed through the Windows App SDK, and Python users should install the onnxruntime-winml package(will be published soon). We encourage users who can upgrade to the latest Windows 11 to utilize the WinML build to take advantage of this enhancement.
Now on Windows some global object will be not destroyed if we detect that the process is being shutting down(#24891) . It will not cause memory leak as when a process ends all the memory will be returned to the operating system. This change can reduce the chance of having crashes on process exit.
Now ONNX Runtime has the ability to automatically discovery computing devices and select the best EPs to download and register. The EP downloading feature currently only works on Windows 11 version 25H2 or later.
ROCM EP was removed from the source tree. Users are recommended to use Migraphx or Vitis AI EPs from AMD. A new EP, Nvidia TensorRT RTX, was added.
EMDSK is upgraded from 4.0.4 to 4.0.8
Added WGSL template support.
SDK Update: Added support for QNN SDK 2.37.
Enhanced performance for SGEMM, IGEMM, and Dynamic Quantized MatMul operations, especially for Conv2D operators on hardware that supports SME2 (Scalable Matrix Extension v2).
Contributors to ONNX Runtime include members across teams at Microsoft, along with our community members:
@1duo, @Akupadhye, @amarin16, @AndreyOrb, @ankan-ban, @ankitm3k, @anujj, @aparmp-quic, @arnej27959, @bachelor-dou, @benjamin-hodgson, @Bonoy0328, @chenweng-quic, @chuteng-quic, @clementperon, @co63oc, @daijh, @damdoo01-arm, @danyue333, @fanchenkong1, @gedoensmax, @genarks, @gnedanur, @Honry, @huaychou, @ianfhunter, @ishwar-raut1, @jing-bao, @joeyearsley, @johnpaultaken, @jordanozang, @JulienMaille, @keshavv27, @kevinch-nv, @khoover, @krahenbuhl, @kuanyul-quic, @mauriciocm9, @mc-nv, @minfhong-quic, @mingyueliuh, @MQ-mengqing, @NingW101, @notken12, @omarhass47, @peishenyan, @pkubaj, @qc-tbhardwa, @qti-jkilpatrick, @qti-yuduo, @quic-ankus, @quic-ashigarg, @quic-ashwshan, @quic-calvnguy, @quic-hungjuiw, @quic-tirupath, @qwu16, @ranjitshs, @saurabhkale17, @schuermans-slx, @sfatimar, @stefantalpalaru, @sunnyshu-intel, @TedThemistokleous, @thevishalagarwal, @toothache, @umangb-09, @vatlark, @VishalX, @wcy123, @xhcao, @xuke537, @zhaoxul-qti
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
This release introduces new API's for Model Editor, Auto EP infrastructure, and AOT Compile
MatMulNBits, enabling matrix multiplication with weights quantized to 8 bits.MatMulNBits.Contributors to ONNX Runtime include members across teams at Microsoft, along with our community members:
Yulong Wang, Jian Chen, Changming Sun, Satya Kumar Jandhyala, Hector Li, Prathik Rao, Adrian Lizarraga, Jiajia Qin, Scott McKay, Jie Chen, Tianlei Wu, Edward Chen, Wanming Lin, xhcao, vraspar, Dmitri Smirnov, Jing Fang, Yifan Li, Caroline Zhu, Jianhui Dai, Chi Lo, Guenther Schmuelling, Ryan Hill, Sushanth Rajasankar, Yi-Hong Lyu, Ankit Maheshkar, Artur Wojcik, Baiju Meswani, David Fan, Enrico Galli, Hans, Jambay Kinley, John Paul, Peishen Yan, Yateng Hong, amarin16, chuteng-quic, kunal-vaishnavi, quic-hungjuiw, Alessio Soldano, Andreas Hussing, Ashish Garg, Ashwath Shankarnarayan, Chengdong Liang, Clément Péron, Erick Muñoz, Fanchen Kong, George Wu, Haik Silm, Jagadish Krishnamoorthy, Justin Chu, Karim Vadsariya, Kevin Chen, Mark Schofield, Masaya, Kato, Michael Tyler, Nenad Banfic, Ningxin Hu, Praveen G, Preetha Veeramalai, Ranjit Ranjan, Seungtaek Kim, Ti-Tai Wang, Xiaofei Han, Yueqing Zhang, co63oc, derdeljan-msft, genmingz@AMD, jiangzhaoming, jing-bao, kuanyul-quic, liqun Fu, minfhong-quic, mingyue, quic-tirupath, quic-zhaoxul, saurabh, selenayang888, sfatimar, sheetalarkadam, virajwad, zz002, Ștefan Talpalaru
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Extend CMAKE_CUDA_FLAGS with all Blackwell compute capacity #23928 - @yf711
Chat mode introduced breaking changes in the API (see migration guide).
top_k on CPU.NMS, RoiAlign, NonZero) to TensorRT by default.trt_op_types_to_exclude to exclude specific ops from TensorRT assignment.--use_qnn static_lib.SkipLayerNormalization, MatMulNBits, FusedGemm, FusedConv, EmbedLayerNormalization, BiasGelu, Attention, DynamicQuantizeMatMul, FusedMatMul, QuickGelu, SkipSimplifiedLayerNormalizationChatGLM, Baichuan2, Phi-4, etc.Phi-4 pre/post-processing support for text, vision, and audio.tokenizer.json.ImageCodec now links to native APIs if available; otherwise, falls back to built-in libraries.All the prebuilt Windows packages now require VC++ Runtime version >= 14.40(instead of 14.38). If your VC++ runtime version is lower than that, you may see a crash when ONNX Runtime was initializing. See https://github.com/microsoft/STL/wiki/Changelog#vs-2022-1710 for more details.
Updated minimum iOS and Android SDK requirements to align with React Native 0.76:
All macOS packages now require macOS version >= 13.3.
CMake Version: Increased the minimum required CMake version from 3.26 to 3.28. Added support for CMake 4.0. Python Version: Increased the minimum required Python version from 3.8 to 3.10 for building ONNX Runtime from source. Improved VCPKG support
Added the following cmake options for WebGPU EP
Added cmake option onnxruntime_BUILD_QNN_EP_STATIC_LIB for building with QNN EP as a static library. Removed cmake option onnxruntime_USE_PREINSTALLED_EIGEN.
Fixed a build issue with Visual Studio 2022 17.3 (#23911)
onnxruntime_USE_CUDA_NHWC_OPS by default for CUDA builds.nsync from dependencies.Updated Node.js installation script to support network proxy usage (#23231)
Contributors to ONNX Runtime include members across teams at Microsoft, along with our community members:
Changming Sun, Yulong Wang, Tianlei Wu, Jian Chen, Wanming Lin, Adrian Lizarraga, Hector Li, Jiajia Qin, Yifan Li, Edward Chen, Prathik Rao, Jing Fang, shiyi, Vincent Wang, Yi Zhang, Dmitri Smirnov, Satya Kumar Jandhyala, Caroline Zhu, Chi Lo, Justin Chu, Scott McKay, Enrico Galli, Kyle, Ted Themistokleous, dtang317, wejoncy, Bin Miao, Jambay Kinley, Sushanth Rajasankar, Yueqing Zhang, amancini-N, ivberg, kunal-vaishnavi, liqun Fu, Corentin Maravat, Peishen Yan, Preetha Veeramalai, Ranjit Ranjan, Xavier Dupré, amarin16, jzm-intel, kailums, xhcao, A-Satti, Aleksei Nikiforov, Ankit Maheshkar, Javier Martinez, Jianhui Dai, Jie Chen, Jon Campbell, Karim Vadsariya, Michael Tyler, PARK DongHa, Patrice Vignola, Pranav Sharma, Sam Webster, Sophie Schoenmeyer, Ti-Tai Wang, Xu Xing, Yi-Hong Lyu, genmingz@AMD, junchao-zhao, sheetalarkadam, sushraja-msft, Akshay Sonawane, Alexis Tsogias, Ashrit Shetty, Bilyana Indzheva, Chen Feiyue, Christian Larson, David Fan, David Hotham, Dmitry Deshevoy, Frank Dong, Gavin Kinsey, George Wu, Grégoire, Guenther Schmuelling, Indy Zhu, Jean-Michaël Celerier, Jeff Daily, Joshua Lochner, Kee, Malik Shahzad Muzaffar, Matthieu Darbois, Michael Cho, Michael Sharp, Misha Chornyi, Po-Wei (Vincent), Sevag H, Takeshi Watanabe, Wu, Junze, Xiang Zhang, Xiaoyu, Xinpeng Dou, Xinya Zhang, Yang Gu, Yateng Hong, mindest, mingyue, raoanag, saurabh, shaoboyan091, sstamenk, tianf-fff, wonchung-microsoft, xieofxie, zz002
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Prevent int32 quantized bias from clipping by adjusting the weight's scale (#22020) - @adrianlizarraga
Big thank you to the release manager @yf711, along with @adrianlizarraga, @HectorSVC, @jywu-msft, and everyone else who helped to make this patch release process a smooth one!
All ONNX Runtime Training packages have been deprecated. ORT 1.19.2 was the last release for which onnxruntime-training (PyPI), onnxruntime-training-c…
Release Manager: @apsonawane
Full release notes for ONNX Runtime generate() API v0.5.0 can be found here.
Full release notes for ONNX Runtime Extensions v0.13 can be found here.
Big thank you to the release manager @apsonawane, as well as @snnn, @jchen351, @sheetalarkadam, and everyone else who made this release possible!
Tianlei Wu, Yi Zhang, Yulong Wang, Scott McKay, Edward Chen, Adrian Lizarraga, Wanming Lin, Changming Sun, Dmitri Smirnov, Jian Chen, Jiajia Qin, Jing Fang, George Wu, Caroline Zhu, Hector Li, Ted Themistokleous, mindest, Yang Gu, jingyanwangms, liqun Fu, Adam Pocock, Patrice Vignola, Yueqing Zhang, Prathik Rao, Satya Kumar Jandhyala, Sumit Agarwal, Xu Xing, aciddelgado, duanshengliu, Guenther Schmuelling, Kyle, Ranjit Ranjan, Sheil Kumar, Ye Wang, kunal-vaishnavi, mingyueliuh, xhcao, zz002, 0xdr3dd, Adam Reeve, Arne H Juul, Atanas Dimitrov, Chen Feiyue, Chester Liu, Chi Lo, Erick Muñoz, Frank Dong, Jake Mathern, Julius Tischbein, Justin Chu, Xavier Dupré, Yifan Li, amarin16, anujj, chenduan-amd, saurabh, sfatimar, sheetalarkadam, wejoncy, Akshay Sonawane, AlbertGuan9527, Bin Miao, Christian Bourjau, Claude, Clément Péron, Emmanuel, Enrico Galli, Fangjun Kuang, Hann Wang, Indy Zhu, Jagadish Krishnamoorthy, Javier Martinez, Jeff Daily, Justin Beavers, Kevin Chen, Krishna Bindumadhavan, Lennart Hannink, Luis E. P., Mauricio A Rovira Galvez, Michael Tyler, PARK DongHa, Peishen Yan, PeixuanZuo, Po-Wei (Vincent), Pranav Sharma, Preetha Veeramalai, Sophie Schoenmeyer, Vishnudas Thaniel S, Xiang Zhang, Yi-Hong Lyu, Yufeng Li, goldsteinn, mcollinswisc, mguynn-intc, mingmingtasd, raoanag, shiyi, stsokolo, vraspar, wangshuai09
Full changelog: v1.19.2...v1.20.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
ORT 1.19.2 is a small patch release, fixing some broken workflows and introducing bug fixes.
@prathikr, @mszhanyi, @edgchen1, @tianleiwu, @wangyems, @aciddelgado, @mindest, @snnn, @baijumeswani, @MaanavD
Thanks to everyone who helped ship this release smoothly!
Full Changelog: https://github.com/microsoft/onnxruntime/compare/v1.19.0...v1.19.2
Remove calls to deprecated api’s
TensorRT
CUDA
CPU
QNN
OpenVINO
DirectML
Changming Sun, Baiju Meswani, Scott McKay, Edward Chen, Jian Chen, Wanming Lin, Tianlei Wu, Adrian Lizarraga, Chester Liu, Yi Zhang, Yulong Wang, Hector Li, kunal-vaishnavi, pengwa, aciddelgado, Yifan Li, Xu Xing, Yufeng Li, Patrice Vignola, Yueqing Zhang, Jing Fang, Chi Lo, Dmitri Smirnov, mingyueliuh, cloudhan, Yi-Hong Lyu, Ye Wang, Ted Themistokleous, Guenther Schmuelling, George Wu, mindest, liqun Fu, Preetha Veeramalai, Justin Chu, Xiang Zhang, zz002, vraspar, kailums, guyang3532, Satya Kumar Jandhyala, Rachel Guo, Prathik Rao, Maximilian Müller, Sophie Schoenmeyer, zhijiang, maggie1059, ivberg, glen-amd, aamajumder, Xavier Dupré, Vincent Wang, Suryaprakash Shanmugam, Sheil Kumar, Ranjit Ranjan, Peishen Yan, Frank Dong, Chen Feiyue, Caroline Zhu, Adam Louly, Ștefan Talpalaru, zkep, winskuo-quic, wejoncy, vividsnow, vivianw-amd, moyo1997, mcollinswisc, jingyanwangms, Yang Gu, Tom McDonald, Sunghoon, Shubham Bhokare, RuomeiMS, Qingnan Duan, PeixuanZuo, Pavan Goyal, Nikolai Svakhin, KnightYao, Jon Campbell, Johan MEJIA, Jake Mathern, Hans, Hann Wang, Enrico Galli, Dwayne Robinson, Clément Péron, Chip Kerchner, Chen Fu, Carson M, Adam Reeve, Adam Pocock.
Big thank you to everyone who contributed to this release!
Full Changelog: https://github.com/microsoft/onnxruntime/compare/v1.18.1...v1.19.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
…iOS cocoapods are being deprecated. Please use the onnxruntime-android Android package, and onnxruntime-c/onnxruntime-objc cocoapods, which support ON…
--rv64, --riscv_toolchain_root, and --riscv_qemu_path.--use_binskim_compliant_compile_flags. Note: All our release binaries are built with this flag, but when building ONNX Runtime from source, this flag is default OFF.TensorRT
CUDA
QNN
OpenVINO
AUTO:GPU,CPU will only create GPU blob, not CPU blob.DirectML
Yi Zhang, Yulong Wang, Adrian Lizarraga, Changming Sun, Scott McKay, Tianlei Wu, Peng Wang, Hector Li, Edward Chen, Dmitri Smirnov, Patrice Vignola, Guenther Schmuelling, Ye Wang, Chi Lo, Wanming Lin, Xu Xing, Baiju Meswani, Peixuan Zuo, Vincent Wang, Markus Tavenrath, Lei Cao, Kunal Vaishnavi, Rachel Guo, Satya Kumar Jandhyala, Sheil Kumar, Yifan Li, Jiajia Qin, Maximilian Müller, Xavier Dupré, Yi-Hong Lyu, Yufeng Li, Alejandro Cid Delgado, Adam Louly, Prathik Rao, wejoncy, Zesong Wang, Adam Pocock, George Wu, Jian Chen, Justin Chu, Xiaoyu, guyang3532, Jingyan Wang, raoanag, Satya Jandhyala, Hariharan Seshadri, Jiajie Hu, Sumit Agarwal, Peter Mcaughan, Zhijiang Xu, Abhishek Jindal, Jake Mathern, Jeff Bloomfield, Jeff Daily, Linnea May, Phoebe Chen, Preetha Veeramalai, Shubham Bhokare, Wei-Sheng Chin, Yang Gu, Yueqing Zhang, Guangyun Han, inisis, ironman, Ivan Berg, Liqun Fu, Yu Luo, Rui Ren, Sahar Fatima, snadampal, wangshuai09, Zhenze Wang, Andrew Fantino, Andrew Grigorev, Ashwini Khade, Atanas Dimitrov, AtomicVar, Belem Zhang, Bowen Bao, Chen Fu, Dhruv Matani, Fangrui Song, Francesco, Frank Dong, Hans Chen, He Li, Heflin Stephen Raj, Jambay Kinley, Masayoshi Tsutsui, Matttttt, Nanashi, Phoebe Chen, Pranav Sharma, Segev Finer, Sophie Schoenmeyer, TP Boudreau, Ted Themistokleous, Thomas Boby, Xiang Zhang, Yongxin Wang, Zhang Lei, aamajumder, danyue, Duansheng Liu, enximi, fxmarty, kailums, maggie1059, mindest, mo-ja, moyo1997 Big thank you to everyone who contributed to this release!
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Update copying API header files to make Linux logic consistent with Windows (#19736) - @mszhanyi
General:
Build System & Packages:
Core:
CUDA EP:
TensorRT EP:
Web:
Windows:
Kernel Optimizations:
Models:
This patch release also includes additional fixes by @spampana95 and @enximi. Big thank you to all our contributors!
This patch release includes the following updates:
This patch release includes the following updates:
Updated calls to deprecated TensorRT APIs (e.g., enqueue_v2 → enqueue_v3).
In the next release, we will totally drop support for Windows ARM32.
python -m pip install cerberus flatbuffers h5py numpy>=1.16.6 onnx packaging protobuf sympy setuptools>=41.4.0
pip install -i https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/ORT/pypi/simple/ onnxruntime-training
pip install torch-ort
python -m torch_ort.configure
Installation instructions can also be accessed here.Contributors to ONNX Runtime include members across teams at Microsoft, along with our community members: Changming Sun, Yulong Wang, Tianlei Wu, Yi Zhang, Jian Chen, Jiajia Qin, Adrian Lizarraga, Scott McKay, Wanming Lin, pengwa, Hector Li, Chi Lo, Dmitri Smirnov, Edward Chen, Xu Xing, satyajandhyala, Rachel Guo, PeixuanZuo, RandySheriffH, Xavier Dupré, Patrice Vignola, Baiju Meswani, Guenther Schmuelling, Jeff Bloomfield, Vincent Wang, cloudhan, zesongw, Arthur Islamov, Wei-Sheng Chin, Yifan Li, raoanag, Caroline Zhu, Sheil Kumar, Ashwini Khade, liqun Fu, xhcao, aciddelgado, kunal-vaishnavi, Aditya Goel, Hariharan Seshadri, Ye Wang, Adam Pocock, Chen Fu, Jambay Kinley, Kaz Nishimura, Maximilian Müller, Yang Gu, guyang3532, mindest, Abhishek Jindal, Justin Chu, Numfor Tiapo, Prathik Rao, Yufeng Li, cao lei, snadampal, sophies927, BoarQing, Bowen Bao, George Wu, Jiajie Hu, MistEO, Nat Kershaw (MSFT), Sumit Agarwal, Ted Themistokleous, ivberg, zhijiang, Christian Larson, Frank Dong, Jeff Daily, Nicolò Lucchesi, Pranav Sharma, Preetha Veeramalai, Cheng Tang, Xiang Zhang, junchao-loongson, petermcaughan, rui-ren, shaahji, simonjub, trajep, Adam Louly, Akshay Sonawane, Artem Shilkin, Atanas Dimitrov, AtanasDimitrovQC, BODAPATIMAHESH, Bart Verhagen, Ben Niu, Benedikt Hilmes Brian Lambert, David Justice, Deoksang Kim, Ella Charlaix, Emmanuel Ferdman, Faith Xu, Frank Baele, George Nash, hans00, computerscienceiscool, Jake Mathern, James Baker, Jiangzhuo, Kevin Chen, Lennart Hannink, Lukas Berbuer, Mike Guo, Milos Puzovic, Mustafa Ateş Uzun, Peishen Yan, Ran Gal, Ryan Hill, Steven Roussey, Suryaprakash Shanmugam, Vadym Stupakov, Yiming Hu, Yueqing Zhang, Yvonne Chen, Zhang Lei, Zhipeng Han, aimilefth, gunandrose4u, kailums, kushalpatil07, kyoshisuki, luoyu-intel, moyo1997, tbqh, weischan-quic, wejoncy, winskuo-quic, wirthual, yuwenzho
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →