NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
NuGet · #2167 most downloaded on NuGet
This package contains native shared library artifacts for all supported platforms of ONNX Runtime.
Last release 26 days ago
10 Sep 2026
Release timing varies
gaps range from 2 weeks to 3 months
Rarely documented
notes for 13 of the last 60 stable releases
4 versions withdrawn
withdrawn after publishing
127 years old
68 releases · first in 1900
Nothing published for this version
Nothing published for this version
This release brings kernel performance improvements, expanded operator support, and reliability fixes to the ONNX Runtime WebGPU Plugin EP.
This release brings kernel performance improvements, expanded operator support, and reliability fixes to the ONNX Runtime WebGPU Plugin EP.
MatMulNBits wide-tile execution using subgroup shuffle (#31703)Split when every output segment is vec4-aligned (#32251)EngramGate, NGramHashMapping) (#32106, #32268)MatMulNaiveProgram's pipeline cache key cover everything it bakes into WGSL (#32048)copy_tensors misuse instead of terminating the process, and rejected foreign GPU handles in the built-in data transfer (#32315, #32317)onnxruntime_providers compilation from its archive (#32162)wgsl_template Python tests in CI (#32214)Thanks to everyone who contributed to this release:
@4n4ny4, @daijh, @edgchen1, @hariharans29, @jchen10, @kunal-vaishnavi, @qjia7, @rvandermeulen, @sushraja-msft, @tianleiwu, @xhcao, @xiaofeihan1
This release covers commits affecting WebGPU Plugin EP code and packaging.
This summary was drafted with AI assistance from commit history and PR metadata, and reviewed before publishing.
One column per quarter.
Nothing published for this version
ONNX Runtime WebGPU Plugin EP 0.3.0 expands model and data-type coverage, improves generative-model performance, and strengthens configuration, reliab
ONNX Runtime WebGPU Plugin EP 0.3.0 expands model and data-type coverage, improves generative-model performance, and strengthens configuration, reliability, and release tooling.
These release notes were drafted with AI assistance.
int64 for Add, Cast, Clip, Concat, Equal, Gather, Min, Max, ReduceSum, Reshape, Sub, Tile, and Where; uint8 for Cast, Expand, Gather, and Reshape; and int32/uint32 for CumSum and Tile. (#28804, #29392, #29830, #29834, #29839, #29844, #29847, #29854, #29861, #29897, #31049, #31702, #31709, #31714)GatherBlockQuantized support and integrated ONNX 1.22 with opset 27. (#29054, #28754)GatherBlockQuantized, and fixed WebGPU data-transfer callbacks on Windows x86. (#28704, #29030, #29255, #29595, #31568)Thank you to everyone who contributed to this release:
@AngelGalindo7, @daijh, @danielsongmicrosoft, @edgchen1, @fanchenkong1, @feich-ms, @guschmue, @haoxli, @hariharans29, @Honry, @huningxin, @jchen10, @Jiawei-Shao, @miaobin, @mingmingtasd, @mirounga, @mustjab, @nicholascelestin, @prathikr, @qjia7, @Reranko05, @Shivani767, @skottmckay, @ssam18, @sushraja-msft, @tairenpiao, @tianleiwu, @titaiwangms, @wuisabel-gif, @xhcao, and @xiaofeihan1.
Scope: commits affecting ONNX Runtime WebGPU Plugin EP code, tests, build integration, and packaging since plugin-ep-webgpu/v0.2.1.
Major performance work for attention-heavy LLMs.
Major performance work for attention-heavy LLMs.
max_k_step was enabled for NVIDIA (#28511).Qwen3 and Gemma 4 model-path improvements.
LinearAttention and quantized-path optimizations.
Reliability and hardening fixes.
past_state == present_state buffer handling (#28753).Graph-capture and buffer-management improvements.
Note: This section was AI-generated. It may have inaccuracies.
Thanks to everyone who contributed to the WebGPU EP (human contributors, alphabetical):
@apsonawane, @daijh, @edgchen1, @feich-ms, @GopalakrishnanN, @guschmue, @hariharans29, @HectorSVC, @jchen10, @qjia7, @tianleiwu, @xiaofeihan1, @xenova, @yuslepukhin.
Note: This list was compiled on a best-effort basis from PRs that touched WebGPU EP-specific paths and
intentionally includes human contributors only, so it may not capture every contribution. If yours was
missed, the omission is unintentional. Your work is no less appreciated.
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →