NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
NuGet · #2167 most downloaded on NuGet
This package contains native shared library artifacts for all supported platforms of ONNX Runtime.
Last release 28 days ago
10 Sep 2026
Release timing varies
gaps range from 2 weeks to 3 months
Rarely documented
notes for 13 of the last 60 stable releases
4 versions withdrawn
withdrawn after publishing
127 years old
68 releases · first in 1900
Release Def: Branch: refs/heads/rel-1.30.0 Commit: f2c39fe2f838cf35ce7da92824f5a5e3ee6e88a7 Build: https://aiinfra.visualstudio.com/Lotus/_build/resul
Release Def: Branch: refs/heads/rel-1.30.0 Commit: f2c39fe2f838cf35ce7da92824f5a5e3ee6e88a7 Build: https://aiinfra.visualstudio.com/Lotus/_build/results?buildId=1412493
ONNX Runtime 1.30.0 expands generative AI inference, improves CPU and GPU performance, adds Go bindings, and strengthens runtime reliability. These notes cover changes since ONNX Runtime 1.29.1.
-Donnxruntime_USE_FP4_QMOE=OFF (#32096, #32163).block_size=32. Set -Donnxruntime_USE_FPA_INTB_GEMM_FULL=ON when building from source to retain the full kernel set, including BF16, zero-point, bias, larger-block-size, and native Hopper variants (#32324).Gemm and MatMul execution is gated on hardware acceleration. CPU-assigned FP16 nodes without a matching kernel now fall back to FP32 (#32301, #32197).Split, Scan, GatherND, ScatterND, SpaceToDepth/DepthToSpace, Crop, Conv, Normalizer, and pooling (#29461, #31668, #32034, #32039, #32076, #32157, #32160, #32161, #32345, #32349).BifurcationDetector inputs, generation subgraph shapes, and QEmbed segment inputs. BeamSearch buffer expansion now uses dynamic shape storage (#31648, #31701, #32009, #32078, #32144).TreeEnsemble node references and bounded subtree comparison, rejected non-finite CPU RoiAlign coordinates, and required ImageScaler bias to match the channel count (#32031, #32043, #32011, #32002).MatMulFpQ4 shape inputs, and checked MLAS blockwise quantization/dequantization index ranges (#31682, #32032, #32007).MatMulNBits, RemovePadding, RotaryEmbedding, SparseAttention, Whisper beam search, NMS, QDQ, and GatherElements (#31643, #31994, #31995, #31996, #31998, #32014, #32029, #32030).CudaAsyncBuffer staging storage alive across CUDA graph replay (#31968, #32121).js-yaml, joi, fast-uri, and the Next.js end-to-end fixture (#32397, #32486, #32488, #32505, #32508).KernelContext::GetPreallocatedOutput (#29726, #32089).EngramGate and NGramHashMapping, and expanded kernel coverage for Qwen-3.5 operators (#32268, #32106).is_causal attribute to PagedAttention (#32515, #32225).VarlenCausalConvWithState for continuous batching and compact variable-length causal-convolution state updates (#32168, #32290).GatedDeltaNet operator and BFloat16 support for CUDA GatedDeltaNet (#32282, #32307).ORT_QMOE_FP4_DEEPGEMM=1, default off). This path is disabled on Windows (#32122, #32485).MatMul shapes and refined fpA-intB GEMV support checks (#31478, #32338).TopK, ArgMax, and ArgMin performance for wide last axes, and accelerated low-lane INT64 CumSum (#32404, #32092, #32238).Slice fast path for contiguous subregions and removed pinned-buffer use from Split and Concat fast paths (#28902, #32410).ReduceMean kernels and fixed ScatterElements reduction dispatch by element type and signed-zero handling in Abs (#32326, #29879, #31477).Gather support and optimized MatMulNBits wide tiles with subgroup shuffle (#31714, #31703).Split when all output segments are vec4-aligned, and selected pooling paths based on occupancy (#29820, #32251, #32313).onnxruntime_perf_test -i options to WebGPU (#32074, #31971, #32316).GPUDevice instances, corrected MatMul pipeline-cache keys and the 1D-dispatch shader fast path, and changed copy_tensors misuse to report errors instead of terminating the process (#32259, #32048, #32343, #32315).SkipLayerNormalization support and corrected output-rank validation and fallback data-type support checks (#32377, #31708, #32067, #32293).kernel_shape and output_padding lengths during DirectML kernel setup (#31999).QuantizeLinear rounding, prevented CPU TensorScatter index overflow, serialized ScatterND string updates, and widened Compress loop counters (#32452, #32012, #32033, #32008).LpNormalization inputs, zero-element BiasGelu/FastGelu, extreme Trilu diagonals, and empty reduction axes (#32020, #31698, #32013, #32156).Slice starts rank (#31670, #31678, #32018, #32044).Gemm (#32038, #32143, #32426, #32435).NodeAttrHelper string-default lifetimes (#32051, #32138, #32019).Env, and clarified how external-initializer paths interact with EP context paths (#32502, #32503, #32442).RunAsync arguments until completion (#32041, #32015).Conv3DNaive shader (#32469, #32357).Thanks to our 57 human contributors for this release!
@4n4ny4, @apsonawane, @arnej27959, @baijumeswani, @bmehta001, @chilo-ms, @crvineeth97, @daijh, @danfiedler-msft, @danielsongmicrosoft, @dannyota, @DKAIN-py, @edgchen1, @ericcraw, @eserscor, @fanchenkong1, @hanbitmyths, @hariharans29, @hdharpure9922, @Honry, @jambayk, @javier-intel, @jchen10, @jiafatom, @jnagi-intel, @justinchuby, @kadyrbekovhamit-cyber, @kunal-vaishnavi, @Lapis0x0, @LOGO127, @Manogna-Sree, @martin-klacer-arm, @mei1127, @miaobin, @mirounga, @miyanyan, @MohamedElashri, @mustjab, @Nikhi00718, @Noperi0r, @Novestars, @pkubaj, @preetha-intel, @qjia7, @rvandermeulen, @sanaa-hamel-microsoft, @skottmckay, @sushraja-msft, @swetha097, @sylvesterkaczmarek, @tianleiwu, @titaiwangms, @toothache, @xadupre, @xhcao, @xiaofeihan1, @Zestion
Full Changelog: rel-1.29.1...rel-1.30.0
Release highlights were prepared with AI assistance.
One column per quarter.
Release Def: Branch: refs/heads/rel-1.29.0 Commit: 2e2543fbe9fae542f921d47a72d21d5a4ef0b710 Build: https://aiinfra.visualstudio.com/Lotus/_build/resul
Release Def: Branch: refs/heads/rel-1.29.0 Commit: 2e2543fbe9fae542f921d47a72d21d5a4ef0b710 Build: https://aiinfra.visualstudio.com/Lotus/_build/results?buildId=1354516
ORT_DISABLE_TELEMETRY=1 before initialization disables non-Windows telemetry for the process (#27379, #29872).onnxruntime/python/tools/tensorrt dashboard tooling was removed. This does not affect the TensorRT Execution Provider APIs (#29395).k attribute against the number of experts and fixed a CPU TensorScatter security issue (#29907, #29916).Range, and CropAndResize (#29254, #29255, #29265, #29579, #29595, #29605, #29871, #31636, #31671, #31675, #31676, #31684).OrtApi::GetValue and validated DML constant tensor byte sizes (#29157, #31665).adm-zip for onnxruntime-node (#29827, #29926, #31192).ORT_INTRA_OP_NUM_THREADS and ORT_INTER_OP_NUM_THREADS. Explicit thread settings still take precedence, and 0 preserves machine-sized defaults (#29688).EpContext nodes, and wired maximum-shape inference into workspace estimation (#29607, #29799, #31613).MRotaryEmbedding contrib operator for Qwen mRoPE variants (#29261, #31728).onnxruntime_perf_test through --data_shape, plus verbose graph-transformer tracing and broader inference-session error-path coverage (#29555, #29558, #29569, #29571).PagedAttention with quantized KV cache, XQA decode, MLA, QK-Norm, and head-sink support (#29912).Attention CUDA kernel and enabled cuDNN SDPA for contrib Attention (#29715, #29717).attention_bias support to the GroupQueryAttention unfused path and state_window support to LinearAttention and CausalConvWithState for MTP (#29525, #31157).MatMulBlockQuantizedFp4Weight and MatMulBlockQuantizedFp8Weight, plus block-scaled tensor-core/GEMV decode paths, packed FP4 decode, M-tiling, and folded W8A8 activation QDQ (#29818, #29850, #29896, #31155, #31481).LinearAttentionGate, GatedRMSNorm, and GatedAdd contrib operators (#31158, #31835).AllReduce, AllGather, and AllToAll (#31571).GatherBlockQuantized (#31693).GatherBlockQuantized and LpNormalization, reused the shared WASM loader for Blob-backed external data, and fixed per-axis QDQ and MatMulNBits edge cases (#29475, #29801, #31151, #31152, #31197).M from the input tensor at compute time (#31189).Cos and int32 support to CPU Trilu, and fixed int8 QLinearSoftmax saturation and AvgPool ceil_mode/count_include_pad behavior (#28975, #29476, #29629, #29728).TfIdfVectorizer weight indexing and skipped MinLength logits-processor construction when eos_token_id is negative (#29604, #31649).GetOverridableInitializerNames() (#29349, #29589, #29616).Split axes now produce an error instead of being accepted (#31149).ceil_mode, allowed DFT to ignore excess input data, and fixed a webpack/Terser release-build crash (#29627, #29680, #31652).CUDA package architecture selections are now aligned across plugin EP, Python, C API, TensorRT, and Node.js pipelines. Windows arm64 is only available in CUDA plugin EP (#31992):
| OS | CUDA | CUDA architectures (all in -real form) |
|---|---|---|
| Linux x64 | 12.8 | 60;70;75;80;86;89;90a;120a |
| Linux x64 | 13.x | 75;80;86;89;90a;120a |
| Linux aarch64 | 13.x | 89;90a;120a;121a |
| Windows x64 | 12.8 | 61;75;86;89;120a |
| Windows x64 | 13.x | 75;80;86;89;120a |
| Windows arm64 | 13.x | 120a;121a |
Reduced CUDA compilation time and memory usage by splitting generated SM80 MoE, fpA_intB, and MatMulNBits translation units and adding two-level workspace estimation (#29614, #29699, #29811, #31834, #31837).
Fixed CUDA 13 plugin and packaging builds on Windows, Linux, and Windows ARM64, including MSVC/TMA compatibility and CI memory limits (#31608, #31609, #31615, #31616, #31617, #31622, #31729, #31748).
Fixed MLAS AVX2 builds on toolchains without AVX-VNNI assembler support, GCC 15 -Werror builds, and an MSVC C1001 issue in the W2 AVX-512-VNNI dispatch path (#28767, #29679, #29885).
Improved Windows compatibility by skipping DXGI discovery when Win32k system calls are unavailable and delay-loading shell32 (#29755, #30889).
Fixed Dawn parallel-build races, and GPU discovery in build/test environments (#29858, #29866).
Thanks to our 63 contributors for this release!
@adrastogi, @ahsan-ca, @AngelGalindo7, @ankitm3k, @apsonawane, @blazingphoenix7, @bmehta001, @chilo-ms, @claude, @daijh, @ducviet00, @edgchen1, @elwhyjay, @eserscor, @GopalakrishnanN, @guptaishaan, @hariharans29, @Honry, @huningxin, @jchen10, @jiafatom, @jiangzhuo, @Jiawei-Shao, @JonathanC-ARM, @justinchuby, @kjg0724, @kunal-vaishnavi, @kylo5aby, @Laan33, @martin-klacer-arm, @mastryukov1990, @mcollinswisc, @miaobin, @mingmingtasd, @mirounga, @mustjab, @n1harika, @namgyu-youn, @neilmsft, @nenad1002, @nicholascelestin, @OscarFree, @prathikr, @qjia7, @quic-muchhsu, @Sammy-Dabbas, @sanaa-hamel-microsoft, @shiyi9801, @skottmckay, @tairenpiao, @TedThemistokleous, @the0cp, @tianleiwu, @titaiwangms, @velonica0, @wangw-1991, @wuisabel-gif, @xadupre, @xhcao, @xiaofeihan1, @xiaoyu-work, @yen-shi, @zlma7001
Full Changelog: v1.28.0...v1.29.0
Release Def: Branch: refs/heads/rel-1.28.0 Commit: da9b5e364c465de65c49d91e696cd6485270757f Build: https://aiinfra.visualstudio.com/Lotus/_build/resul
Release Def: Branch: refs/heads/rel-1.28.0 Commit: da9b5e364c465de65c49d91e696cd6485270757f Build: https://aiinfra.visualstudio.com/Lotus/_build/results?buildId=1324944
nvrtc is no longer linked, which significantly reduces the required CUDA redistributable footprint (#29252, #29808, #29705, #29620).OrtModelPackageApi now lives in the experimental C API and may change in future releases (#28746, #29142, #28990).wgsl-gen implementation (#29141, #28355).CUDA_QUANT_PREPROCESS is off by default (#29687).bind_input causing an out-of-bounds write (#28839)TensorAt for sub-byte packed types (#28973)Col2Im inputs to prevent heap over-read (#28706)CropAndResize against malformed crop_size tensors (#28766)BeamSearch vocab_size against logits width (#28774)WhisperDecoderSubgraph::CreateInitialFeeds (#29239)SparseAttention CSR indices/key lengths and rejected zero-dimension block_row_indices (#29015, #29242)mask_index to valid bounds (#29449)MaxpoolWithMask kernel rank against input spatial rank (#29253)EmbedLayerNorm/SkipLayerNorm shapes exceeding 32-bit output indexing (#29264)DecoderAttention/MultiHeadAttention shape inference and negative-axis handling in ExpandDims shape inference (#29268, #29448)TreeEnsemble target id validation and added input validation to LinearClassifier (#29293, #29060)DynamicQuantizeLSTM zero-point/scale validation typos (#29462)Loop/Scan output concatenation (#29397)raw_data to {0, 1} on unpack (#29238)Resize, PadFusion, and LoRA handling (#28779, #28780, #28801)WithOutputTensor in the Rust bindings (#29251)MlasConvPrepare working-buffer products and ConvTranspose pad computation with SafeInt (#29444, #29446)SamplingState::Init that could cause a heap buffer overflow (#29443)B/scales/zero-points shape in MatMulNBits::PrePack (#29445)ConstantOfShape output size against the input initializer before constant folding (#28751)Pad (int64/int32 truncation), Slice, and GatherBlockQuantized (#28721, #28704, #28718)subprocess (#28776, #28775)shell-quote, esbuild, tmp, ws, protobufjs, js-yaml, tar, markdown-it, @babel/core (#29022, #29044, #29055, #29057, #29061, #29062, #29063, #29079, #29090, #29156)external_data into session options (#28271, #28989, #29501)CompileModel validation to accept zero-input OrtModel graphs (#28771)OrtErrorCode documentation, single-sourced the values so StatusCode stays in sync, and added OrtErrorCode::ORT_DEVICE_RESET (#29018, #29065, #29748)EpDeviceUsage event, and ORT version logging (#28794)model_external_initializers_file_folder_path is now honored for file-path model loads (#29459)HOST_ACCESSIBLE OrtValue allocation (#28038)CudaQuantizer to onnxruntime.quantization (#29509)Flatten as a Direct8Bit op in the Python QDQ static quantizer (#28340)MaxPool during FP8 static quantization and fixed the FP8 (FLOAT8E4M3FN) scale reference distribution (#28488, #29350)TensorArray custom op (#28335)Attention & LLM decode
cudnn_frontend to 1.24 and enabled cuDNN SDPA for MHA/GQA (#28849)MoE & quantized GEMM
PrePack hook, symmetric with MatMulNBits, and fixed the prepack to always use the SM80 layout (#28749, #28978, #28965)MatMulNBits GEMV (#28980, #29167, #29166, #29170)MatMulNBits (#29451)block_size=32, and fused bias (#29622, #29499, #29585)Coverage, dependencies & fixes
ArgMax/ArgMin/ReduceSum and fixed LogSoftmax on the plugin EP (#29620)Softplus and Softsign up to opset 22 (#28982)cudaErrorInvalidValue on Blackwell (sm_120) (#29706)libcudart.so.13 hard dependency that broke import on CPU-only Linux (#29202, #29590)PrePack, and plugin EP allocator deleter lifetime (#29658, #29663)moe_kernels.cu (#29295)Equal/Sub/Where/ReduceSum under the enable_int64 flag (#29236, #29392)max_num_pending_dispatches configurable (#28761, #28894)GatherBlockQuantized (#29054)Cast to int64 by default and switched to naive reduction (#28804, #28174)do_rotary (#29247, #29002)head_size out-of-bounds race, the past_state == present_state buffer aliasing case, and GatherBlockQuantized dispatch failure for empty indices (#29593, #28753, #29030)round_prefer_ceil/round_prefer_floor (#28757)MatMulNBitsMlpFusion when the kernel is unavailable (#29089)Where and And builders (#28597)"model_path" must not be empty error (#29394)PadNodeGroupSelector::Check when dq_nodes is empty (#28733)Gather handler for the transpose optimizer (#28755)MatMulNBits (#29025)process_ext_address (#29248)sbgemm_neon_kernel and fixed failing KleidiAI NHWC unit tests (#28394, #29010)MatMulNBits (#29163)SpaceToDepth and int8 for DepthToSpace (#29154)is_causal bottom-right alignment for external KV cache (#29050, #28958)ReverseSequence returning zeros for zero-length sequences (#28759)RandomForestClassifier binary classification predictions (#28685)Gemm → QGemm fusion when alpha != 1 with bias, and validated DQ scale/zero-point shapes before QGemm fusion (#28131, #28714)Resize with non-nearest interpolation modes under ORT_ENABLE_ALL (#28454)Shape → Gather → TopK rank-1 regression (#28778, #29084)TransposeOptimizer type error on zero-point-less DequantizeLinear (#29192)SimplifiedLayerNorm fusion with a node-produced Pow exponent (#29196)MaxPool (#29201)Conv/ConvTranspose rank from weights when the input shape is unknown (#29149)STFT complex input frame offsets (#28961)wgsl-gen and removed the dynamic WGSL generator path (#28355, #29141)wasm_Release / build-wasm CI (#29040)HOST_ACCESSIBLE OrtValue allocation (#28038)requests package (#28825)outputHandlesArr in the OrtTrainingSession.evalStep JNI binding (#29576)WithOutputTensor (#29251)cpuinfo_deinitialize() (#28245)PIP_INDEX_URL piping for container build stages (#29803, #29823)LinearRegressor test coverage (#28979, #29083)Thanks to our 103 contributors for this release!
@adrastogi, @ankitm3k, @apsonawane, @ArsalanShakil, @bachelor-dou, @bopeng1234, @cbourjau, @chilo-ms, @chunghow-qti, @chwarr, @claude, @Craigacp, @crvineeth97, @dabhattimsft, @daijh, @danielsongmicrosoft, @derdeljan-msft, @edgchen1, @elwhyjay, @ericcraw, @eserscor, @fanchenkong1, @feich-ms, @fs-eire, @FuZoe, @gaugarg-nv, @gblong1, @GopalakrishnanN, @grybouilli, @guschmue, @haoxli, @hariharans29, @hnsyprst, @Honry, @intbf, @ishwar-raut1, @jambayk, @Jaswanth51, @jatinwadhwa921, @javier-intel, @jiafatom, @Jiawei-Shao, @JonathanC-ARM, @justinchuby, @jwludzik, @Kotomi-Du, @kpkbandi, @lucka-me, @maoger, @martin-klacer-arm, @maxwbuckley, @MayureshV1, @mc-nv, @mcollinswisc, @mdvoretc-intel, @mingyueliuh, @mklimenk, @n1harika, @namgyu-youn, @neilmsft, @orlmon01, @Osamaali313, @prathikr, @preetha-intel, @psakhamoori, @qiurui144, @qjia7, @qti-hungjuiw, @qti-yuduo, @quic-muchhsu, @raagrawal, @RajeevSekar, @Reranko05, @Rishi-Dave, @RyanMetcalfeInt8, @Sammy-Dabbas, @sanaa-hamel-microsoft, @sayanshaw24, @selenayang888, @sfatimar, @sgbihu, @Shivani767, @ssam18, @susbhere, @sushraja-msft, @tairenpiao, @the0cp, @tianleiwu, @titaiwangms, @umireon, @vthaniel, @wangw-1991, @wenqinI, @xadupre, @xenova, @xhcao, @xiaofeihan1, @xieofxie, @xuke537, @yen-shi, @yinli-systems, @yuslepukhin, @ZackyLake
Full Changelog: v1.27.1...v1.28.0
Release Def: Branch: refs/heads/rel-1.27.1 Commit: df2ba1cf8108aa63627cf4cdf8f807880b938616 Build: https://aiinfra.visualstudio.com/Lotus/_build/resul
Release Def: Branch: refs/heads/rel-1.27.1 Commit: df2ba1cf8108aa63627cf4cdf8f807880b938616 Build: https://aiinfra.visualstudio.com/Lotus/_build/results?buildId=1300459
This is a patch release on top of v1.27.0, containing targeted bug fixes, a CUDA QMoE decode-path optimization, and CI/build infrastructure fixes.
igemm regression in the KleidiAI path (#28571)azcopy (#29274)brew install applesimutils failure by trusting the wix/brew tap (#29450)mac-cpu-packing-jobs.yml (#29575)Thanks to our 8 contributors for this release!
@tianleiwu, @chilo-ms, @edgchen1, @adrastogi, @damdoo01-arm, @JonathanC-ARM, @martin-klacer-arm, @sanaa-hamel-microsoft
Full Changelog: v1.27.0...v1.27.1
Release Def: Branch: refs/heads/rel-1.27.0 Commit: 8f0278c77bf44b0cc83c098c6c722b92a36ac4b5 Build: https://aiinfra.visualstudio.com/Lotus/_build/resul
Release Def: Branch: refs/heads/rel-1.27.0 Commit: 8f0278c77bf44b0cc83c098c6c722b92a36ac4b5 Build: https://aiinfra.visualstudio.com/Lotus/_build/results?buildId=1263613
n.b. This release is targeting ONNX 1.21. ONNX 1.22 will be supported in ORT 1.28.
n.b. This changelog was generated via LLM. Only the contributor list has been verified. As always, only trust the commit history.
SoftmaxCrossEntropyLoss via label bounds validation (#28004)OneHot input validation and output-size computation (#28014)Expand and capped constant-folding output sizes (#28055)Tile kernel (#28070)MaxpoolWithMask::Compute (#28223)BitShift UB for shift amounts greater than or equal to bit width (#28272)seqlens_k vs cos_cache) (#28277)WordConvEmbedding to prevent OOB reads (#28279)CropBase scale handling (#28399)torch.load() calls to weights_only=True (#28421)OrtEp::OnSessionInitializationEnd() callback (#28319)kOrtEpDevice_EpMetadataKey_OSDriverVersion example and docs (#28282)quantize_static (#28221)ActivationRestrictedAsymmetric quantization option (#28237)block_size attribute support to QDQ quantization (#28522)FusedAdam optimizer in ORT Training (#28233)ConvTranspose-22 support (#27710)Split (split attribute path) and scalar Gather indices (#28270, #28278)FusedConv, Identity, Ceil, Tile, Cast(bool), Sin, Cos, and GatherND support (#28289, #28293, #28595, #28596, #28598)next, postcss, tmp, qs, body-parser, and other npm packages) (#27705, #27894, #28304, #28547, #28644, #28683, [#28694]present_key/present_value outputs in GQA and Gemma4 support (#28242)Session and InferenceSession (#27802)sympy an optional runtime dependency (#28141)py.typed marker to the onnxruntime package (#28438)UserLoggingFunction is used (#28314)OrtModelEditorApi ownership transfer (#28123)allowzero=1 handling for chained zero-size tensors (#28455)Resize nearest-mode rounding bug for negative halfway values (#28345)Thanks to our 106 contributors for this release!
@adrastogi, @adrianlizarraga, @AIFrameworksIntegration, @AlekseiNikiforovIBM, @angelser, @angelserMS, @ankitm3k, @anzzraju1997-glitch, @apsonawane, @arajendra, @ayappanec, @bachelor-dou, @badranX, @baijumeswani, @BoarQing, @bopeng1234, @bsosnader, @cbourjau, @chilo-ms, @chunghow-qti, @chwarr, @Craigacp, @daijh, @derdeljan-msft, @dparikh79, @edgchen1, @elwhyjay, @ericcraw, @eserscor, @feich-ms, @fs-eire, @gaugarg-nv, @gblong1, @GopalakrishnanN, @gramalingam, @guschmue, @hariharans29, @HectorSVC, @intbf, @ishwar-raut1, @Jaswanth51, @jatinwadhwa921, @javier-intel, @jchen10, @jiafatom, @Jiawei-Shao, @jnagi-intel, @jocelyn-stericker, @JonathanC-ARM, @justinchuby, @jwludzik, @kevinch-nv, @Kotomi-Du, @kpkbandi, @KV2773, @Laan33, @lhrios, @maxwbuckley, @MayureshV1, @mdvoretc-intel, @mingyueliuh, @mklimenk, @mustjab, @n1harika, @nazanin-beheshti, @orlmon01, @prathikr, @preetha-intel, @psakhamoori, @qiurui144, @qjia7, @qti-ashwshan, @qti-hungjuiw, @qti-yuduo, @rajatmonga, @RajeevSekar, @Rishi-Dave, @rvandermeulen, @RyanMetcalfeInt8, @SamuelLess, @sanaa-hamel-microsoft, @sfatimar, @sgbihu, @shiyi9801, @simonbyrne, @skottmckay, @susbhere, @sushraja-msft, @tairenpiao, @TejalKhade28, @theHamsta, @tianleiwu, @titaiwangms, @umangb-09, @velonica0, @vraspar, @vthaniel, @wenqinI, @xadupre, @xenova, @xhan65, @xiaofeihan1, @xieofxie, @yuslepukhin, @ZackyLake, @zejianzhang1982, @zz002
Full Changelog: v0.1.4...v1.27.0
Release Def: Branch: refs/heads/rel-1.26.0 Commit: 8c546c37b43caaca1fa25db430dab94b901cf277 Build: https://aiinfra.visualstudio.com/Lotus/_build/resul
Release Def: Branch: refs/heads/rel-1.26.0 Commit: 8c546c37b43caaca1fa25db430dab94b901cf277 Build: https://aiinfra.visualstudio.com/Lotus/_build/results?buildId=1209449
n.b. The following was generated via LLM from Git history. Only the contributor list has been verified.
onnxruntime-<os>-<arch>-gpu_cuda13-<version>.<ext>.ort model loads (#28164).setattr configuration with an allowlist (#28083).roundingType -> outputShapeRounding (#28172).@tianleiwu, @yuslepukhin, @edgchen1, @vraspar, @hariharans29, @skottmckay, @eserscor, @xadupre, @sanaa-hamel-microsoft, @elwhyjay, @Rishi-Dave, @titaiwangms, @adrianlizarraga, @jatinwadhwa921, @jchen10, @Jiawei-Shao, @maxwbuckley, @preetha-intel, @qjia7, @qti-hungjuiw, @RajeevSekar, @umangb-09, @adrastogi, @akote123, @amd-genmingz, @ankitm3k, @apsonawane, @bachelor-dou, @baijumeswani, @bopeng1234, @chilo-ms, @chwarr, @Craigacp, @dccarmo, @derdeljan-msft, @ericcraw, @fdwr, @fs-eire, @gaugarg-nv, @gblong1, @GopalakrishnanN, @Honry, @intbf, @ishwar-raut1, @Jaswanth51, @javier-intel, @JonathanC-ARM, @julia-thorn, @justinchuby, @jwludzik, @Kevin-Taha, @Kotomi-Du, @MayureshV1, @mdvoretc-intel, @miaobin, @milpuz01, @mingyueliuh, @mklimenk, @n1harika, @prathikr, @psakhamoori, @qti-yuduo, @quic-calvnguy, @RyanMetcalfeInt8, @sfatimar, @sgbihu, @ShirasawaSama, @ssam18, @susbhere, @sushraja-msft, @TejalKhade28, @theHamsta, @TomCrypto, @TsofnatMaman, @velonica0, @vthaniel, @wenqinI, @xhan65, @xhcao
Release Def: Branch: refs/heads/rel-1.25.1 Commit: 8a77e459420f58fb946fd9067285cfa719f10bdd Build: https://aiinfra.visualstudio.com/Lotus/_build/resul
Release Def: Branch: refs/heads/rel-1.25.1 Commit: 8a77e459420f58fb946fd9067285cfa719f10bdd Build: https://aiinfra.visualstudio.com/Lotus/_build/results?buildId=1190117
n.b. This changelog is LLM generated. Only the contributor listing has been verified.
SetRawDataInTensorProto in NVIDIA TensorRT RTX tests (#28065)Thanks to our 7 contributors for this release:
@guschmue, @sanaa-hamel-microsoft, @apsonawane, @eserscor, @ishwar-raut1, @qjia7, @theHamsta
Full Changelog: v1.25.0...v1.25.1
Release Def: Branch: refs/heads/rel-1.25.0 Commit: 7a71bc575b189cdedea7fa2c0f87389f870bd10e Build: https://aiinfra.visualstudio.com/Lotus/_build/resul
Release Def: Branch: refs/heads/rel-1.25.0 Commit: 7a71bc575b189cdedea7fa2c0f87389f870bd10e Build: https://aiinfra.visualstudio.com/Lotus/_build/results?buildId=1175068
Release Def: Branch: refs/heads/rel-1.24.4 Commit: 2d924974ef147392ced8409d36bd6d2e7fcc8a74 Build: https://aiinfra.visualstudio.com/Lotus/_build/resul
Release Def: Branch: refs/heads/rel-1.24.4 Commit: 2d924974ef147392ced8409d36bd6d2e7fcc8a74 Build: https://aiinfra.visualstudio.com/Lotus/_build/results?buildId=1124809
Release Def: Branch: refs/heads/rel-1.24.3 Commit: 3a728b75062256951b6e19ce718907cf1a1d4cf0 Build: https://aiinfra.visualstudio.com/Lotus/_build/resul
Release Def: Branch: refs/heads/rel-1.24.3 Commit: 3a728b75062256951b6e19ce718907cf1a1d4cf0 Build: https://aiinfra.visualstudio.com/Lotus/_build/results?buildId=1108362
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Compare
Compare
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →