Skip to content

Bump Microsoft.ML.OnnxRuntime from 1.28.0 to 1.29.0 - #7

Open
dependabot[bot] wants to merge 1 commit into
mainfrom
dependabot/nuget/src/VoiceMemoryDemo.App/Microsoft.ML.OnnxRuntime-1.29.0
Open

dependabot[bot] wants to merge 1 commit into
mainfrom
dependabot/nuget/src/VoiceMemoryDemo.App/Microsoft.ML.OnnxRuntime-1.29.0

Conversation

@dependabot

@dependabot dependabot Bot commented on behalf of github Aug 21, 2026 •

Copy link
Copy Markdown

Updated Microsoft.ML.OnnxRuntime from 1.28.0 to 1.29.0.

Release notes

Sourced from Microsoft.ML.OnnxRuntime's releases.

1.29.0

Announcements & Breaking Changes

  • onnxruntime-web has announced the deprecation of WebGL and JSEP. The native WebGPU EP is the recommended path going forward. See the deprecation and migration plans for details (#​29716, #​31683).
  • POSIX telemetry is now available on Linux, macOS, Android, and iOS when ONNX Runtime is built with telemetry enabled. It does not change the public ABI, WebAssembly remains telemetry-free, and setting ORT_DISABLE_TELEMETRY=1 before initialization disables non-Windows telemetry for the process (#​27379, #​29872).
  • The unused internal onnxruntime/python/tools/tensorrt dashboard tooling was removed. This does not affect the TensorRT Execution Provider APIs (#​29395).

Security Fixes

Path, bounds, and input validation

Supply chain and tooling

  • Updated npm lockfiles, refreshed the Next.js end-to-end fixture lockfile for security advisories, and upgraded adm-zip for onnxruntime-node (#​29827, #​29926, #​31192).

New Features

Core APIs & Runtime

  • Default intra-op and inter-op thread-pool sizes can now be set with ORT_INTRA_OP_NUM_THREADS and ORT_INTER_OP_NUM_THREADS. Explicit thread settings still take precedence, and 0 preserves machine-sized defaults (#​29688).
  • Added weightless-model support for all initializer types, allowed zero-input EpContext nodes, and wired maximum-shape inference into workspace estimation (#​29607, #​29799, #​31613).
  • Added ONNX-domain support for rotary embedding and a fused MRotaryEmbedding contrib operator for Qwen mRoPE variants (#​29261, #​31728).
  • Added multi-shape profiling to onnxruntime_perf_test through --data_shape, plus verbose graph-transformer tracing and broader inference-session error-path coverage (#​29555, #​29558, #​29569, #​29571).

Execution Provider ABI & Plugin EPs

  • WebGPU now supports device-free compile-only sessions for offline graph transformation (#​29681).
  • Expanded CUDA plugin EP packaging and testing, including Windows ARM64 package and size options, updated package outputs, and aligned architecture selections across Python, C API, TensorRT, Node.js, and plugin packages (#​31635, #​31722, #​31992).
  • Improved plugin lifecycle handling by unloading failed EP library loads and fixing allocator-deleter lifetime (#​29634, #​29770).

Execution Provider Updates

NVIDIA CUDA EP

Attention and decoding

  • Added PagedAttention with quantized KV cache, XQA decode, MLA, QK-Norm, and head-sink support (#​29912).
  • Extended quantized KV-cache support with attention sinks, independent and per-channel scales, sliding-window cache support, and a fused K/V dequantization launch (#​29900, #​29904, #​31480).
  • Added a cuDNN SDPA decode tier to the standard ONNX Attention CUDA kernel and enabled cuDNN SDPA for contrib Attention (#​29715, #​29717).
  • Added attention_bias support to the GroupQueryAttention unfused path and state_window support to LinearAttention and CausalConvWithState for MTP (#​29525, #​31157).
  • Fixed LinearAttention on GPUs with limited shared memory (#​31982).

MoE and quantized GEMM

1.28.1

This is a patch release on top of v1.28.0, containing support for device-free WebGPU compilation, improved compatibility with sandboxed Windows processes, and targeted graph-validation fixes.

WebGPU EP

  • Added support for device-free compile-only sessions, enabling offline graph transformation and optimized-model serialization without access to GPU hardware (#​29681)

Bug Fixes

  • Prevented an access violation in Windows processes under Win32k lockdown by skipping DXGI device discovery (#​29755)
  • Allowed zero-input EPContext nodes, aligning their schema with support for compiling zero-input models (#​29799)
  • Hardened FastGelu fusion to skip malformed Mul and Pow patterns (#​32016)
  • Added validation for in-memory external initializer references, rejecting unregistered or mismatched data before graph transformation (#​32042)

Contributors

Thanks to our 4 contributors for this release!

@​apsonawane, @​shiyi9801, @​adrastogi, @​mingmingtasd

Full Changelog: v1.28.0...v1.28.1

Commits viewable in compare view.

@dependabot dependabot Bot added .NET Pull requests that update .NET code dependencies Pull requests that update a dependency file labels Aug 21, 2026
@dependabot
dependabot Bot force-pushed the dependabot/nuget/src/VoiceMemoryDemo.App/Microsoft.ML.OnnxRuntime-1.29.0 branch 3 times, most recently from b902283 to 007bd87 Compare August 21, 2026 11:14
---
updated-dependencies:
- dependency-name: Microsoft.ML.OnnxRuntime
  dependency-version: 1.29.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
@dependabot
dependabot Bot force-pushed the dependabot/nuget/src/VoiceMemoryDemo.App/Microsoft.ML.OnnxRuntime-1.29.0 branch from 007bd87 to 1b9a1bb Compare August 21, 2026 12:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Pull requests that update a dependency file .NET Pull requests that update .NET code

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants