Sync with Microsoft ONNX Runtime - 28092026 - #1312
Merged
Merged
Conversation
### Motivation and Context AutoEP can select and expose multiple `OrtEpDevice` instances for a single session. Some of those devices may be backed by the same EP factory, so querying each device for custom-op domains can return repeated domains. ORT gathers all EP-provided custom-op domains into one session-local collection and registers them with the session's custom registry before model load. That registry expects each domain to be unique. If the same domain is registered twice, registration fails with "Domain already set in registry", even though the duplicate came from equivalent AutoEP devices rather than user configuration. This mainly affects multi-device AutoEP sessions. Single-device AutoEP sessions do not hit the issue because only one copy of the domain is collected. ### Description Deduplicate custom-op domains collected from AutoEP devices before registering them with the session-local custom registry. The change keeps only one copy of each EP-provided custom-op domain while still skipping domains already supplied through session options. This ports the AutoEP custom-op domain deduplication from onnxruntime/onnxruntime-qnn#608.
### Description Packed GroupQueryAttention shape inference divided the query hidden size by `num_heads + 2 * kv_num_heads` without validating the attributes first. Malformed head counts could therefore cause division by zero or signed overflow while resolving a model. Validate both counts before the calculation, reject grouped-head overflow, require concrete hidden sizes to divide evenly, and preserve symbolic hidden dimensions. ### Testing - `cmake --build build\Windows\Release --config Release --target onnxruntime_graph --parallel` - `cmake --build build\Windows\Release --config Release --target onnxruntime_provider_test --parallel` - `onnxruntime_provider_test.exe --gtest_filter=GroupQueryAttentionTest.CausalMaskDefaultsToEnabled_CPU` --------- Co-authored-by: Daniel Song <danielsong@microsoft.com>
…soft#32786) ### Description In the 2-bit `ParQuantizeLinearStd` helper, return once the leading partial byte has covered the whole scale interval. Before this change, such an interval also went through the trailing partial-byte path, which re-quantized the preceding elements in the same byte with the current scale and zero point. Adds a per-axis INT2 regression test where each scale interval is a single element. ### Motivation and Context With float input and per-axis quantization on the last axis (for example a `[K, N]` tensor with one scale per column), each scale interval holds one element, so intervals start and end inside packed bytes. The first two elements of each byte are then re-quantized with the scale and zero point of a later element in that byte, without any error. Intervals with more than one element, FP16 input and INT4 outputs are not affected. Repro with onnxruntime 1.30.0: ```python import numpy as np import onnxruntime as ort from onnx import TensorProto, helper, numpy_helper nodes = [ helper.make_node("QuantizeLinear", ["x", "scale"], ["q"], axis=1, output_dtype=TensorProto.UINT2), helper.make_node("DequantizeLinear", ["q", "one"], ["y"]), ] inits = [ numpy_helper.from_array(np.array([1, 1, 4, 1, 1, 1, 4, 1], np.float32), "scale"), numpy_helper.from_array(np.array(1, np.float32), "one"), ] graph = helper.make_graph(nodes, "g", [helper.make_tensor_value_info("x", TensorProto.FLOAT, [2, 8])], [helper.make_tensor_value_info("y", TensorProto.FLOAT, [2, 8])], inits) model = helper.make_model(graph, opset_imports=[helper.make_opsetid("", 25)], ir_version=10) so = ort.SessionOptions() so.graph_optimization_level = ort.GraphOptimizationLevel.ORT_DISABLE_ALL sess = ort.InferenceSession(model.SerializeToString(), so, providers=["CPUExecutionProvider"]) print(sess.run(None, {"x": np.ones((2, 8), np.float32)})[0]) # Expected: [[1 1 0 1 1 1 0 1], [1 1 0 1 1 1 0 1]] # Actual: [[0 0 0 1 0 0 0 1], [0 0 0 1 0 0 0 1]] ``` Found while working on microsoft#32735, which routes more cases through this helper.
### Description - Batch JSEP WebGPU Concat inputs using: ```ts maxInputsPerDispatch = maxStorageBuffersPerShaderStage - 1 ``` - Reuse the output buffer across dispatches and preserve concat-axis offsets. - Add an 8-input regression case to the WebGPU suite. ### Motivation and Context Concat bound every input and the output in one shader. At eight inputs, this exceeded the device’s eight-storage-buffer limit, silently skipped execution, and returned zeros. <!-- START COPILOT CODING AGENT SUFFIX --> - Fixes microsoft#32757 --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: tianleiwu <30328909+tianleiwu@users.noreply.github.com>
…ns. (microsoft#32768) Split the webgpu build & test stage into 2 stages. This lets us use better machines for build. In total, this cuts the action time between 30 and 45 minutes. <img width="515" height="247" alt="image" src="https://github.com/user-attachments/assets/2088079a-1060-4cde-baf8-9e04585eb578" /> -> <img width="1024" height="299" alt="image" src="https://github.com/user-attachments/assets/aeb1d7d8-283b-437e-af39-5708ce406167" /> --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: eserscor <247253654+eserscor@users.noreply.github.com> Copilot-Session: 409e405a-fdd1-4c73-b54a-4676012fb9e5
## Description Adds an unquantized WebGPU `MoE` kernel and expands WebGPU `QMoE` support for the configurations needed by packed token-major inference, including Qwen 3.8 Flash. The existing operator schemas already accept both representations: - Packed 2D input: `(num_tokens, hidden_size)` - Dense 3D input: `(batch_size, sequence_length, hidden_size)` `MoE` and `QMoE` are token-local, so concatenating variable-length requests does not require a separate varlen operator or a cumulative sequence-length input. Each router tensor contains one row per input token, and the output preserves the input shape. ## Changes ### WebGPU MoE - Implements and registers the unquantized WebGPU `MoE` kernel for float and float16. - Supports packed 2D and dense 3D inputs. - Supports ReLU, GELU, SiLU, identity, and SwiGLU activations. - Supports separate FC3 gating and interleaved or concatenated fused SwiGLU layouts. - Handles empty inputs, chunked multi-token execution, and hidden sizes that are not multiples of the workgroup size. - Uses segmented-buffer-safe tensor access throughout the WebGPU shader pipeline, including stacked expert weights and shape-preserving outputs. ### WebGPU QMoE - Adds all schema activation types and both fused SwiGLU layouts. - Adds separate FC3 projection and gating. - Adds aggregation-only `router_weights` while continuing to use `router_probs` for Top-K selection, including selected Top-K normalization when `normalize_routing_weights=1`. - Adds block-wise explicit zero-point support. - Makes stacked `MatMulNBits` zero-point indexing follow the selected expert, matching weight and scale indexing. - Preserves the optimized single-token decode path. - Rejects unsupported row-wise and optimized single-token explicit-zero-point combinations with clear errors. - Validates that `num_experts` fits the adapter's workgroup-size and invocation limits before creating the gate shader. ### Packed semantics and tests - Clarifies packed 2D and dense 3D behavior in the authoritative schemas and generated contrib documentation. - Adds a packed ragged-batch test representing request lengths `[2, 1, 2]`. - Uses distinct token rows and nonzero weights so the packed test detects hidden-state reordering or cross-request interaction. - Disables CPU fallback in every WebGPU-specific test so registration or kernel-claim regressions cannot pass through another EP. - Adds focused coverage for float/float16 MoE, dense and packed inputs, activations, FC3, SwiGLU fusion layouts, expert-specific zero points, normalized `k > 1` router weights, non-aligned gather tails, the 2,048-token chunk boundary, adapter expert limits, empty inputs, and the optimized QMoE decode path. ## Provider-specific scope This change targets portable WebGPU functionality. CUDA-specific `fp4`, `nvfp4`, `fp8`, and `wfp4afp8` formats, CUDA expert-weight prepacking, and sparse mixer execution remain unsupported on WebGPU and are rejected explicitly. The normal symmetric INT4 Qwen path remains supported. Block-wise asymmetric INT4/INT8 weights, separate router weights, and FC3-based gating are now also represented by the WebGPU implementation and tests. The official Qwen 3.8 Flash FP8 checkpoint uses 128x128 block-scaled FP8 expert weights. CUDA `QMoE` currently supports FP8 weights with one global scale per expert, while its existing block-scaled modes are `nvfp4` and `wfp4afp8`. Supporting the official block-FP8 checkpoint therefore requires a separate CUDA/schema change and is intentionally left to a focused follow-up rather than expanding this WebGPU PR. ## Validation - Windows Release WebGPU provider and `onnxruntime_provider_test` targets build successfully with Dawn and the D3D12 Agility SDK. - WGSL template tests pass on Windows and Linux CI. - Deterministic nonzero WebGPU MoE and optimized single-token QMoE tests pass locally. - `clang-format --dry-run --Werror` passes for changed C++ files. - `git diff --check` passes. - Editor diagnostics report no errors. - Generated `OperatorKernels.md` content is synchronized with CI output. Multi-token QMoE dispatch reaches the local D3D12 adapter, but repeated local runs can end in `DXGI_ERROR_DEVICE_REMOVED`; cross-adapter runtime validation is covered by WebGPU CI. --------- Co-authored-by: Sayan Shaw <sayanshaw@microsoft.com> Co-authored-by: Justin Chu <justinchuby@users.noreply.github.com> Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…rosoft#32767) # WebGPU GatedDeltaNet Compact State Updates ## Changes This PR adds compact state-update capture to the native WebGPU `com.microsoft.GatedDeltaNet` kernel, which previously rejected `state_update_capacity > 0`. - Validates `state_update_capacity` and the required `[batch]` `capture_count` input. - Allocates the packed float32 `state_update` output and includes its buffers in WebGPU binding accounting. - Captures the clamped per-row prefix of decay, shared-key, and delta factors in the recurrent shader. These factors replay `S = decay * S + outer(key, delta)`. - Preserves inverse-GQA key sharing and existing value-channel tiling. - Uses the recurrent path while capture is active because parallel prefill does not emit per-token transitions. - Honors `state_update_active=0` by clearing the entire output; uncaptured entries remain unspecified when capture is active. - Updates the registered schema text and generated contrib-operator documentation to advertise WebGPU compact-state support. ## Test Coverage Tests now compare only the captured decay/key/delta ranges instead of assuming uncaptured storage is zero. The inactive test still verifies that the complete output is zero-filled. Coverage includes ragged capacity-7 cases with zero, partial, and full capture counts, plus an FP16 Qwen-like case (`Hq=2`, `Hv=6`, `head_size_qk=head_size_v=128`). The larger case exercises the FP16 aliases, value tiling, shared-key head mapping, and binding-packing path used by GenAI. ## DFlash2 Prerequisite This PR adds the ORT kernel prerequisite; it does **not** enable end-to-end WebGPU DFlash2 by itself. GenAI also needs: - builder export with nonzero `state_update_capacity` ([microsoft#2453](microsoft/onnxruntime-genai#2453), [microsoft#2585](microsoft/onnxruntime-genai#2585)); and - compact-state replay after partial acceptance ([microsoft#2454](microsoft/onnxruntime-genai#2454)), including WebGPU `ReplayStateUpdates` ([microsoft#2606](microsoft/onnxruntime-genai#2606)). Before advertising WebGPU DFlash2 support, an end-to-end GenAI run should cover full, partial, and zero draft acceptance, match non-speculative output/state, and confirm that neither target execution nor replay silently falls back from WebGPU. --------- Co-authored-by: Prathik Rao <prathikrao@microsoft.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
### Description Add an AVX2/FMA/F16C MLAS kernel for FP16 LayerNorm and RMSNorm on x86. The kernel widens FP16 inputs in vectors, performs the reductions and affine transform in FP32 (with a double-precision mean accumulation for LayerNorm), and narrows the result back to FP16 in vectors. The CPU LayerNorm implementation dispatches to this kernel for rows of at least 16 elements when the processor supports the required instructions. Other architectures, older x86 processors, short rows, BF16, and unsupported output-statistics types retain the existing scalar path. ### Motivation and Context The existing FP16 CPU path converts every input and output element individually around scalar normalization math. Transformer workloads commonly execute LayerNorm and RMSNorm over rows of 1K-4K elements, where conversion and scalar arithmetic dominate execution time. The benchmark processes 4,194,304 FP16 elements per invocation with rows parallelized across OpenMP workers, mirroring ORT's per-row thread-pool parallelism. It was pinned to one 48-core Intel Xeon Platinum 8480C socket with workers bound to cores. The table reports the median speedup from three process runs; graph/session overhead is intentionally excluded. | Operation | Row size | 1 thread | 8 threads | 16 threads | 32 threads | |---|---:|---:|---:|---:|---:| | LayerNorm | 1024 | 15.66x | 15.39x | 15.85x | 14.89x | | RMSNorm | 1024 | 12.07x | 11.91x | 11.96x | 10.63x | | LayerNorm | 4096 | 14.96x | 15.72x | 14.41x | 13.31x | | RMSNorm | 4096 | 11.71x | 11.42x | 10.84x | 9.61x | At 32 threads, the optimized kernels reach median throughputs of 54.0-57.6 billion elements/s for LayerNorm and 81.2-89.4 billion elements/s for RMSNorm. Validation: - Built `onnxruntime_mlas_test` and `onnxruntime_provider_test` in a CPU Release configuration. - `MlasLayerNormF16Test.MatchesFloatReference` passes for LayerNorm with/without bias and RMSNorm at row sizes 16, 31, and 128. - `RMSNormalizationOpTest.RMSNorm_float16_OptimizedRow` passes through the CPU execution provider. - Existing `RMSNormalizationOpTest.*` suite passes (26 tests before adding the dedicated optimized-row case). - Lintrunner passes on all changed files. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
### Description The CPU N-D Col2Im loop computes image and column indices for each value. Compute the valid range on the last axis once per kernel tap, then add each valid row. Advance the input row pointer without another index calculation. Keep the order of additions. For windows that can overlap, use the original loop if the output contains NaN. This keeps NaN payloads. Skip the NaN scan when the windows cannot overlap, so large gaps in the output do not add a scan cost. This applies to FP32 CPU ConvTranspose with 1D or 3D spatial inputs and other users of `Col2imNd`. The 2D helper from microsoft#32698 and GEMM do not change. ### Motivation and Context Measurements through the C API on a Core Ultra 7 255H, WSL2, GCC 13.3, Release, CPU EP, one thread. The base is main at `056897df`. The table gives the median of seven rounds with fixed CPU affinity and alternating run order. These are repeated runs after session setup. Timing includes GEMM, output initialization, and bias. It excludes session setup and prepacking of constant weights. | Case | main → change, ms | Speedup | |---|---:|---:| | ConvTranspose 1D, k16 / s8 | 1.363 → 0.470 | 2.90× | | ConvTranspose 3D, k3×3×3 / s1 | 13.821 → 1.614 | 8.56× | | Col2Im 3D, k3×3×3 / s1 | 12.811 → 0.940 | 13.63× | | Synthetic 3D decoder, 3 layers | 21.516 → 2.575 | 8.35× | For ConvTranspose, the 1D input is `[1,16,2048]`, with 16 output channels and padding 4 on each side. The 3D input is `[1,8,16,24,32]`, with 8 output channels and padding 1 on each side. Col2Im uses column shape `[1,216,12288]` and image shape `[16,24,32]`. The decoder is a synthetic graph with input `[1,8,8,16,24]`; it is not a trained model. Overlapping NaN inputs require a second pass. The 3D Col2Im case takes 13.403 → 17.677 ms with mixed special values. An output with infinity and no NaN does not need a second pass. Small-input timing varies; the unchanged 2D control and same-binary checks also vary by a few percent. The Math and CPU ConvTranspose/Col2Im tests pass. Two new tests cover row bounds, addition order, and special values with and without overlap. Additional local checks cover groups, dilation, asymmetric padding, 1 and 4 threads, and output equality. Checks with fixed and runtime weights also pass, with no bias and with bias followed by ReLU. The extracted function also passes ASan and UBSan. CPU `ConvGrad` gradient outputs also match the base with training operators enabled. A full training loop was not tested.
## Summary - add an `adapterIndex` provider option for selecting a zero-based physical Dawn adapter - enumerate adapters matching the requested backend and power preference in native Dawn builds - reject unsupported external-Dawn/wasm combinations, out-of-range indices, invalid values, and conflicting context reuse - document the distinction between physical `adapterIndex` and ORT `deviceId` ## Testing - `clang-format --dry-run --Werror` on all changed C++ files - fresh current-main native Dawn configuration completed - `webgpu_context.cc` and `webgpu_provider_factory.cc` compiled successfully against Dawn v20260818.211311 - full local plugin link was blocked by unrelated existing GCC 13 `-Werror=maybe-uninitialized` failures in WebGPU Gemm/InstanceNorm; PR CI covers the supported build matrix
…osoft#32763) **Description** Enable com.microsoft::PagedAttention on the WebGPU EP with is_causal=0 when local_window_size <= 0 . This is needed by block-drafting models such as DFlash2, where all query tokens submitted in one speculative block must attend bidirectionally to the complete live context. Previously, the WebGPU kernel rejected every non-causal configuration and then hard-coded causal behavior when constructing WebgpuAttentionParameters . **Changes** - Store and propagate the is_causal attribute through WebGPU PagedAttention into the existing FlashAttention is_unidirectional shader specialization. - Preserve the existing causal and causal local-window behavior. - Continue to reject is_causal=0 with local_window_size > 0 ; query-relative left-window masking for that combination is intentionally out of scope. - Generalize the C++ reference implementation to compare causal prefix visibility with non-causal full-context visibility. - Add focused WebGPU parity tests covering: - direct paged decode with a multi-token speculative block; - fused paged prefill; - gather-then-FlashAttention fallback; - ragged requests with different query and past lengths. - Replace the Python non-causal rejection test with reference parity coverage and add a dedicated test for the unsupported non-causal local-window combination. No WGSL redesign is required: the dense and paged FlashAttention shaders already support causal and non-causal specializations. This change only restores the missing host-side attribute plumbing. **Masking behavior** - is_causal=1 : query row i attends through past_length + i . - is_causal=0 : every query row attends through past_length + query_length , including later tokens in the same submitted block. - is_causal=0 && local_window_size>0 : remains explicitly unsupported. **Validation** - git diff --check - Python syntax compilation for test_paged_attention.py - Source-level checks for attribute propagation and focused test coverage --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
### Description Extends `ConvBNFusion` to support fusing `ConvTranspose -> BatchNormalization`. The existing fusion path only handled `Conv -> BatchNormalization`. This change adds `ConvTranspose` as a supported target and folds BatchNormalization parameters into ConvTranspose when the required initializers are constant and the channel shapes are compatible. For ConvTranspose, the weight layout is `[C_in, C_out/group, ...]`, so the BatchNormalization scale is applied by absolute output channel instead of using the existing Conv weight-axis scaling path. The existing Conv fusion behavior is unchanged. This also adds graph transformer tests for: - ConvTranspose without bias - ConvTranspose with bias - grouped ConvTranspose - no fusion when ConvTranspose weight is not constant ### Motivation and Context `Conv -> BatchNormalization` is already fused by `ConvBNFusion`, but the same optimization was not applied to `ConvTranspose -> BatchNormalization`. As a result, eligible ConvTranspose models kept an extra BatchNormalization node after graph optimization. Fusing the BatchNormalization into ConvTranspose removes the redundant node and makes ConvTranspose consistent with the existing Conv fusion behavior while preserving numerical equivalence. Fixes microsoft#14270
This pull request strengthens input validation and error handling for the `ConvTransposeWithDynamicPads` operator. The main focus is on ensuring that the required dynamic pads tensor is always provided, both at schema and runtime, and that missing or malformed inputs are handled gracefully. Additionally, a new test is added to verify the operator's behavior when the pads input is missing. **Input validation and schema enforcement:** * The operator schema for `ConvTransposeWithDynamicPads` is updated to make the `Pads` input required instead of optional, ensuring that all usage must provide this tensor. * The shape inference function now immediately returns if fewer than three inputs are provided, preventing access to missing inputs. * In the operator implementation, a check is added to explicitly return an error if the `Pads` tensor is missing when dynamic padding is expected. * The `getInputData` method in `InferenceContextImpl` now safely returns `nullptr` if the requested input index is out of bounds, preventing potential crashes. **Testing improvements:** * A new unit test, `ConvTransposeWithDynamicPads_MissingPadsRejected`, is added to verify that the operator correctly rejects models where the required `Pads` input is missing. **Test infrastructure:** * Includes necessary headers in the test file to support the new test case.
### Description Constrain the Python documentation ONNX dependency to versions before 1.23. `docs/python/examples/plot_pipeline.py` imports `onnx.tools.net_drawer`, which was removed in ONNX 1.23. The documentation requirements now resolve to ONNX 1.21 or 1.22 only. ### Motivation and Context The Python API documentation workflow currently fails while building the Sphinx gallery because the unbounded `onnx >= 1.21.0` requirement resolves to ONNX 1.23, where `onnx.tools.net_drawer` is unavailable. This restores generation of the Python API-doc artifact required by the GitHub Pages publication workflow. Signed-off-by: Damien Dooley <damien.dooley@arm.com>
### Description <!-- Describe your changes. --> Fix warning from recent change in microsoft#32729 > LINK : error C2220: the following warning is treated as an error [N:\_work\1\b\RelWithDebInfo\onnxruntime_provider_test.vcxproj] LINK : warning C4744: 'static char const absl::lts_20250814::base_internal::FastTypeTag<class std::basic_string<char,struct std::char_traits<char>,class std::allocator<char> > >::kDummyVar' has different type in 'n:\_work\1\b\relwithdebinfo\_deps\googletest-src\googletest\src\gtest-all.cc' and 'n:\_work\1\s\onnxruntime\test\providers\webgpu\webgpu_context_test.cc': 'unsigned char' and 'signed char' [N:\_work\1\b\RelWithDebInfo\onnxruntime_provider_test.vcxproj] ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> Fix packaging pipeline
) ### Description The CUDA Plugin EP factory handed **one BFC device arena per GPU** to every `CreateAllocator` caller: every session, plus the environment's shared allocator. After a CUDA graph is captured, the session frees the graph's intermediate chunks back to the arena, but the captured graph still reads and writes those addresses on every replay. Because the arena was shared, another session could be given those chunks and overwrite them between replays. With two graph-captured sessions on one device (a Qwen3.8-27B target plus a DFlash2 drafter in onnxruntime-genai), this silently corrupted generation. MMLU-Pro dropped 5.5 pp and speculative acceptance fell from 87.7% to 69.8%. The bundled CUDA EP is not affected because it creates a separate arena for each session. This PR gives each `CreateAllocator` call for device memory its own arena, which matches the bundled EP. It replaces the temporary graph-disable gate in microsoft#32801. #### Changes | File | Change | |------|--------| | `cuda_ep_factory.h/.cc` | Replace the single `device_arena` and its refcount with `std::vector<DeviceArena> device_arenas`. Each entry tracks its own `has_quarantine` / `abandoned` state. `CreateAllocatorImpl` creates a new arena on every call and now honors that session's arena options, which the shared arena used to ignore with a warning. `ReleaseAllocatorImpl` destroys only the matching arena, or leaks it on purpose if it is quarantined or abandoned. Stream run-end reset and quarantine/abandon after an undrained release now visit every arena on the device. For abandon, this is the same scope the single shared arena had before. `GetDeviceArenaForDevice()` had no callers and is removed. | | `cuda_plugin_arena_test.cc` | New `DeviceAllocator_IsSessionScoped` test: an allocation from session A must not show up in session B's arena stats or the environment allocator's stats. It **fails on the current shared-arena plugin** (`NumAllocs` 3 vs 2) and passes with this change. | | `docs/cuda_plugin_ep/*.md` | Describe the per-session device arena and why CUDA graphs require it. | The pinned arena and the CUDA mempool allocator are still shared per device and are unchanged. ### Motivation and Context #### Root-cause isolation (MMLU-Pro, 200 questions, 8,192 max new tokens, DFlash2 drafts=6, concurrency 8) | Configuration | Correct | Cap hits | Acceptance | |---|---:|---:|---:| | Bundled CUDA EP, graphs on | 168 | 3 | 87.7% | | Plugin (current), graphs on | 157 | 64 | 69.8% | | Plugin, target session graph only | 156 | 66 | 69.4% | | Plugin, drafter session graph only | 167 | 7 | 86.2% | | Plugin, per-session arena prototype, both graphs on | 169 | 4 | 87.8% | | Same prototype binary, shared arena (control) | 156 | 64 | 70.1% | Toggling only the arena sharing, with the same binary and both graphs on, flips the result. That identifies the shared arena as the cause. #### This PR (graphs on for both sessions) | Metric | Bundled | Plugin with this PR | |---|---:|---:| | MMLU-Pro correct | 168/200 | **168/200** | | Cap hits | 3 | 3 | | Acceptance | 87.7% | 87.8% | | Paired vs bundled | — | 1 win / 1 loss, McNemar p=1.0, 183/200 token-exact | | Paired vs current plugin | — | 13 wins / 2 losses, p=0.0074 | #### Performance Batch 1, 32,768 prompt tokens, 1,024 generated tokens, DFlash2 draft width 6, 3 fresh-process repetitions, all arms run at the same time on separate H200s: | Runtime | Median decode tok/s | Median prefill tok/s | Median TTFT | Peak workload | |---|---:|---:|---:|---:| | Bundled CUDA EP | 314.58 | 3,767.2 | 8.698 s | 31,590 MiB | | Plugin, graphs disabled (microsoft#32801 gate) | 305.65 (0.972x) | 3,822.9 | 8.572 s | 30,900 MiB | | **Plugin, this PR** | **317.01 (1.008x)** | 3,780.7 | 8.667 s | 31,526 MiB | Re-enabling graphs safely recovers the 2.8% decode loss from the gate. Peak memory is 64 MiB below the bundled EP. #### Tests - `CudaPluginArenaTest.*`, `CudaPluginUserStreamGraphTest.*`, `CudaPluginDeviceDiscoveryTest.*`: 51/51 passed on an H200 (Release, SM90). - `clang-format --dry-run --Werror` and `git diff --check` pass.
…oft#32810) ### Description Adds optional M (row) chunking to the CUDA `MatMulNBits` fpA_intB path to bound memory that grows with `M`, targeting large-vocabulary LM heads during chunked prefill. It is off by default and size-gated, so small layers are never split. | Setting | Meaning | |---|---| | `ep.cuda.matmul_nbits_m_chunk_size` (session config) / `ORT_MATMULNBITS_M_CHUNK_SIZE` (env) | Max rows of `A` per fpA_intB launch. `0`/unset disables. Session config wins. Also works for the CUDA plugin EP. | | `ORT_MATMULNBITS_FORCE_CHUNKED=1` (existing) | Now also bypasses the M-chunking size gate (it already bypassed the fallback's size guards). | A node is chunked only when **`M > chunk`** and **`M * (N + K) * sizeof(T) > 256 MiB`** (the tactic profiler's `A` and `C` buffers for an unchunked `M`; same budget as the existing N-chunked fallback). For Qwen3.8-27B in fp16/bf16: | node | chunked when `M` > | |---|---| | LM head (N=248320, K=5120) | 529 | | MLP gate/up/down | 5957 | | any node with `N + K <= 16384` | >= 8192 (effectively never) | ### Changes | File | Change | |---|---| | `matmul_nbits.h` / `.cc` | Parse/validate the chunk size; per-node size gate (`FpAIntBRowsPerLaunch`); fpA_intB GEMM/GEMV loop over row chunks sharing one chunk-sized workspace, with at most two tactic lookups (full chunks + trailing chunk, which may use the GEMV below 16 rows); cap constructor profiling at `max(chunk, 256 MiB / ((N + K) * sizeof(T)))`; Level-2 `DeclareWorkspaceRequirements` uses the chunk rows; empty input (`M == 0`) now returns before runtime `g_idx` validation, which may sync the stream. | | `onnxruntime_session_options_config_keys.h` | New key `kOrtSessionOptionsCudaMatMulNBitsMChunkSize`. | | `docs/contrib_ops/cuda/matmul_nbits.md` | New §6.2 (condition, measurements, memory breakdown); env-var table. | | tests | C++ chunked-correctness test, Level-2/runtime size-gate test (`MatMulNBitsWorkspace.MChunkSizeGate`), Python parity / invalid-value / empty-input tests. | The dequantize + cuBLAS fallback and the fused GEMV are unchanged: neither has `M`-proportional scratch (the fallback's dequant buffer is already bounded by N chunking). ### Motivation and Context For an LM head, the fpA_intB tactic profiler allocates an `M x N` output buffer per profiled bucket, and a runtime `M` above the largest profiled bucket (2048 by default) triggers lazy profiling of that larger bucket at run time. Measured on H200 for the Qwen3.8-27B LM head alone (FP16, INT4, block 32, `kSameAsRequested` arena, NVML sampled from a separate process): | `M` | chunk | session-creation peak | run peak | session creation | first run | steady run | |---|---|---|---|---|---|---| | 2048 | off | 3112 MiB | 3080 MiB | 17.5 s | 584 ms | 78.4 ms | | 2048 | 256 | 2234 MiB | 3080 MiB | 8.2 s | 619 ms | 79.0 ms | | 8192 | off | | 10762 MiB | 18.5 s | 21716 ms | 464-480 ms | | 8192 | 2048 | | 5990 MiB | 18.3 s | 2239 ms | 316-324 ms | - Above the profiled range, chunking saves 4.7 GiB of peak and makes the first run ~10x faster; the remainder is mostly the `Y` output. - Within it, only the session-creation peak drops. The run peak is set by `Y`, both weight copies, and the CUDA context (breakdown in §6.2). ### Limitations / follow-ups - Default is off pending end-to-end benchmarks. A default that chunks to the largest profiled bucket when the size gate passes looks favorable from the data above. - The Level-1 (partition-time) estimate ignores chunking and stays a conservative upper bound. - Not exercised end to end with the plugin EP or CUDA graph capture. For graph capture, warmup must cover the trailing-chunk bucket, the same requirement as today. ### Testing Release build, CUDA 13.0, `CMAKE_CUDA_ARCHITECTURES=90-real`, `onnxruntime_USE_FPA_INTB_GEMM=ON` (compact), H200: - `onnxruntime_provider_test --gtest_filter='MatMulNBits*.*:MatMul2Bits*.*'`: 125 passed. - `onnxruntime_provider_test --gtest_filter='CUDA_EP_Unittest.*'` (with `onnxruntime_ENABLE_CUDA_EP_INTERNAL_TESTS=ON`): 291 passed, including `MatMulNBitsWorkspace.MChunkSizeGate` and the existing workspace agreement / zero-M tests. - `python -m pytest onnxruntime/test/python/quantization/test_op_matmulnbits_prepacked_cuda.py`: 7 passed, 9 skipped (offline weight packer not in the build). - `ORT_FPA_INTB_DEBUG=1` trace confirms a small node stays a single launch while the LM head at `M=2048`, chunk 256 runs 256-row chunks.
### Description Faling test on Windows on ARM fixed ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> Signed-off-by: melkap01 <melike.kaptan@arm.com>
### Description This PR fixes the file placement for a shared linear attention file. ### Motivation and Context The [original PR](microsoft#32512) placed the file in the wrong location.
ai-fw-intg
requested review from
Jaswanth51,
ankitm3k,
jatinwadhwa921 and
vthaniel
September 27, 2026 20:38
hdharpure9922
self-requested a review
September 28, 2026 03:54
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Automated daily backmerge from ORT main to ovep-develop. No conflicts detected. Do NOT squash or rebase - use merge commit only.