Skip to content

Sync with Microsoft ONNX Runtime - 28092026 - #1312

Merged
hdharpure9922 merged 24 commits into
ovep-developfrom
sync_msft_28092026
Sep 28, 2026
Merged

hdharpure9922 merged 24 commits into
ovep-developfrom
sync_msft_28092026

Conversation

@ai-fw-intg

Copy link
Copy Markdown

Automated daily backmerge from ORT main to ovep-develop. No conflicts detected. Do NOT squash or rebase - use merge commit only.

kuanyul-qti and others added 24 commits September 24, 2026 14:44
### Motivation and Context
AutoEP can select and expose multiple `OrtEpDevice` instances for a
single session. Some of those devices may be backed by the same EP
factory, so querying each device for custom-op domains can return
repeated domains.

ORT gathers all EP-provided custom-op domains into one session-local
collection and registers them with the session's custom registry before
model load. That registry expects each domain to be unique. If the same
domain is registered twice, registration fails with "Domain already set
in registry", even though the duplicate came from equivalent AutoEP
devices rather than user configuration.

This mainly affects multi-device AutoEP sessions. Single-device AutoEP
sessions do not hit the issue because only one copy of the domain is
collected.

### Description
Deduplicate custom-op domains collected from AutoEP devices before
registering them with the session-local custom registry.

The change keeps only one copy of each EP-provided custom-op domain
while still skipping domains already supplied through session options.

This ports the AutoEP custom-op domain deduplication from
onnxruntime/onnxruntime-qnn#608.
### Description

Packed GroupQueryAttention shape inference divided the query hidden size
by `num_heads + 2 * kv_num_heads` without validating the attributes
first. Malformed head counts could therefore cause division by zero or
signed overflow while resolving a model.

Validate both counts before the calculation, reject grouped-head
overflow, require concrete hidden sizes to divide evenly, and preserve
symbolic hidden dimensions.

### Testing

- `cmake --build build\Windows\Release --config Release --target
onnxruntime_graph --parallel`
- `cmake --build build\Windows\Release --config Release --target
onnxruntime_provider_test --parallel`
- `onnxruntime_provider_test.exe
--gtest_filter=GroupQueryAttentionTest.CausalMaskDefaultsToEnabled_CPU`

---------

Co-authored-by: Daniel Song <danielsong@microsoft.com>
…soft#32786)

### Description

In the 2-bit `ParQuantizeLinearStd` helper, return once the leading
partial byte has covered the whole scale interval. Before this change,
such an interval also went through the trailing partial-byte path, which
re-quantized the preceding elements in the same byte with the current
scale and zero point.

Adds a per-axis INT2 regression test where each scale interval is a
single element.

### Motivation and Context

With float input and per-axis quantization on the last axis (for example
a `[K, N]` tensor with one scale per column), each scale interval holds
one element, so intervals start and end inside packed bytes. The first
two elements of each byte are then re-quantized with the scale and zero
point of a later element in that byte, without any error. Intervals with
more than one element, FP16 input and INT4 outputs are not affected.

Repro with onnxruntime 1.30.0:

```python
import numpy as np
import onnxruntime as ort
from onnx import TensorProto, helper, numpy_helper

nodes = [
    helper.make_node("QuantizeLinear", ["x", "scale"], ["q"], axis=1, output_dtype=TensorProto.UINT2),
    helper.make_node("DequantizeLinear", ["q", "one"], ["y"]),
]
inits = [
    numpy_helper.from_array(np.array([1, 1, 4, 1, 1, 1, 4, 1], np.float32), "scale"),
    numpy_helper.from_array(np.array(1, np.float32), "one"),
]
graph = helper.make_graph(nodes, "g", [helper.make_tensor_value_info("x", TensorProto.FLOAT, [2, 8])],
                          [helper.make_tensor_value_info("y", TensorProto.FLOAT, [2, 8])], inits)
model = helper.make_model(graph, opset_imports=[helper.make_opsetid("", 25)], ir_version=10)
so = ort.SessionOptions()
so.graph_optimization_level = ort.GraphOptimizationLevel.ORT_DISABLE_ALL
sess = ort.InferenceSession(model.SerializeToString(), so, providers=["CPUExecutionProvider"])
print(sess.run(None, {"x": np.ones((2, 8), np.float32)})[0])
# Expected: [[1 1 0 1 1 1 0 1], [1 1 0 1 1 1 0 1]]
# Actual:   [[0 0 0 1 0 0 0 1], [0 0 0 1 0 0 0 1]]
```

Found while working on microsoft#32735, which routes more cases through this
helper.
### Description

- Batch JSEP WebGPU Concat inputs using:
  ```ts
  maxInputsPerDispatch = maxStorageBuffersPerShaderStage - 1
  ```
- Reuse the output buffer across dispatches and preserve concat-axis
offsets.
- Add an 8-input regression case to the WebGPU suite.

### Motivation and Context

Concat bound every input and the output in one shader. At eight inputs,
this exceeded the device’s eight-storage-buffer limit, silently skipped
execution, and returned zeros.

<!-- START COPILOT CODING AGENT SUFFIX -->

- Fixes microsoft#32757

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: tianleiwu <30328909+tianleiwu@users.noreply.github.com>
…ns. (microsoft#32768)

Split the webgpu build & test stage into 2 stages. This lets us use
better machines for build. In total, this cuts the action time between
30 and 45 minutes.

<img width="515" height="247" alt="image"
src="https://github.com/user-attachments/assets/2088079a-1060-4cde-baf8-9e04585eb578"
/>

->

<img width="1024" height="299" alt="image"
src="https://github.com/user-attachments/assets/aeb1d7d8-283b-437e-af39-5708ce406167"
/>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: eserscor <247253654+eserscor@users.noreply.github.com>
Copilot-Session: 409e405a-fdd1-4c73-b54a-4676012fb9e5
## Description

Adds an unquantized WebGPU `MoE` kernel and expands WebGPU `QMoE`
support for the configurations needed by packed token-major inference,
including Qwen 3.8 Flash.

The existing operator schemas already accept both representations:

- Packed 2D input: `(num_tokens, hidden_size)`
- Dense 3D input: `(batch_size, sequence_length, hidden_size)`

`MoE` and `QMoE` are token-local, so concatenating variable-length
requests does not require a separate varlen operator or a cumulative
sequence-length input. Each router tensor contains one row per input
token, and the output preserves the input shape.

## Changes

### WebGPU MoE

- Implements and registers the unquantized WebGPU `MoE` kernel for float
and float16.
- Supports packed 2D and dense 3D inputs.
- Supports ReLU, GELU, SiLU, identity, and SwiGLU activations.
- Supports separate FC3 gating and interleaved or concatenated fused
SwiGLU layouts.
- Handles empty inputs, chunked multi-token execution, and hidden sizes
that are not multiples of the workgroup size.
- Uses segmented-buffer-safe tensor access throughout the WebGPU shader
pipeline, including stacked expert weights and shape-preserving outputs.

### WebGPU QMoE

- Adds all schema activation types and both fused SwiGLU layouts.
- Adds separate FC3 projection and gating.
- Adds aggregation-only `router_weights` while continuing to use
`router_probs` for Top-K selection, including selected Top-K
normalization when `normalize_routing_weights=1`.
- Adds block-wise explicit zero-point support.
- Makes stacked `MatMulNBits` zero-point indexing follow the selected
expert, matching weight and scale indexing.
- Preserves the optimized single-token decode path.
- Rejects unsupported row-wise and optimized single-token
explicit-zero-point combinations with clear errors.
- Validates that `num_experts` fits the adapter's workgroup-size and
invocation limits before creating the gate shader.

### Packed semantics and tests

- Clarifies packed 2D and dense 3D behavior in the authoritative schemas
and generated contrib documentation.
- Adds a packed ragged-batch test representing request lengths `[2, 1,
2]`.
- Uses distinct token rows and nonzero weights so the packed test
detects hidden-state reordering or cross-request interaction.
- Disables CPU fallback in every WebGPU-specific test so registration or
kernel-claim regressions cannot pass through another EP.
- Adds focused coverage for float/float16 MoE, dense and packed inputs,
activations, FC3, SwiGLU fusion layouts, expert-specific zero points,
normalized `k > 1` router weights, non-aligned gather tails, the
2,048-token chunk boundary, adapter expert limits, empty inputs, and the
optimized QMoE decode path.

## Provider-specific scope

This change targets portable WebGPU functionality. CUDA-specific `fp4`,
`nvfp4`, `fp8`, and `wfp4afp8` formats, CUDA expert-weight prepacking,
and sparse mixer execution remain unsupported on WebGPU and are rejected
explicitly.

The normal symmetric INT4 Qwen path remains supported. Block-wise
asymmetric INT4/INT8 weights, separate router weights, and FC3-based
gating are now also represented by the WebGPU implementation and tests.

The official Qwen 3.8 Flash FP8 checkpoint uses 128x128 block-scaled FP8
expert weights. CUDA `QMoE` currently supports FP8 weights with one
global scale per expert, while its existing block-scaled modes are
`nvfp4` and `wfp4afp8`. Supporting the official block-FP8 checkpoint
therefore requires a separate CUDA/schema change and is intentionally
left to a focused follow-up rather than expanding this WebGPU PR.

## Validation

- Windows Release WebGPU provider and `onnxruntime_provider_test`
targets build successfully with Dawn and the D3D12 Agility SDK.
- WGSL template tests pass on Windows and Linux CI.
- Deterministic nonzero WebGPU MoE and optimized single-token QMoE tests
pass locally.
- `clang-format --dry-run --Werror` passes for changed C++ files.
- `git diff --check` passes.
- Editor diagnostics report no errors.
- Generated `OperatorKernels.md` content is synchronized with CI output.

Multi-token QMoE dispatch reaches the local D3D12 adapter, but repeated
local runs can end in `DXGI_ERROR_DEVICE_REMOVED`; cross-adapter runtime
validation is covered by WebGPU CI.

---------

Co-authored-by: Sayan Shaw <sayanshaw@microsoft.com>
Co-authored-by: Justin Chu <justinchuby@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…rosoft#32767)

# WebGPU GatedDeltaNet Compact State Updates

## Changes

This PR adds compact state-update capture to the native WebGPU
`com.microsoft.GatedDeltaNet` kernel, which previously rejected
`state_update_capacity > 0`.

- Validates `state_update_capacity` and the required `[batch]`
`capture_count` input.
- Allocates the packed float32 `state_update` output and includes its
buffers in WebGPU binding
  accounting.
- Captures the clamped per-row prefix of decay, shared-key, and delta
factors in the recurrent
  shader. These factors replay `S = decay * S + outer(key, delta)`.
- Preserves inverse-GQA key sharing and existing value-channel tiling.
- Uses the recurrent path while capture is active because parallel
prefill does not emit per-token
  transitions.
- Honors `state_update_active=0` by clearing the entire output;
uncaptured entries remain
  unspecified when capture is active.
- Updates the registered schema text and generated contrib-operator
documentation to advertise
  WebGPU compact-state support.

## Test Coverage

Tests now compare only the captured decay/key/delta ranges instead of
assuming uncaptured storage is
zero. The inactive test still verifies that the complete output is
zero-filled.

Coverage includes ragged capacity-7 cases with zero, partial, and full
capture counts, plus an FP16
Qwen-like case (`Hq=2`, `Hv=6`, `head_size_qk=head_size_v=128`). The
larger case exercises the
FP16 aliases, value tiling, shared-key head mapping, and binding-packing
path used by GenAI.

## DFlash2 Prerequisite

This PR adds the ORT kernel prerequisite; it does **not** enable
end-to-end WebGPU DFlash2 by itself.
GenAI also needs:

- builder export with nonzero `state_update_capacity`
([microsoft#2453](microsoft/onnxruntime-genai#2453),
[microsoft#2585](microsoft/onnxruntime-genai#2585)); and
- compact-state replay after partial acceptance
([microsoft#2454](microsoft/onnxruntime-genai#2454)),
including WebGPU `ReplayStateUpdates`
([microsoft#2606](microsoft/onnxruntime-genai#2606)).

Before advertising WebGPU DFlash2 support, an end-to-end GenAI run
should cover full, partial, and
zero draft acceptance, match non-speculative output/state, and confirm
that neither target execution
nor replay silently falls back from WebGPU.

---------

Co-authored-by: Prathik Rao <prathikrao@microsoft.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
### Description

Add an AVX2/FMA/F16C MLAS kernel for FP16 LayerNorm and RMSNorm on x86.
The kernel widens FP16 inputs in vectors, performs the reductions and
affine transform in FP32 (with a double-precision mean accumulation for
LayerNorm), and narrows the result back to FP16 in vectors.

The CPU LayerNorm implementation dispatches to this kernel for rows of
at least 16 elements when the processor supports the required
instructions. Other architectures, older x86 processors, short rows,
BF16, and unsupported output-statistics types retain the existing scalar
path.

### Motivation and Context

The existing FP16 CPU path converts every input and output element
individually around scalar normalization math. Transformer workloads
commonly execute LayerNorm and RMSNorm over rows of 1K-4K elements,
where conversion and scalar arithmetic dominate execution time.

The benchmark processes 4,194,304 FP16 elements per invocation with rows
parallelized across OpenMP workers, mirroring ORT's per-row thread-pool
parallelism. It was pinned to one 48-core Intel Xeon Platinum 8480C
socket with workers bound to cores. The table reports the median speedup
from three process runs; graph/session overhead is intentionally
excluded.

| Operation | Row size | 1 thread | 8 threads | 16 threads | 32 threads
|
|---|---:|---:|---:|---:|---:|
| LayerNorm | 1024 | 15.66x | 15.39x | 15.85x | 14.89x |
| RMSNorm | 1024 | 12.07x | 11.91x | 11.96x | 10.63x |
| LayerNorm | 4096 | 14.96x | 15.72x | 14.41x | 13.31x |
| RMSNorm | 4096 | 11.71x | 11.42x | 10.84x | 9.61x |

At 32 threads, the optimized kernels reach median throughputs of
54.0-57.6 billion elements/s for LayerNorm and 81.2-89.4 billion
elements/s for RMSNorm.

Validation:
- Built `onnxruntime_mlas_test` and `onnxruntime_provider_test` in a CPU
Release configuration.
- `MlasLayerNormF16Test.MatchesFloatReference` passes for LayerNorm
with/without bias and RMSNorm at row sizes 16, 31, and 128.
- `RMSNormalizationOpTest.RMSNorm_float16_OptimizedRow` passes through
the CPU execution provider.
- Existing `RMSNormalizationOpTest.*` suite passes (26 tests before
adding the dedicated optimized-row case).
- Lintrunner passes on all changed files.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
### Description

The CPU N-D Col2Im loop computes image and column indices for each
value. Compute the valid range on the last axis once per kernel tap,
then add each valid row. Advance the input row pointer without another
index calculation.

Keep the order of additions. For windows that can overlap, use the
original loop if the output contains NaN. This keeps NaN payloads. Skip
the NaN scan when the windows cannot overlap, so large gaps in the
output do not add a scan cost.

This applies to FP32 CPU ConvTranspose with 1D or 3D spatial inputs and
other users of `Col2imNd`. The 2D helper from microsoft#32698 and GEMM do not
change.

### Motivation and Context

Measurements through the C API on a Core Ultra 7 255H, WSL2, GCC 13.3,
Release, CPU EP, one thread. The base is main at `056897df`. The table
gives the median of seven rounds with fixed CPU affinity and alternating
run order. These are repeated runs after session setup. Timing includes
GEMM, output initialization, and bias. It excludes session setup and
prepacking of constant weights.

| Case | main → change, ms | Speedup |
|---|---:|---:|
| ConvTranspose 1D, k16 / s8 | 1.363 → 0.470 | 2.90× |
| ConvTranspose 3D, k3×3×3 / s1 | 13.821 → 1.614 | 8.56× |
| Col2Im 3D, k3×3×3 / s1 | 12.811 → 0.940 | 13.63× |
| Synthetic 3D decoder, 3 layers | 21.516 → 2.575 | 8.35× |

For ConvTranspose, the 1D input is `[1,16,2048]`, with 16 output
channels and padding 4 on each side. The 3D input is `[1,8,16,24,32]`,
with 8 output channels and padding 1 on each side. Col2Im uses column
shape `[1,216,12288]` and image shape `[16,24,32]`. The decoder is a
synthetic graph with input `[1,8,8,16,24]`; it is not a trained model.

Overlapping NaN inputs require a second pass. The 3D Col2Im case takes
13.403 → 17.677 ms with mixed special values. An output with infinity
and no NaN does not need a second pass. Small-input timing varies; the
unchanged 2D control and same-binary checks also vary by a few percent.

The Math and CPU ConvTranspose/Col2Im tests pass. Two new tests cover
row bounds, addition order, and special values with and without overlap.
Additional local checks cover groups, dilation, asymmetric padding, 1
and 4 threads, and output equality. Checks with fixed and runtime
weights also pass, with no bias and with bias followed by ReLU. The
extracted function also passes ASan and UBSan.

CPU `ConvGrad` gradient outputs also match the base with training
operators enabled. A full training loop was not tested.
## Summary
- add an `adapterIndex` provider option for selecting a zero-based
physical Dawn adapter
- enumerate adapters matching the requested backend and power preference
in native Dawn builds
- reject unsupported external-Dawn/wasm combinations, out-of-range
indices, invalid values, and conflicting context reuse
- document the distinction between physical `adapterIndex` and ORT
`deviceId`

## Testing
- `clang-format --dry-run --Werror` on all changed C++ files
- fresh current-main native Dawn configuration completed
- `webgpu_context.cc` and `webgpu_provider_factory.cc` compiled
successfully against Dawn v20260818.211311
- full local plugin link was blocked by unrelated existing GCC 13
`-Werror=maybe-uninitialized` failures in WebGPU Gemm/InstanceNorm; PR
CI covers the supported build matrix
…osoft#32763)

**Description**

Enable  com.microsoft::PagedAttention  on the WebGPU EP with
 is_causal=0  when  local_window_size <= 0 .

This is needed by block-drafting models such as DFlash2, where all query
tokens submitted in one speculative block must attend bidirectionally to
the complete live context. Previously, the WebGPU kernel rejected every
non-causal configuration and then hard-coded causal behavior when
constructing  WebgpuAttentionParameters .

**Changes**

- Store and propagate the  is_causal  attribute through WebGPU
PagedAttention into the existing FlashAttention  is_unidirectional 
shader specialization.
 - Preserve the existing causal and causal local-window behavior.
- Continue to reject  is_causal=0  with  local_window_size > 0 ;
query-relative left-window masking for that combination is intentionally
out of scope.
- Generalize the C++ reference implementation to compare causal prefix
visibility with non-causal full-context visibility.
 - Add focused WebGPU parity tests covering:
  - direct paged decode with a multi-token speculative block;
  - fused paged prefill;
  - gather-then-FlashAttention fallback;
  - ragged requests with different query and past lengths.
- Replace the Python non-causal rejection test with reference parity
coverage and add a dedicated test for the unsupported non-causal
local-window combination.

No WGSL redesign is required: the dense and paged FlashAttention shaders
already support causal and non-causal specializations. This change only
restores the missing host-side attribute plumbing.

**Masking behavior**

 -  is_causal=1 : query row  i  attends through  past_length + i .
-  is_causal=0 : every query row attends through  past_length +
query_length , including later tokens in the same submitted block.
 -  is_causal=0 && local_window_size>0 : remains explicitly unsupported.

**Validation**

 -  git diff --check 
 - Python syntax compilation for  test_paged_attention.py 
- Source-level checks for attribute propagation and focused test
coverage

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
### Description

Extends `ConvBNFusion` to support fusing `ConvTranspose ->
BatchNormalization`.

The existing fusion path only handled `Conv -> BatchNormalization`. This
change adds `ConvTranspose` as a supported target and folds
BatchNormalization parameters into ConvTranspose when the required
initializers are constant and the channel shapes are compatible.

For ConvTranspose, the weight layout is `[C_in, C_out/group, ...]`, so
the BatchNormalization scale is applied by absolute output channel
instead of using the existing Conv weight-axis scaling path. The
existing Conv fusion behavior is unchanged.

This also adds graph transformer tests for:

- ConvTranspose without bias
- ConvTranspose with bias
- grouped ConvTranspose
- no fusion when ConvTranspose weight is not constant

### Motivation and Context

`Conv -> BatchNormalization` is already fused by `ConvBNFusion`, but the
same optimization was not applied to `ConvTranspose ->
BatchNormalization`. As a result, eligible ConvTranspose models kept an
extra BatchNormalization node after graph optimization.

Fusing the BatchNormalization into ConvTranspose removes the redundant
node and makes ConvTranspose consistent with the existing Conv fusion
behavior while preserving numerical equivalence.

Fixes microsoft#14270
This pull request strengthens input validation and error handling for
the `ConvTransposeWithDynamicPads` operator. The main focus is on
ensuring that the required dynamic pads tensor is always provided, both
at schema and runtime, and that missing or malformed inputs are handled
gracefully. Additionally, a new test is added to verify the operator's
behavior when the pads input is missing.

**Input validation and schema enforcement:**

* The operator schema for `ConvTransposeWithDynamicPads` is updated to
make the `Pads` input required instead of optional, ensuring that all
usage must provide this tensor.
* The shape inference function now immediately returns if fewer than
three inputs are provided, preventing access to missing inputs.
* In the operator implementation, a check is added to explicitly return
an error if the `Pads` tensor is missing when dynamic padding is
expected.
* The `getInputData` method in `InferenceContextImpl` now safely returns
`nullptr` if the requested input index is out of bounds, preventing
potential crashes.

**Testing improvements:**

* A new unit test, `ConvTransposeWithDynamicPads_MissingPadsRejected`,
is added to verify that the operator correctly rejects models where the
required `Pads` input is missing.

**Test infrastructure:**

* Includes necessary headers in the test file to support the new test
case.
### Description

Constrain the Python documentation ONNX dependency to versions before
1.23.

`docs/python/examples/plot_pipeline.py` imports `onnx.tools.net_drawer`,
which was removed in ONNX 1.23. The documentation requirements now
resolve to ONNX 1.21 or 1.22 only.

### Motivation and Context

The Python API documentation workflow currently fails while building the
Sphinx gallery because the unbounded `onnx >= 1.21.0` requirement
resolves to ONNX 1.23, where `onnx.tools.net_drawer` is unavailable.

This restores generation of the Python API-doc artifact required by the
GitHub Pages publication workflow.

Signed-off-by: Damien Dooley <damien.dooley@arm.com>
### Description
<!-- Describe your changes. -->
Fix warning from recent change in microsoft#32729

> LINK : error C2220: the following warning is treated as an error
[N:\_work\1\b\RelWithDebInfo\onnxruntime_provider_test.vcxproj]
LINK : warning C4744: 'static char const
absl::lts_20250814::base_internal::FastTypeTag<class
std::basic_string<char,struct std::char_traits<char>,class
std::allocator<char> > >::kDummyVar' has different type in
'n:\_work\1\b\relwithdebinfo\_deps\googletest-src\googletest\src\gtest-all.cc'
and
'n:\_work\1\s\onnxruntime\test\providers\webgpu\webgpu_context_test.cc':
'unsigned char' and 'signed char'
[N:\_work\1\b\RelWithDebInfo\onnxruntime_provider_test.vcxproj]


### Motivation and Context
<!-- - Why is this change required? What problem does it solve?
- If it fixes an open issue, please link to the issue here. -->
Fix packaging pipeline
)

### Description

The CUDA Plugin EP factory handed **one BFC device arena per GPU** to
every `CreateAllocator` caller: every session, plus the environment's
shared allocator. After a CUDA graph is captured, the session frees the
graph's intermediate chunks back to the arena, but the captured graph
still reads and writes those addresses on every replay. Because the
arena was shared, another session could be given those chunks and
overwrite them between replays.

With two graph-captured sessions on one device (a Qwen3.8-27B target
plus a DFlash2 drafter in onnxruntime-genai), this silently corrupted
generation. MMLU-Pro dropped 5.5 pp and speculative acceptance fell from
87.7% to 69.8%. The bundled CUDA EP is not affected because it creates a
separate arena for each session.

This PR gives each `CreateAllocator` call for device memory its own
arena, which matches the bundled EP. It replaces the temporary
graph-disable gate in microsoft#32801.

#### Changes

| File | Change |
|------|--------|
| `cuda_ep_factory.h/.cc` | Replace the single `device_arena` and its
refcount with `std::vector<DeviceArena> device_arenas`. Each entry
tracks its own `has_quarantine` / `abandoned` state.
`CreateAllocatorImpl` creates a new arena on every call and now honors
that session's arena options, which the shared arena used to ignore with
a warning. `ReleaseAllocatorImpl` destroys only the matching arena, or
leaks it on purpose if it is quarantined or abandoned. Stream run-end
reset and quarantine/abandon after an undrained release now visit every
arena on the device. For abandon, this is the same scope the single
shared arena had before. `GetDeviceArenaForDevice()` had no callers and
is removed. |
| `cuda_plugin_arena_test.cc` | New `DeviceAllocator_IsSessionScoped`
test: an allocation from session A must not show up in session B's arena
stats or the environment allocator's stats. It **fails on the current
shared-arena plugin** (`NumAllocs` 3 vs 2) and passes with this change.
|
| `docs/cuda_plugin_ep/*.md` | Describe the per-session device arena and
why CUDA graphs require it. |

The pinned arena and the CUDA mempool allocator are still shared per
device and are unchanged.

### Motivation and Context

#### Root-cause isolation (MMLU-Pro, 200 questions, 8,192 max new
tokens, DFlash2 drafts=6, concurrency 8)

| Configuration | Correct | Cap hits | Acceptance |
|---|---:|---:|---:|
| Bundled CUDA EP, graphs on | 168 | 3 | 87.7% |
| Plugin (current), graphs on | 157 | 64 | 69.8% |
| Plugin, target session graph only | 156 | 66 | 69.4% |
| Plugin, drafter session graph only | 167 | 7 | 86.2% |
| Plugin, per-session arena prototype, both graphs on | 169 | 4 | 87.8%
|
| Same prototype binary, shared arena (control) | 156 | 64 | 70.1% |

Toggling only the arena sharing, with the same binary and both graphs
on, flips the result. That identifies the shared arena as the cause.

#### This PR (graphs on for both sessions)

| Metric | Bundled | Plugin with this PR |
|---|---:|---:|
| MMLU-Pro correct | 168/200 | **168/200** |
| Cap hits | 3 | 3 |
| Acceptance | 87.7% | 87.8% |
| Paired vs bundled | — | 1 win / 1 loss, McNemar p=1.0, 183/200
token-exact |
| Paired vs current plugin | — | 13 wins / 2 losses, p=0.0074 |

#### Performance

Batch 1, 32,768 prompt tokens, 1,024 generated tokens, DFlash2 draft
width 6, 3 fresh-process repetitions, all arms run at the same time on
separate H200s:

| Runtime | Median decode tok/s | Median prefill tok/s | Median TTFT |
Peak workload |
|---|---:|---:|---:|---:|
| Bundled CUDA EP | 314.58 | 3,767.2 | 8.698 s | 31,590 MiB |
| Plugin, graphs disabled (microsoft#32801 gate) | 305.65 (0.972x) | 3,822.9 |
8.572 s | 30,900 MiB |
| **Plugin, this PR** | **317.01 (1.008x)** | 3,780.7 | 8.667 s | 31,526
MiB |

Re-enabling graphs safely recovers the 2.8% decode loss from the gate.
Peak memory is 64 MiB below the bundled EP.

#### Tests

- `CudaPluginArenaTest.*`, `CudaPluginUserStreamGraphTest.*`,
`CudaPluginDeviceDiscoveryTest.*`: 51/51 passed on an H200 (Release,
SM90).
- `clang-format --dry-run --Werror` and `git diff --check` pass.
…oft#32810)

### Description

Adds optional M (row) chunking to the CUDA `MatMulNBits` fpA_intB path
to bound memory that grows with `M`, targeting large-vocabulary LM heads
during chunked prefill. It is off by default and size-gated, so small
layers are never split.

| Setting | Meaning |
|---|---|
| `ep.cuda.matmul_nbits_m_chunk_size` (session config) /
`ORT_MATMULNBITS_M_CHUNK_SIZE` (env) | Max rows of `A` per fpA_intB
launch. `0`/unset disables. Session config wins. Also works for the CUDA
plugin EP. |
| `ORT_MATMULNBITS_FORCE_CHUNKED=1` (existing) | Now also bypasses the
M-chunking size gate (it already bypassed the fallback's size guards). |

A node is chunked only when **`M > chunk`** and **`M * (N + K) *
sizeof(T) > 256 MiB`** (the tactic profiler's `A` and `C` buffers for an
unchunked `M`; same budget as the existing N-chunked fallback). For
Qwen3.8-27B in fp16/bf16:

| node | chunked when `M` > |
|---|---|
| LM head (N=248320, K=5120) | 529 |
| MLP gate/up/down | 5957 |
| any node with `N + K <= 16384` | >= 8192 (effectively never) |

### Changes

| File | Change |
|---|---|
| `matmul_nbits.h` / `.cc` | Parse/validate the chunk size; per-node
size gate (`FpAIntBRowsPerLaunch`); fpA_intB GEMM/GEMV loop over row
chunks sharing one chunk-sized workspace, with at most two tactic
lookups (full chunks + trailing chunk, which may use the GEMV below 16
rows); cap constructor profiling at `max(chunk, 256 MiB / ((N + K) *
sizeof(T)))`; Level-2 `DeclareWorkspaceRequirements` uses the chunk
rows; empty input (`M == 0`) now returns before runtime `g_idx`
validation, which may sync the stream. |
| `onnxruntime_session_options_config_keys.h` | New key
`kOrtSessionOptionsCudaMatMulNBitsMChunkSize`. |
| `docs/contrib_ops/cuda/matmul_nbits.md` | New §6.2 (condition,
measurements, memory breakdown); env-var table. |
| tests | C++ chunked-correctness test, Level-2/runtime size-gate test
(`MatMulNBitsWorkspace.MChunkSizeGate`), Python parity / invalid-value /
empty-input tests. |

The dequantize + cuBLAS fallback and the fused GEMV are unchanged:
neither has `M`-proportional scratch (the fallback's dequant buffer is
already bounded by N chunking).

### Motivation and Context

For an LM head, the fpA_intB tactic profiler allocates an `M x N` output
buffer per profiled bucket, and a runtime `M` above the largest profiled
bucket (2048 by default) triggers lazy profiling of that larger bucket
at run time. Measured on H200 for the Qwen3.8-27B LM head alone (FP16,
INT4, block 32, `kSameAsRequested` arena, NVML sampled from a separate
process):

| `M` | chunk | session-creation peak | run peak | session creation |
first run | steady run |
|---|---|---|---|---|---|---|
| 2048 | off | 3112 MiB | 3080 MiB | 17.5 s | 584 ms | 78.4 ms |
| 2048 | 256 | 2234 MiB | 3080 MiB | 8.2 s | 619 ms | 79.0 ms |
| 8192 | off | | 10762 MiB | 18.5 s | 21716 ms | 464-480 ms |
| 8192 | 2048 | | 5990 MiB | 18.3 s | 2239 ms | 316-324 ms |

- Above the profiled range, chunking saves 4.7 GiB of peak and makes the
first run ~10x faster; the remainder is mostly the `Y` output.
- Within it, only the session-creation peak drops. The run peak is set
by `Y`, both weight copies, and the CUDA context (breakdown in §6.2).

### Limitations / follow-ups

- Default is off pending end-to-end benchmarks. A default that chunks to
the largest profiled bucket when the size gate passes looks favorable
from the data above.
- The Level-1 (partition-time) estimate ignores chunking and stays a
conservative upper bound.
- Not exercised end to end with the plugin EP or CUDA graph capture. For
graph capture, warmup must cover the trailing-chunk bucket, the same
requirement as today.

### Testing

Release build, CUDA 13.0, `CMAKE_CUDA_ARCHITECTURES=90-real`,
`onnxruntime_USE_FPA_INTB_GEMM=ON` (compact), H200:

- `onnxruntime_provider_test
--gtest_filter='MatMulNBits*.*:MatMul2Bits*.*'`: 125 passed.
- `onnxruntime_provider_test --gtest_filter='CUDA_EP_Unittest.*'` (with
`onnxruntime_ENABLE_CUDA_EP_INTERNAL_TESTS=ON`): 291 passed, including
`MatMulNBitsWorkspace.MChunkSizeGate` and the existing workspace
agreement / zero-M tests.
- `python -m pytest
onnxruntime/test/python/quantization/test_op_matmulnbits_prepacked_cuda.py`:
7 passed, 9 skipped (offline weight packer not in the build).
- `ORT_FPA_INTB_DEBUG=1` trace confirms a small node stays a single
launch while the LM head at `M=2048`, chunk 256 runs 256-row chunks.
### Description
Faling test on Windows on ARM fixed 



### Motivation and Context
<!-- - Why is this change required? What problem does it solve?
- If it fixes an open issue, please link to the issue here. -->

Signed-off-by: melkap01 <melike.kaptan@arm.com>
### Description

This PR fixes the file placement for a shared linear attention file.

### Motivation and Context

The [original PR](microsoft#32512)
placed the file in the wrong location.

@hdharpure9922 hdharpure9922 left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@hdharpure9922
hdharpure9922 merged commit ce7b893 into ovep-develop Sep 28, 2026
7 of 10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.