Opt-in joint/two-stage grid search (arXiv 2505.17595) for the
optimized-RTN path, carried by RTNConfig and SignRoundConfig:
- asym: joint (scale, integer zero-point) two-stage search per group;
asymmetric layers had no optimized initializer before (plain min/max
RTN only). sym: two-stage signed scale search replacing the incumbent
clip search; the winning clamp convention is searched per group.
- zero-shot (iters=0): both searches weighted by the activation imatrix,
collected automatically. iters>0: the search anchors the SignRound
tuning grid (range margins pinned, tuning proceeds on rounding values);
the anchor consumes the imatrix when the tuning path collects one
(SignRoundV2 enable_alg_ext), unweighted otherwise.
- backends: extension Triton kernels (full sweep + shared-multiplier
coarse pass), torch.compile, eager fallback; a failing backend is
disabled for the rest of the process with a warning. Benchmark (2M
groups, g128, 64+32 candidates, RTX 3090): eager 23.8s, compile 0.43s
(~55x), Triton 0.33s (~72x).
- MoE: same-shape expert projections quantize in one stacked call,
bit-identical to per-module.
- AR_NEUQI_COARSE/AR_NEUQI_FINE: backend-aware defaults (256/64 on the
Triton/compile backends, 64/32 with eager); AR_NEUQI_BACKEND pins a
backend for debugging.
Ported onto the datatype-refactor lifecycle (intel#2360): the search and the
tuning anchor engage inside _IntWeightQuantizer.create_state via
WeightQuantizationSpec.enable_neuqi (zero-shot asym keeps the searched
(scale, zp) verbatim for exact qdq reproduction; sym zero-shot swaps the
optimized_init source; the tuned path anchors tensor_min/tensor_max and
freezes the range margins). The AWQ clip composes on sym zero-shot and is
overridden with a warning on the anchor paths.
Validation (full-vocab fp32 KL vs the BF16 base, openwebtext 8x8192,
Qwen3.8-27B W4A16 g128, single RTX 3090; details in docs/neuqi_acc.md):
sym iters=0 -4.1% mean KL and -32% wall vs the incumbent search
(0.02734/9m vs 0.02852/13m19s); SignRoundV2 iters=50 sym is
quality-neutral vs the same recipe without NeUQI (0.02191 vs 0.02203 -
the anchor is free); asym 0.02579 (iters=0), 0.02106 (iters=50),
0.01921 (iters=200).
Signed-off-by: avtc <tarasenkov@gmail.com>
Description
This PR adds the NeUQI grid search (arXiv 2505.17595) to the optimized-RTN path, opt-in via
--enable_neuqi(CLI, orenable_neuqi=Truein the API; carried by bothRTNConfigandSignRoundConfig):iters=0) both searches are weighted by the activation imatrix, collected automatically when the flag is set (asymmetric layers previously never consumed one — plain min/max RTN was their only path). Withiters > 0the anchor runs unweighted; tuning itself is data-driven.iters > 0: frozen-init anchor. The search runs once at wrapper init and anchors the SignRound tuning grid — range margins are pinned to the winner, tuning proceeds on the rounding values only. Measured quality-neutral atiters=50on sym with SignRoundV2 (0.02191 vs 0.02203 without), so the anchor is a no-cost initialization once tuning runs.auto_round_extension), a torch.compile implementation, or plain eager PyTorch — and falls back to the next one if a backend fails at runtime (the failed one is disabled for the rest of the process, with a warning).AR_NEUQI_BACKEND=triton|compile|eagerpins a backend, mainly for debugging. Measured on the joint search (2M groups, group size 128, 64+32 candidates, RTX 3090): eager 23.8 s, torch.compile 0.43 s (~55x faster than eager), Triton 0.33 s (~72x).AR_NEUQI_COARSE/AR_NEUQI_FINEcontrol the grid. Defaults are 256/64 when Triton or torch.compile serves the search and 64/32 with eager — the eager cost is proportional to the grid size and measured quality is flat between the two (Qwen3.8-27B W4A16 g128, iters=0: mean KL 0.02570 at 64/32 vs 0.02579 at 256/64). Explicit values always win.Validation — full-vocabulary fp32 KL vs the BF16 base, openwebtext 8x8192 positions; Qwen3.8-27B, W4A16, group size 128, single RTX 3090. Protocol and full statistics: docs/neuqi_acc.md. Sorted by mean KL (lower is better):
All
iters > 0rows use SignRoundV2 (enable_alg_ext).Type of Change
New feature
Related Issues
Relates to #2010 — NeUQI support was requested there.
Checklist Before Submitting