Conversation
Complexity-10 AArch64 NEON encoding with four delayed-decision states repeatedly copies one narrow state lane across the state rows. Preserve every fallback and state transition while scheduling that exact copy four rows at a time. Constraint: Owner authorized prior valid research performance evidence after the exact-production campaign remained invalid and ungraded Rejected: Import the research commit | it contains experiments, diagnostics, and unrelated candidates Confidence: high Scope-risk: narrow Directive: Keep the four-row 32-bit schedule and scalar remainder bound to c10/four-state AArch64 code generation Tested: Exact local and Graviton premeasurement differentials; OPUS_CHECK_ASM; ASan/UBSan; CMake/autotools/Meson; AArch64 codegen; Kimi K3 PASS Not-tested: Fresh exact-production performance did not complete; prior qualified research performance was used by explicit owner decision
Author
|
I tested this, output is exactly the same across many different files. |
This was referenced Sep 16, 2026
silk: Arm NEON for the float SILK analysis path (inner_product, warped autocorrelation, energy)
#481
Open
Author
|
Tested head: The earlier directly qualified Graviton4 comparison against stock found the useful Candidate-3 result: production c10 encode improved thread CPU by +4.0713% and wall by +4.0826%, with c10 p99.9 +4.7248% and an optimized-symbol diagnostic of +17.3194%. Positive percentages mean faster/lower cost. The c5/c9 controls were approximately zero. This used 28 blocks and 22,400 frames per label per preset, with exact packet, final-range, and output identity. This remains the benchmark baseline. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Complexity-10 AArch64 NEON encoding with four delayed-decision states repeatedly copies one narrow state lane across the state rows. Preserve every fallback and state transition while scheduling that exact copy four rows at a time.
On an m4 mac processor this increases transcoding speed by 8%, on a aws c8g graviton, about 4%.