Skip to content

Restore the simple loop in fastfloat_strncasecmp - #412

Merged
lemire merged 1 commit into
mainfrom
strncasecmp-simple
Sep 22, 2026
Merged

lemire merged 1 commit into
mainfrom
strncasecmp-simple

Conversation

@lemire

@lemire lemire commented Sep 22, 2026

Copy link
Copy Markdown
Member

Summary

Reverts the SWAR case-insensitive compare introduced in #356 / #362 back to the plain 3/5-character loop, and removes the unused generic fastfloat_strncasecmp SWAR variant (−125 lines).

fastfloat_strncasecmp is only reached from parse_infnan, i.e. for nan / inf / infinity inputs, so it cannot speed up ordinary number parsing. What #356 did instead was grow the inlined inf/nan tail of the parser, which perturbed the hot path.

Investigation

The per-commit dashboard dropped by ~4% at #356 and #362 did not recover it (as expected: #362 only touched the cpp20_and_in_constexpr() branch).

Reproduced on an Intel Xeon (Gold 6548N), realbenchmark, 5 interleaved rounds, medians, at the #356 merge commit vs its parent:

GCC 14 base → #356 clang 21 base → #356
canada ASCII 64 1151 → 1135 MB/s, 314.5 → 321.4 i/f 918 → 901 MB/s, 381.7 → 389.7 i/f
mesh ASCII 64 862 → 851 MB/s, 176.3 → 179.1 i/f 714 → 696 MB/s, 225.7 → 230.8 i/f
UTF-16 flat flat

The new code never executes on these inputs; the extra instructions/float come from codegen side effects:

  • clang: parse_infnan was a separate 0x17f-byte function before; after optimize fastfloat_strncasecmp #356 it gets inlined and from_chars_float_advanced<double,char> grows 2307 → 2891 bytes.
  • GCC: parse_infnan was already inlined; the function only grew by 10 instructions, but the register allocator reshuffled the live path and added spills inside the digit loop.

On Apple M4 Max (Apple clang 17, hardware counters) the same change costs ±1 i/f — with 31 GPRs the added dead code creates no register pressure, which is why the PR author's M1 numbers looked flat.

Effect on current main

Measured against main at 6373592 (after #410), same machine, 5 interleaved rounds:

GCC 14 main → this PR clang 21 main → this PR
canada ASCII 64 199.82 → 199.82 i/f, 37.7 → 36.6 c/f 276.1 → 275.7 i/f
canada UTF-16 64 242.79 → 242.79 i/f 298.6 → 296.6 i/f
mesh ASCII 64 95.33 → 95.60 i/f 158.4 → 158.4 i/f
mesh UTF-16 64 109.68 → 109.79 i/f 168.2 → 166.8 i/f

#410 already moved the !pns.valid tail off the hot path, which absorbed the #356 regression on x86, so this PR is performance-neutral today (identical instruction counts with GCC, ±1–2 i/f with clang; clang's parse_infnan goes back out of line). It remains a simplification and removes code that only ever cost performance in the benchmarks.

Also tried: marking parse_infnan fastfloat_never_inline on top of this — clearly worse with GCC (+5–7 i/f), so not included.

Tests

  • ctest 15/15 (C++17), basictest with FASTFLOAT_CONSTEXPR_TESTS=ON (C++20)
  • extra static_asserts of the C++20 constexpr path for nan, NAN, -Inf, INFINITY, and the negative cases (nax, inx) — the branch fix early return error in fastfloat_strncasecmp #362 had to fix
  • -fsyntax-only -Wall -Wextra under C++11/14/20
  • clang-format: no new violations

PR #356 replaced the 3/5-character case-insensitive compare used by
parse_infnan with SWAR helpers (fastfloat_strncasecmp3/5 plus a generic
memcpy-based fastfloat_strncasecmp). That code only runs for "nan",
"inf" and "infinity" inputs, so it cannot help ordinary parsing, but it
grew the inlined inf/nan tail of the parser: on x86-64 (GCC 14 and
clang 21) the per-commit benchmarks dropped by 1.5-2.5% because the
extra dead code changed inlining decisions and register allocation on
the hot path (+7 instructions/float on canada.txt).

Go back to the plain loop, which the compiler unrolls for the constant
lengths 3 and 5, and drop the now-unused generic SWAR function. The
`actual < 256` guard of the original is unnecessary: OR-ing 0x20 into
a value can only equal an ASCII lowercase letter when the value is that
letter or its uppercase form, so the plain OR is exact for every
character type.
@lemire
lemire merged commit e9e04de into main Sep 22, 2026
54 of 66 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant