Skip to content

fix: lone high surrogate decoded as U+0000 instead of an error - #482

Open
ddxd wants to merge 1 commit into
simd-lite:mainfrom
leocaolab:fix/lone-high-surrogate
Open

ddxd wants to merge 1 commit into
simd-lite:mainfrom
leocaolab:fix/lone-high-surrogate

Conversation

@ddxd

@ddxd ddxd commented Sep 22, 2026

Copy link
Copy Markdown

Fixes #481.

get_unicode_codepoint returned Ok((0, src_offset)) for a high surrogate not followed by a \u escape, and for a surrogate pair with invalid hex digits in either half. The 0 was the old "bytes written = 0 means failure" sentinel from handle_unicode_codepoint. Since 29d9998 the function returns a code point, so it now means U+0000 and "\ud800" parsed as "\0".

This change returns Err(ErrorType::InvalidUnicodeCodepoint) in both places. Every parse_str implementation already maps an error from here to InvalidUnicodeCodepoint at the escape's position, so no caller changes are needed.

Test: lone_high_surrogate_is_an_error covers "\ud800", "\ud800x", ["\udbff"], a surrogate inside a string, "\ud800\n", "\ud800\uzzzz", "\ud800A" and "\udc00" through both to_tape and to_borrowed_value, and checks that a valid pair still decodes. It fails without the fix (accepted: "\ud800") and passes with it.

cargo test passes on aarch64 (macOS, NEON) and x86_64 (Linux, AVX2), and cargo fmt --check is clean.

Side note, not changed here: the if o == 0 checks after handle_unicode_codepoint in the avx2/sse42/neon/simd128/portable parse_str can no longer trigger, because codepoint_to_utf8 writes at least one byte. I left them in to keep this diff minimal. Happy to remove them in a follow-up if you prefer.

🤖 Generated with Claude Code

get_unicode_codepoint returned Ok((0, _)) for a high surrogate that is
not followed by a `\u` escape, and for a surrogate pair with invalid hex
digits. The 0 is a leftover sentinel: when this code returned the number
of bytes written, 0 meant failure and callers checked `o == 0`. Since
29d9998 it returns the code point instead, so the 0 became U+0000 and
"\ud800" parsed as "\0" without an error.

Return InvalidUnicodeCodepoint in both places, and add a regression test
covering the lone-surrogate forms (fails before this change).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Lone high surrogate "\ud800" is decoded as U+0000 instead of an error

1 participant