Skip to content

[patch] Compare natural strings in place, allocating nothing per comparison - #64

Merged
matt-edmondson merged 1 commit into
mainfrom
fix/allocation-free-compare-56
Oct 5, 2026
Merged

matt-edmondson merged 1 commit into
mainfrom
fix/allocation-free-compare-56

Conversation

@matt-edmondson

Copy link
Copy Markdown
Contributor

Fixes #56

Summary

#57 already removed the per-call Regex, but Compare still built both strings' chunks before comparing them: a List<Chunk> and a StringBuilder per string, plus a string per chunk. That came to about 616 bytes per comparison. This PR implements the issue's preferred fix, a character scanner that walks both strings chunk by chunk and allocates nothing.

  • Digits are still read by code point with CharUnicodeInfo and compared by value. Non-ASCII digits, non-BMP digits, leading zeros and long digit runs behave exactly as before.
  • Text chunks are compared with string.CompareOrdinal(x, xStart, y, yStart, length) plus a length tie-break. That matches the previous string.Compare(..., Ordinal) on the extracted chunks.
  • When a number meets text, the number's first digit (as its ASCII equivalent) is compared with the text's first character, as Keep NaturalStringComparer transitive across digit scripts #57 specified.
  • The Chunk record and using System.Text are gone. The public API is unchanged.

Verification

  • New test Compare_AllocatesNothing: it runs 14,000 comparisons over ASCII, non-ASCII, surrogate-pair, leading-zero and number-against-text inputs and asserts GC.GetAllocatedBytesForCurrentThread() didn't move. With the library change stashed it fails (8624000 bytes allocated over 14000 comparisons). With the change it passes.
  • Full suite: 20/20. The library builds with 0 warnings for every target framework.
  • Differential check (scratch, not committed): I compared the sign of Compare from the previous implementation and this one on 2,000,000 random pairs. The pairs mixed ASCII, Arabic-Indic, Devanagari and mathematical digits, letters, punctuation, emoji and lone surrogates. There were 0 mismatches.
  • Timing (Release, List.Sort of 50,000 IMG_<n>_<n>.jpg names): 532 ms before, 115 ms after.

🤖 Generated with Claude Code

https://claude.ai/code/session_015zULxdhphqscknnLMm1Bk1


Generated by Claude Code

…arison

Compare split both strings into a List<Chunk> through a StringBuilder before
comparing, so every comparison allocated two lists, a builder per string and
a string per chunk: about 616 bytes. A sort makes O(n log n) comparisons.

It now walks both strings a chunk at a time by index. Digits are read by
their value with CharUnicodeInfo, text chunks compare with
string.CompareOrdinal, and a number against text compares its first digit's
ASCII equivalent with the text's first character. The ordering is unchanged:
a differential run of 2,000,000 random pairs (ASCII, Arabic-Indic,
Devanagari and mathematical digits, leading zeros, lone surrogates, emoji)
against the previous implementation found no difference in sign. Sorting
50,000 file names went from 532 ms to 115 ms in Release.

Fixes #56

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015zULxdhphqscknnLMm1Bk1
@sonarqubecloud

sonarqubecloud Bot commented Oct 5, 2026

Copy link
Copy Markdown

@matt-edmondson
matt-edmondson merged commit ace1b1c into main Oct 5, 2026
14 checks passed
@matt-edmondson
matt-edmondson deleted the fix/allocation-free-compare-56 branch October 5, 2026 14:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

NaturalStringComparer builds a new Regex on every Compare call, so sorting large lists is slow and allocation-heavy

2 participants