Skip to content

Copy relocated leaf pages as bytes during compaction - #74

Closed
venkat1701 wants to merge 3 commits into
perf/background-compactionfrom
perf/compaction-byte-copy
Closed

venkat1701 wants to merge 3 commits into
perf/background-compactionfrom
perf/compaction-byte-copy

Conversation

@venkat1701

Copy link
Copy Markdown
Collaborator

Part of #50. Stacked on #73: merge that first; this PR's base then moves to main once the #73 branch is deleted.

Compaction decoded every moved leaf into objects, copied them, and encoded them again. That was most of the roughly 2 GiB a pass allocated, and the GC pauses that came with it were what was left of the commit tail after #73.

Now a moved page from a tree without nested references is copied as raw bytes:

  • read it
  • verify it against its old page id, so a corrupt page fails instead of being re-sealed with a fresh, valid checksum
  • check the height in its header; only leaves are copied
  • re-stamp the header with the new page id and checksum, log it and write it, the same way the Materializer does

Branches and leaves that hold nested trees are still rebuilt. This also covers nested single-leaf trees (most small edges' membership trees), which are reached through a reference whose height the walker doesn't know; the page header tells it. Old and new pages are now paired after the copies are written, since copied pages never go through the node cache.

Results, HStore only, 3 alternating runs each, #73 against #73 plus this PR, load 9 to 21, all checksums identical:

during a full compaction default segments 16 MiB segments
allocated 1,783 → 1,244 MiB (-30%) 2,141 → 1,459 MiB (-32%)
GC pauses 1,326 → 340 ms (-74%) 720 → 270 ms (-63%)
commit p99 692 → 267 ms (-61%) 348 → 195 ms (-44%)
commit max 771 → 343 ms 418 → 271 ms
compaction time 3.8 → 2.3 s 4.1 → 2.3 s

Space reclaimed and disk size after compaction are unchanged.

I also tried capping background passes at 64 MiB of live data. With 16 MiB segments that made passes bigger (16 to 17 segments instead of 1), and the mixed.write tail got worse. Real bounding needs moving part of a segment per pass, so I dropped that change.

Tested:

  • full ./mvnw install passes (103 tests); each commit compiles on its own
  • all 11 CompactionTest cases pass, including the snapshot-to-swap merge, two branches with reclaim and reopen, and the crash matrix of the swap
  • new aCorruptPageIsNotCopiedWithAFreshChecksum: a membership leaf corrupted on disk (while its decoded copy is still cached) makes relocation fail. I checked that it fails if verification is skipped.

Compaction decoded every moved leaf into objects, copied them and encoded
them again. A leaf of a tree without nested references can be copied as
it is: verify it against its old id, re-stamp the header with the new
page id and checksum, and write it. Only branches and leaves holding
nested trees are still rebuilt. Moves are paired after the copies are
written, since copied pages never go through the node cache.
A corrupted membership leaf makes relocation fail instead of being
re-sealed with a valid checksum; the test fails if verification is
skipped.
copyPage wrote the new page id straight into the image that PageStore.read
handed back. That only works while read returns a private buffer. Once pages
are read through the segment map they come back read-only, and every
compaction failed with "Attempt to write a read-only segment". Copy the
page into its own array first.
@venkat1701

Copy link
Copy Markdown
Collaborator Author

Closing this. #81 replaced tree-rewriting compaction with moving pages through the page directory, so most of this stack no longer applies. A review also found two data-loss bugs in #73 (a retirement generation older than a commit that still used the moved pages, and bulk-loaded pages left behind in a victim). The pieces that still help, the posting list decode, records reused by mapRefs, and the checkpoint crash test, are in #82.

@venkat1701 venkat1701 closed this Oct 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant