Repository navigation
Keep commits fast while compaction runs - #76
Closed
venkat1701 wants to merge 7 commits into
Closed
venkat1701 wants to merge 7 commits into
venkat1701 wants to merge 7 commits into
Conversation
The liveness walk and relocation decode every page they reach. On a cold cache all of those nodes were cached, survived the next young collection and were copied into the old generation, which showed up as one long GC pause while commits waited. Compaction now uses pages that are already cached but no longer adds the ones it decodes.
The swap appended one WAL record per copied page while holding the commit lock, so commits waited for tens of thousands of records to be encoded and written. Recovery matches records to their commit by transaction id, so the records are now logged right after the copies are written and the swap only appends the new roots. A relocation that swaps nothing logs an abort.
Relocation kept every copied page image in a list until the whole walk had finished, so a full pass held about as much heap as the live data it moved. The copies sit at fresh locations that no committed tree points to until the swap, so they are written straight away. Pages rebuilt for relocation are no longer added to the node cache either.
Relocation built the whole rewritten tree for a slot in memory before writing any of it, which for the atoms tree means most of its leaves. Each rebuilt node is now written as soon as its children are, so only the current path stays on the heap.
A checkpoint fsynced every segment file, the feed, the catalog and the WAL while holding the commit lock, so commits stalled for as long as the disk took to flush. Right after a compaction that is hundreds of MiB. The checkpoint now only captures its position under the lock: the current generation, history, segment list, WAL end and feed size. The syncs and the catalog write happen afterwards while commits continue. Everything up to the captured generation was written before the capture, so the syncs cover it, and later commits replay from the WAL as before.
The liveness walk decoded every leaf of trees whose values can point to nested trees, such as the posting lists of the indexes, only to find that most of them point nowhere. On a cold cache that decoded every inline posting list and allocated about 3 GiB per full pass. Leaves are now marked in the unused flags byte of the page header when none of their values references a nested tree. The walk reads just the header of an uncached leaf and stops there when the mark is set. Pages written before this change carry no mark and are decoded as before. Relocation still decodes, since it needs checksummed pages to be safe.
Each inline posting list went through a stream, a list that allows nulls and then a second copy in Postings.Inline. It is now built once with List.of, which Inline keeps as it is.
venkat1701
force-pushed
the
perf/compaction-allocation
branch
from
October 8, 2026 20:56
14fa0c0 to
5d6371d
Compare
venkat1701
force-pushed
the
perf/compaction-latency
branch
from
October 8, 2026 20:56
484e77b to
53cbb21
Compare
Collaborator
Author
|
Closing this. #81 replaced tree-rewriting compaction with moving pages through the page directory, so most of this stack no longer applies. A review also found two data-loss bugs in #73 (a retirement generation older than a commit that still used the moved pages, and bulk-loaded pages left behind in a victim). The pieces that still help, the posting list decode, records reused by mapRefs, and the checkpoint crash test, are in #82. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #50. Stacked on #75 (which is stacked on #74 and #73). Merge those first.
After #75, a commit could still wait 100 to 150 ms while a full compaction ran. I traced the long waits to two causes. Neither was the copying itself.
1. Compaction kept a lot of heap alive, so young GCs were slow. With a cold cache, peak live heap during one full compaction was 1.4 GiB:
Fixes:
2. Checkpoints fsynced while holding the commit lock.
compact()runs a checkpoint at the end, and the checkpoint fsynced every segment file, the feed, the catalog and the WAL under the commit lock. Commits stalled for as long as the disk took, about 100 ms after a compaction. This affects periodic checkpoints too. Now:Everything up to the captured generation was written before the capture, so the syncs cover it, and later commits replay from the WAL as before. Relocation also logs its page records before taking the lock now, instead of during the swap.
Then allocation. With the cache no longer absorbing decoded pages, a full liveness walk allocated about 3 GiB, mostly decoding every inline posting list in the index leaves just to learn they point nowhere.
flagsbyte of the page header when none of their values references a nested tree.Results
Peak live heap during one full compaction, cold cache, class histograms (deterministic):
Slowest commits while a full compaction runs: one thread committing single updates back to back, 3 runs each, same machine:
Compaction allocation on the same probe: 3.7 GiB → 1.2 GiB (liveness walk 3.0 GiB → 0.44 GiB).
Benchmark (
compactionworkload, 50 commits/s, open loop, 3 alternating runs per config, all checksums identical): the machine's load average was 8 to 17 for the whole run, so these are noisy.mixed.writep99 107 → 49 ms (default), 56 → 62 ms (segments).Tested:
./mvnw installpasses (114 tests); each commit compiles on its ownlivenessWalkLeavesTheNodeCacheAsItWas: a second walk after a cleared cache gets no hits (fails with the old behaviour: 2,549 hits)checkpointsTakenWhileCommitsRunLoseNothing, both WAL modes: 2,000 commits while checkpoints run in a loop, then a crash; all of them survive. Taking the WAL position after the sync instead of at the capture makes it fail with 1,999 of 2,000.NodeCodecTest: the mark is set exactly when a leaf has no nested tree; branches never get itlivenessCountsEveryPageOfEveryRetainedRootOncenow runs on a cold cache against a walk that decodes every page. Marking every leaf as plain makes it fail.Not done here: relocation and pairing still read whole pages into fresh buffers (about 530 MiB per pass in the benchmark). Reading the exact page size in one buffer would roughly halve that.