Skip to content

Keep commits fast while compaction runs - #76

Closed
venkat1701 wants to merge 7 commits into
perf/compaction-allocationfrom
perf/compaction-latency
Closed

venkat1701 wants to merge 7 commits into
perf/compaction-allocationfrom
perf/compaction-latency

Conversation

@venkat1701

Copy link
Copy Markdown
Collaborator

Part of #50. Stacked on #75 (which is stacked on #74 and #73). Merge those first.

After #75, a commit could still wait 100 to 150 ms while a full compaction ran. I traced the long waits to two causes. Neither was the copying itself.

1. Compaction kept a lot of heap alive, so young GCs were slow. With a cold cache, peak live heap during one full compaction was 1.4 GiB:

  • the liveness walk put every page it decoded into the node cache;
  • relocation held every copied page image until the end of the walk;
  • the rewritten trees were built in memory before any of them was written.

Fixes:

  • compaction reads use pages that are already cached but don't add new ones;
  • copied pages are written as soon as they're made;
  • rebuilt nodes are written as soon as their children are.

2. Checkpoints fsynced while holding the commit lock. compact() runs a checkpoint at the end, and the checkpoint fsynced every segment file, the feed, the catalog and the WAL under the commit lock. Commits stalled for as long as the disk took, about 100 ms after a compaction. This affects periodic checkpoints too. Now:

  • the checkpoint captures its position under the lock: generation, history, segment list, WAL end, feed size;
  • it syncs and writes the catalog afterwards, while commits continue.

Everything up to the captured generation was written before the capture, so the syncs cover it, and later commits replay from the WAL as before. Relocation also logs its page records before taking the lock now, instead of during the swap.

Then allocation. With the cache no longer absorbing decoded pages, a full liveness walk allocated about 3 GiB, mostly decoding every inline posting list in the index leaves just to learn they point nowhere.

  • Leaves now carry a mark in the unused flags byte of the page header when none of their values references a nested tree.
  • The walk reads only the header of such a leaf.
  • Pages written before this change have no mark and are decoded as before. Older binaries don't look at the byte, so files stay compatible both ways.
  • Only the liveness walk uses the mark, since it only steers which segments to compact. Relocation still reads whole, checksummed pages.
  • Inline posting lists are also decoded without a stream and a second list copy, which helps every index read.

Results

Peak live heap during one full compaction, cold cache, class histograms (deterministic):

peak live heap
#75 1,392 MiB
+ cache bypass 231 MiB
+ pages and nodes written as they are made 28.5 MiB (3.5 MiB over idle)

Slowest commits while a full compaction runs: one thread committing single updates back to back, 3 runs each, same machine:

slowest commit GC pauses during compaction
#75 152, 112, 115 ms 388 to 555 ms
+ heap, swap and checkpoint changes 30, 42, 7.5 ms 31 to 54 ms
+ leaf mark and postings decode 8.7, 7.7, 9.4 ms 11 to 14 ms

Compaction allocation on the same probe: 3.7 GiB → 1.2 GiB (liveness walk 3.0 GiB → 0.44 GiB).

Benchmark (compaction workload, 50 commits/s, open loop, 3 alternating runs per config, all checksums identical): the machine's load average was 8 to 17 for the whole run, so these are noisy.

  • Median commit p99 during compaction: 101 → 98 ms (default) and 151 → 135 ms (16 MiB segments), with wide ranges on both sides.
  • The quietest runs show the effect: 12 ms in one A/B run, and 10 ms p99 / 38 ms max in a separate profiled run at load 4.8. HyperGraphDB measured 12 to 15 ms in this window.
  • In that profiled run the one GC during compaction took 35 ms. It copied data that earlier workloads had left in the young generation; compaction itself kept almost nothing alive.
  • mixed.write p99 107 → 49 ms (default), 56 → 62 ms (segments).

Tested:

  • full ./mvnw install passes (114 tests); each commit compiles on its own
  • livenessWalkLeavesTheNodeCacheAsItWas: a second walk after a cleared cache gets no hits (fails with the old behaviour: 2,549 hits)
  • checkpointsTakenWhileCommitsRunLoseNothing, both WAL modes: 2,000 commits while checkpoints run in a loop, then a crash; all of them survive. Taking the WAL position after the sync instead of at the capture makes it fail with 1,999 of 2,000.
  • NodeCodecTest: the mark is set exactly when a leaf has no nested tree; branches never get it
  • livenessCountsEveryPageOfEveryRetainedRootOnce now runs on a cold cache against a walk that decodes every page. Marking every leaf as plain makes it fail.
  • the existing crash tests for compaction (a crash at every crash point during the swap) pass, including the new window where the page records are logged and the commit is not

Not done here: relocation and pairing still read whole pages into fresh buffers (about 530 MiB per pass in the benchmark). Reading the exact page size in one buffer would roughly halve that.

The liveness walk and relocation decode every page they reach. On a cold
cache all of those nodes were cached, survived the next young collection
and were copied into the old generation, which showed up as one long GC
pause while commits waited. Compaction now uses pages that are already
cached but no longer adds the ones it decodes.
The swap appended one WAL record per copied page while holding the commit
lock, so commits waited for tens of thousands of records to be encoded
and written. Recovery matches records to their commit by transaction id,
so the records are now logged right after the copies are written and the
swap only appends the new roots. A relocation that swaps nothing logs an
abort.
Relocation kept every copied page image in a list until the whole walk
had finished, so a full pass held about as much heap as the live data it
moved. The copies sit at fresh locations that no committed tree points
to until the swap, so they are written straight away. Pages rebuilt for
relocation are no longer added to the node cache either.
Relocation built the whole rewritten tree for a slot in memory before
writing any of it, which for the atoms tree means most of its leaves.
Each rebuilt node is now written as soon as its children are, so only
the current path stays on the heap.
A checkpoint fsynced every segment file, the feed, the catalog and the
WAL while holding the commit lock, so commits stalled for as long as the
disk took to flush. Right after a compaction that is hundreds of MiB.

The checkpoint now only captures its position under the lock: the
current generation, history, segment list, WAL end and feed size. The
syncs and the catalog write happen afterwards while commits continue.
Everything up to the captured generation was written before the capture,
so the syncs cover it, and later commits replay from the WAL as before.
The liveness walk decoded every leaf of trees whose values can point to
nested trees, such as the posting lists of the indexes, only to find
that most of them point nowhere. On a cold cache that decoded every
inline posting list and allocated about 3 GiB per full pass.

Leaves are now marked in the unused flags byte of the page header when
none of their values references a nested tree. The walk reads just the
header of an uncached leaf and stops there when the mark is set. Pages
written before this change carry no mark and are decoded as before.
Relocation still decodes, since it needs checksummed pages to be safe.
Each inline posting list went through a stream, a list that allows
nulls and then a second copy in Postings.Inline. It is now built once
with List.of, which Inline keeps as it is.
@venkat1701

Copy link
Copy Markdown
Collaborator Author

Closing this. #81 replaced tree-rewriting compaction with moving pages through the page directory, so most of this stack no longer applies. A review also found two data-loss bugs in #73 (a retirement generation older than a commit that still used the moved pages, and bulk-loaded pages left behind in a victim). The pieces that still help, the posting list decode, records reused by mapRefs, and the checkpoint crash test, are in #82.

@venkat1701 venkat1701 closed this Oct 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant