Skip to content

Copy compaction pages outside the commit lock - #73

Closed
venkat1701 wants to merge 4 commits into
mainfrom
perf/background-compaction
Closed

venkat1701 wants to merge 4 commits into
mainfrom
perf/background-compaction

Conversation

@venkat1701

Copy link
Copy Markdown
Collaborator

Closes #50

The problem: compaction rebuilt, encoded and wrote every relocated page inside one rewrite that held the commit lock, so commits waited for the whole pass (p99 of 2.4 s in the benchmark).

The design (as agreed, "copy in background, swap in"): TransactionManager.relocate works in three steps.

  1. Snapshot under the commit lock for a moment: the latest generation, plus a reserved transaction id.

  2. Copy without the lock:

    • rebuild the paths to every victim page from the snapshot (TreeWalker.relocate, now also recording every page it reaches)
    • write the copies to the active segment right away
    • map each old page to its copy by walking old and new side by side (pairMoves)
  3. Swap under the lock, one branch at a time. TreeWalker.adopt walks the branch's latest roots:

    • a page with a copy is replaced by the copy
    • a page that existed at the snapshot without a copy is kept, since nothing below it moved
    • a page written after the snapshot (by a commit during the copy) is descended into and rebuilt if needed

    Only that last kind costs work under the lock. The swap is a normal commit: it logs the copies under the reserved transaction id and moves relocationFence. Branches created during the copy are swapped too.

rewrite had no other callers and is removed. No file-format change, and nothing changes on the read path. Victim choice, segment states, retiring and reclaiming are as before.

Known costs:

  • The copy step keeps a set of every page id reachable from the snapshot: tiny here, tens of MB on a 100 GB database.
  • With wal_mode = images the copies' page images are appended to the WAL during the swap, so that mode still does that I/O under the lock. Default references mode only logs a small record per page.

Results, HStore only, 3 alternating runs each, on the PR stack (#65/#66/#69) with and without these commits. Machine load was about 6. All checksums agree.

commit latency during a full compaction before after
p99, default segments 882 ms 330 ms
p99, 16 MiB segments 1,175 ms 133 ms
max, 16 MiB segments 1,229 ms 173 ms
p99 without compaction (baseline) 11 to 12 ms 10.5 to 11 ms

Space reclaimed and disk size after compaction are identical (227 / 543 MiB reclaimed). Compaction itself got slightly faster (1.7 to 1.9 s down to 1.4 to 1.7 s).

What's left: the remaining 130 to 330 ms tail is GC, not the lock. A pass allocates 1.7 to 2.1 GiB, causing 2 or 3 pauses of roughly 100 to 350 ms, and each run's worst commit matches its longest pause. Copying pages as raw bytes instead of decoding and re-encoding nodes would cut that. It fits with the allocation work in #56 and #57.

Tested:

  • full ./mvnw install passes (102 tests); each commit compiles on its own
  • new CompactionTest (10 cases), 5 runs in a row:
    • a commit goes through while compaction is copying. The victim check blocks until a commit completes, which would deadlock under the old design.
    • commits that land between the snapshot and the swap end up pointing at the copies. I checked this test by deliberately breaking adopt: it then fails with "2175 pages in the current trees still point into compacted segments".
    • commits on main and a branch during a full compaction are all kept; the old segments are then deleted, and everything reads back, including after reopen
    • a crash at each point of the swap commit (PAGE, WAL_APPEND, DATA_WRITE, COMMIT_APPEND, DATA_SYNC, WAL_SYNC, CATALOG_PUBLISH) loses nothing
  • native Docker image: COMPACT runs through the new path, data unchanged, hstore check passes

…r trees

TreeWalker.relocate can report every page it reaches, pairMoves maps each
old page to its relocated copy, and adopt rebuilds only the nodes written
after the snapshot, swapping in copies for everything that moved.
Compaction used to rebuild and write every relocated page inside one
rewrite that held the commit lock, so commits stopped for the whole pass.
It now relocates and writes the copies from a snapshot without the lock,
then takes the lock only to point each branch's latest roots at the
copies, rebuilding the few nodes commits wrote in the meantime. The swap
commit logs the copies and still advances the relocation fence.
Checks that a commit goes through while compaction copies pages, that
commits landing between the snapshot and the swap end up pointing at the
copies, that two branches keep their changes and the old segments get
deleted, and that a crash at any point of the swap loses nothing.
@venkat1701

Copy link
Copy Markdown
Collaborator Author

Closing this. #81 replaced tree-rewriting compaction with moving pages through the page directory, so most of this stack no longer applies. A review also found two data-loss bugs in #73 (a retirement generation older than a commit that still used the moved pages, and bulk-loaded pages left behind in a victim). The pieces that still help, the posting list decode, records reused by mapRefs, and the checkpoint crash test, are in #82.

@venkat1701 venkat1701 closed this Oct 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Compaction blocks commits until the whole pass finishes

1 participant