Skip to content

Report write and read amplification - #66

Merged
venkat1701 merged 8 commits into
mainfrom
feat/bench-amplification
Oct 11, 2026
Merged

venkat1701 merged 8 commits into
mainfrom
feat/bench-amplification

Conversation

@venkat1701

@venkat1701 venkat1701 commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #61

Stacked on #65, since both change the same harness files. Merge #65 first; this PR's base then moves to main once the #65 branch is deleted on merge.

The Resources table gets two new ratios per workload, plus bytes read:

Ratio Physical bytes Logical bytes (fixed estimate per workload)
write amplification bytes written (from #38) 8 bytes per id plus the value: node-<i> for a new node, 24 bytes for an updated value, (cardinality + 1) × 8 for a new or deleted edge, 16 bytes per removed member
read amplification bytes read from storage 8 bytes per id returned (the checksum counts them)

Correction (pushed after the PR was first opened): I first computed HStore's bytes read as pages read × page size, assuming each read takes a whole page. It doesn't: SegmentFile.read reads the first 4 KiB, then only the rest of the image if it's bigger. That overstated HStore's reads by up to 4 times. So this PR now also:

  • adds PageStore.bytesRead and EngineStats.dataBytesRead, which count the bytes actually read
  • has the adapter use that

JE reports sequential plus random read bytes. Both adapters carry the counts across reopen.

Report measures can now be computed from several fields, which the ratios need.

Smoke runs at scale 1, async, one run each, with the corrected count:

  • With default caches HStore reads nothing from storage on warm reads, since its node cache holds everything.

  • With CACHE_MB=64:

    HStore HyperGraphDB
    read.incidence per op 4 B 1.5 KiB
    read.incidence.cold per op 38 B 10.8 KiB
    read.members per op 4.4 KiB 84 B
    read.twohop per probe 34.9 KiB 1.8 KiB

    Member scans are the outlier: about 157 times the bytes they return, because every small edge has its own membership page. That's Small edges always get their own membership tree #55.

  • Write amplification matches what bytes per operation already showed: HStore is 1.3 to 2× better on batched writes and large edges, and 6.6× worse on single commits (662 against 100), which is Small commits are slow and occasionally stall for seconds #46.

Tested:

  • each commit compiles on its own; full ./mvnw install passes
  • new EngineTest.bytesReadCountWhatComesOffTheSegmentFiles: after reopen, bytes read are at least a header and at most a page per page read
  • old result files render unchanged
  • results without the new fields render without the new rows

HStore reads whole pages, so bytes read are pages read times the page
size. JE reports sequential and random read bytes. Both carry the count
across reopen, like bytes written.
Write workloads get a fixed estimate of the bytes they change, read
workloads 8 bytes per id they return, and read workloads now also record
the bytes they read from storage.
Report measures can now be computed from several fields, which the two
amplification ratios need.
PageStore already counted pages read, but a read covers a first 4 KiB and
then only the rest of the image if it's bigger, so pages times page size
overstated it. EngineStats now reports the bytes actually read.
The adapter multiplied pages read by the page size, which overstated
HStore's reads and read amplification by up to four times.
@venkat1701
venkat1701 deleted the branch main October 11, 2026 09:10
@venkat1701 venkat1701 closed this Oct 11, 2026
@venkat1701 venkat1701 reopened this Oct 11, 2026
@venkat1701
venkat1701 changed the base branch from feat/bench-compaction-ratio to main October 11, 2026 09:12
@venkat1701
venkat1701 merged commit b4c6036 into main Oct 11, 2026
@venkat1701
venkat1701 deleted the feat/bench-amplification branch October 11, 2026 09:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Report write and read amplification in the benchmark

1 participant