Repository navigation
Report write and read amplification - #66
Merged
Merged
Conversation
HStore reads whole pages, so bytes read are pages read times the page size. JE reports sequential and random read bytes. Both carry the count across reopen, like bytes written.
Write workloads get a fixed estimate of the bytes they change, read workloads 8 bytes per id they return, and read workloads now also record the bytes they read from storage.
Report measures can now be computed from several fields, which the two amplification ratios need.
This was referenced Oct 6, 2026
PageStore already counted pages read, but a read covers a first 4 KiB and then only the rest of the image if it's bigger, so pages times page size overstated it. EngineStats now reports the bytes actually read.
The adapter multiplied pages read by the page size, which overstated HStore's reads and read amplification by up to four times.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #61
Stacked on #65, since both change the same harness files. Merge #65 first; this PR's base then moves to
mainonce the #65 branch is deleted on merge.The Resources table gets two new ratios per workload, plus bytes read:
node-<i>for a new node, 24 bytes for an updated value,(cardinality + 1) × 8for a new or deleted edge, 16 bytes per removed memberCorrection (pushed after the PR was first opened): I first computed HStore's bytes read as pages read × page size, assuming each read takes a whole page. It doesn't:
SegmentFile.readreads the first 4 KiB, then only the rest of the image if it's bigger. That overstated HStore's reads by up to 4 times. So this PR now also:PageStore.bytesReadandEngineStats.dataBytesRead, which count the bytes actually readJE reports sequential plus random read bytes. Both adapters carry the counts across reopen.
Report measures can now be computed from several fields, which the ratios need.
Smoke runs at scale 1, async, one run each, with the corrected count:
With default caches HStore reads nothing from storage on warm reads, since its node cache holds everything.
With
CACHE_MB=64:read.incidenceper opread.incidence.coldper opread.membersper opread.twohopper probeMember scans are the outlier: about 157 times the bytes they return, because every small edge has its own membership page. That's Small edges always get their own membership tree #55.
Write amplification matches what bytes per operation already showed: HStore is 1.3 to 2× better on batched writes and large edges, and 6.6× worse on single commits (662 against 100), which is Small commits are slow and occasionally stall for seconds #46.
Tested:
./mvnw installpassesEngineTest.bytesReadCountWhatComesOffTheSegmentFiles: after reopen, bytes read are at least a header and at most a page per page read