Repository navigation
Benchmark deletes and disk space after them - #43
Merged
Merged
Conversation
20% of the hyperedges to delete and 20,000 others that each lose one member. It is drawn after everything else, so the existing dataset and probes don't change.
HyperGraphDB links are immutable, so removing a member replaces the link with one that lacks it. reopen also takes a callback that runs while the store is closed.
The run records disk while closed before the churn, deletes and removes, reads incidence sets again and records disk.churned at the end.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #34
Adds a churn phase after the existing sequence:
diskis now recorded while the store is closed for a reopen, so it still means the loaded store after a clean shutdown.churn.delete: delete 20% of the hyperedges, 1,000 per transaction.churn.remove: remove one member from each of 20,000 other hyperedges with at least three members.read.incidence.churned: incidence sets again. The checksum has to agree, which shows both engines applied the same deletes.disk.churned: bytes on disk after a clean shutdown.The churn set is drawn from the same seed after everything else, so the existing dataset and probes don't change. HyperGraphDB links are immutable, so member removal replaces the link with a new
HGPlainLinkthat lacks the member.Smoke run at scale 1, async, one run each. Checksums agree after churn (645,087 for both).
churn.deletechurn.removediskbefore churndisk.churnedSo HStore deletes quickly but doesn't give the space back. Its files more than double after deleting a fifth of the edges, while JE's cleaner shrinks them. My guess is the 64 retained history generations, plus compaction only running after 256 MiB of writes and moving one segment per pass, but I haven't confirmed it. I'll open a separate issue.
Note that
diskbefore churn is also higher than in the published tables, because the mixed workload from #41 writes about 70 MB first. The published numbers need a full rerun anyway.