Repository navigation
Benchmark two-hop reads - #44
Merged
Merged
Conversation
Counts the distinct nodes reachable through a node's hyperedges with each engine's plain incidence and member calls. The 10,000 probes are drawn after the existing ones, so nothing else in the dataset changes.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #35
Adds
read.twohop: for each of 10,000 probe nodes, count the distinct nodes that share a hyperedge with it.Reader.incidentthenmemberIds, with ids sorted and de-duplicated as longsgetIncidenceSetthen the link's targets, collected into aHashSetof persistent handles (it has no numeric ids)It runs right after
read.comembership, beforelarge.ingest. Once the hyperedge with every node in it exists, every node is two hops from all the others and the workload stops meaning anything.I didn't use
HGBreadthFirstTraversal, since that would measure HyperGraphDB's traversal framework rather than its storage.The probes are drawn after the existing random draws, so nothing else in the dataset changes.
Smoke run at scale 1, async, one run each: both return 281,832. HStore did 25,916 probes/s, HyperGraphDB 13,489 (1.9x), with 10.1 KiB and 13.4 KiB allocated per probe.