Repository navigation
Read pages through a memory map of their segment - #77
Merged
Merged
Conversation
A positional read used to hand back the whole 4 KiB first buffer when the image was smaller, so callers got the start of the next image too. They only looked at the header's length, but it made every read a different size. Now the read returns exactly 80 + payloadLength bytes, and the memory is read-only so nothing can write into a stored image through it.
Every page read used to be a pread into one or two fresh heap buffers. Segment files are now mapped read-only and a read copies the image out of the map, so there is no system call and the raw bytes sit in the OS page cache instead of the Java heap. A sealed segment is remapped as soon as a read needs bytes past the map. The active one is remapped only after it has grown 4 MiB past it, and reads near its end use pread until then. Maps live in an automatic arena: closing a shared arena in a native image needs an experimental option, and an automatic one also stays valid for a thread that is mid-copy when compaction deletes the segment. The copy is on purpose. Decoding straight from the map while the WAL, catalog and feed decode from heap arrays made ByteCursor two to five times slower, because the JIT stops specialising once it sees both kinds of segment. Reading every page of a 24,000-page store once after a reopen took 2.6-3.7 us a page against 7.8-11.3 us with pread, and 1.5-3.2 against 6.7-10.4 us the second time.
Covers when a segment is mapped and remapped, why reads copy out of the map instead of decoding from it, and what the pread path still does for the end of the active segment.
This was referenced Oct 8, 2026
Maps lived in automatic arenas, so they were only unmapped when the garbage collector got to them. A segment deleted by compaction kept its disk space until then, which on a quiet heap could be hours, and every remap of the active segment left the old map in place too. Each map now has its own shared arena, closed when the map is replaced or the segment is closed or deleted. A read that races a close sees the map is gone and reads from the current one. Closing shared arenas in Native Image needs -H:+SharedArenaSupport, which the native profile now turns on.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
First step toward keeping the raw page cache off the Java heap. The server binary is built with
--gc=serial, and right now every page that missesNodeCacheis read withpreadinto one or two new heap buffers. This PR maps segment files read-only and reads pages out of the map instead, so the raw bytes sit in the OS page cache. A follow-up on top of #69 will shrink the decoded cache and rename itnode_cache_mb.What changed
SegmentFilekeeps a read-only map of the file. If the map covers the image,readcopies exactly80 + payloadLengthbytes from it into a heap array. No system call.pread.PageStoreseals a segment when it rolls, and seals everything but the active segment on open.Arena.ofAuto(). Closing anArena.ofShared()in a native image needs-H:+SharedArenaSupport, which is experimental. An automatic arena also keeps the map valid for a thread that is mid-copy when compaction deletes the segment. Segment files are never truncated, so a map can't point past the end of its file.preadpath used to return the whole 4 KiB first buffer when the image was smaller.Why it copies instead of decoding from the map
ByteCursorreads byte by byte throughMemorySegment. In a microbenchmark, decoding varlongs ran at 4.9 ns each from a map and 6.6 ns from a heap array when each was the only kind the JVM saw. Mixed in one JVM, both went to 15-35 ns. In the engine the WAL, catalog and feed always decode from heap arrays, so decoding pages straight from the map would slow down every decoder. Copying the page first kept decoding at 4.8 ns.Numbers
Measured on a loaded laptop (load average 20-33), so only the direct measurements are worth quoting.
pread(main)PageStore.read, every page of a 24,000-page store once after reopenread.membersop with a 32 MB node cacheThe full comparison benchmark swung 0.3x to 3x between identical runs at that load, so I'm not quoting its throughput. Cold incidence reads in it came out lower in some pairs, but the store-level measurement above doesn't show that, and I couldn't separate it from the noise. It needs a rerun on a quiet machine.
Testing
PageStoreTest: 16,000 pages of random sizes written and read back while segments fill, roll and reopen; reads are read-only; finished segments get mapped, and the growing one only after 4 MiB. Breaking the mapping, the seal on roll or the growth remap each fails a test.hstore bench, thenhstore checkon all nine databases it produced: about 130,000 pages verified through mapped reads.checkon the largest took 2.3-2.9 s against 2.5-3.4 s with the 0.1.0 binary.Interaction with the open stack
#74 copies relocated leaves by writing a new page id into the buffer that
readreturned. With reads now read-only, that call has to copy the page first. Whichever of the two merges second needs that small change, and I'll do it when rebasing.