Skip to content

Read pages through a memory map of their segment - #77

Merged
venkat1701 merged 5 commits into
mainfrom
perf/mapped-reads
Oct 11, 2026
Merged

venkat1701 merged 5 commits into
mainfrom
perf/mapped-reads

Conversation

@venkat1701

Copy link
Copy Markdown
Collaborator

First step toward keeping the raw page cache off the Java heap. The server binary is built with --gc=serial, and right now every page that misses NodeCache is read with pread into one or two new heap buffers. This PR maps segment files read-only and reads pages out of the map instead, so the raw bytes sit in the OS page cache. A follow-up on top of #69 will shrink the decoded cache and rename it node_cache_mb.

What changed

  • SegmentFile keeps a read-only map of the file. If the map covers the image, read copies exactly 80 + payloadLength bytes from it into a heap array. No system call.
  • A sealed segment is remapped as soon as a read needs bytes past the map. The active segment is remapped only after it has grown 4 MiB past it; until then the newest pages are read with pread. PageStore seals a segment when it rolls, and seals everything but the active segment on open.
  • Maps are made in Arena.ofAuto(). Closing an Arena.ofShared() in a native image needs -H:+SharedArenaSupport, which is experimental. An automatic arena also keeps the map valid for a thread that is mid-copy when compaction deletes the segment. Segment files are never truncated, so a map can't point past the end of its file.
  • Both paths now return exactly one image, read-only. The pread path used to return the whole 4 KiB first buffer when the image was smaller.

Why it copies instead of decoding from the map

ByteCursor reads byte by byte through MemorySegment. In a microbenchmark, decoding varlongs ran at 4.9 ns each from a map and 6.6 ns from a heap array when each was the only kind the JVM saw. Mixed in one JVM, both went to 15-35 ns. In the engine the WAL, catalog and feed always decode from heap arrays, so decoding pages straight from the map would slow down every decoder. Copying the page first kept decoding at 4.8 ns.

Numbers

Measured on a loaded laptop (load average 20-33), so only the direct measurements are worth quoting.

pread (main) mapped
PageStore.read, every page of a 24,000-page store once after reopen 7.8-11.3 us 2.6-3.7 us
same, second pass 6.7-10.4 us 1.5-3.2 us
heap allocated per read.members op with a 32 MB node cache 16.0 KB 12.5 KB

The full comparison benchmark swung 0.3x to 3x between identical runs at that load, so I'm not quoting its throughput. Cold incidence reads in it came out lower in some pairs, but the store-level measurement above doesn't show that, and I couldn't separate it from the noise. It needs a rerun on a quiet machine.

Testing

  • New PageStoreTest: 16,000 pages of random sizes written and read back while segments fill, roll and reopen; reads are read-only; finished segments get mapped, and the growing one only after 4 MiB. Breaking the mapping, the seal on roll or the growth remap each fails a test.
  • Full suite: 115 tests pass.
  • Built the native server from this branch, ran hstore bench, then hstore check on all nine databases it produced: about 130,000 pages verified through mapped reads. check on the largest took 2.3-2.9 s against 2.5-3.4 s with the 0.1.0 binary.

Interaction with the open stack

#74 copies relocated leaves by writing a new page id into the buffer that read returned. With reads now read-only, that call has to copy the page first. Whichever of the two merges second needs that small change, and I'll do it when rebasing.

A positional read used to hand back the whole 4 KiB first buffer when the
image was smaller, so callers got the start of the next image too. They
only looked at the header's length, but it made every read a different
size. Now the read returns exactly 80 + payloadLength bytes, and the
memory is read-only so nothing can write into a stored image through it.
Every page read used to be a pread into one or two fresh heap buffers.
Segment files are now mapped read-only and a read copies the image out
of the map, so there is no system call and the raw bytes sit in the OS
page cache instead of the Java heap.

A sealed segment is remapped as soon as a read needs bytes past the map.
The active one is remapped only after it has grown 4 MiB past it, and
reads near its end use pread until then. Maps live in an automatic
arena: closing a shared arena in a native image needs an experimental
option, and an automatic one also stays valid for a thread that is
mid-copy when compaction deletes the segment.

The copy is on purpose. Decoding straight from the map while the WAL,
catalog and feed decode from heap arrays made ByteCursor two to five
times slower, because the JIT stops specialising once it sees both kinds
of segment.

Reading every page of a 24,000-page store once after a reopen took
2.6-3.7 us a page against 7.8-11.3 us with pread, and 1.5-3.2 against
6.7-10.4 us the second time.
Covers when a segment is mapped and remapped, why reads copy out of the
map instead of decoding from it, and what the pread path still does for
the end of the active segment.
Maps lived in automatic arenas, so they were only unmapped when the garbage
collector got to them. A segment deleted by compaction kept its disk space
until then, which on a quiet heap could be hours, and every remap of the
active segment left the old map in place too. Each map now has its own shared
arena, closed when the map is replaced or the segment is closed or deleted. A
read that races a close sees the map is gone and reads from the current one.
Closing shared arenas in Native Image needs -H:+SharedArenaSupport, which the
native profile now turns on.
@venkat1701
venkat1701 merged commit d06aea5 into main Oct 11, 2026
@venkat1701
venkat1701 deleted the perf/mapped-reads branch October 11, 2026 09:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant