Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions docs/architecture/overview.md
Original file line number Diff line number Diff line change
Expand Up @@ -227,7 +227,7 @@ Readers take no locks at all.
The following is real output from `hstore init` followed by loading `examples/clinical-claims.hql` and running `CHECKPOINT;` (default 16 KiB pages):

```text
FORMAT 55 B format=2, page-size=16384 (java.util.Properties)
FORMAT 55 B format=3, page-size=16384 (java.util.Properties)
LOCK 0 B FileChannel.tryLock() held while open
hstore.conf 2397 B server configuration template written by `hstore init`
catalog/manifest 20 B name of the latest catalog file
Expand All @@ -242,7 +242,7 @@ semantic/state 53 B persisted HNSW index state (datab

| Path | Owner | Format |
|---|---|---|
| `FORMAT` | `StorageEngine.verifyFormat` | `Properties` with `format=2` (packed extents) and `page-size` (the maximum node size). Opening a format 1 directory, or opening with a different page size, fails. |
| `FORMAT` | `StorageEngine.verifyFormat` | `Properties` with `format=3` (packed extents, sized references) and `page-size` (the maximum node size). Opening a directory written in an older format, or opening with a different page size, fails. |
| `LOCK` | `StorageEngine.open` | An exclusive OS file lock. A second process gets `database ... is opened by another process`. |
| `catalog/generations/%016x.cat` | `CatalogStore.publish` | `int magic 0x54414348`, `int crc32c(payload)`, `long length`, then the `CatalogImage` payload. Written to `*.tmp`, fsynced, atomically renamed, and the directory fsynced. The three newest files are kept. |
| `catalog/manifest` | `CatalogStore` | The file name of the latest `.cat`, written durably the same way. |
Expand Down
2 changes: 1 addition & 1 deletion docs/operations/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,7 +94,7 @@ max_connections = 200
| `cache_nodes` | `HSTORE_CACHE_NODES` | `65536` | Decoded tree nodes kept in the node cache. |
| `checkpoint_wal_mb` | `HSTORE_CHECKPOINT_WAL_MB` | `256` | WAL volume since the last checkpoint that triggers a background checkpoint. |
| `history_limit` | `HSTORE_HISTORY_LIMIT` | `64` | Committed generations kept addressable for `AT GENERATION`, `AS OF`, `HISTORY` and the Studio time slider. |
| `compaction_live_ratio` | `HSTORE_COMPACTION_LIVE_RATIO` | `0.5` | Sealed segments whose live node images fall below this ratio are relocated by compaction: the background pass after every `checkpoint_wal_mb` of written node images, or `COMPACT` on demand. Time-travel history is preserved; retired segments are deleted once `history_limit` has rolled past them. Must satisfy 0 < r < 1; other values fail at startup. |
| `compaction_live_ratio` | `HSTORE_COMPACTION_LIVE_RATIO` | `0.5` | Sealed segments whose live bytes fall below this fraction of their allocated bytes are relocated by compaction: one segment (the emptiest) per background pass after every `checkpoint_wal_mb` of written node images, or all of them with `COMPACT`. Time-travel history is preserved; retired segments are deleted once `history_limit` has rolled past them. Must satisfy 0 < r < 1; other values fail at startup. |
| `feed_retention_generations` | `HSTORE_FEED_RETENTION_GENERATIONS` | `100000` | Committed generations kept in the [change feed](../storage/change-feed.md) for `HISTORY`, `DIFF` and subscribers. After every checkpoint, whole feed segments that lie entirely below `current − N` are deleted. *Holds* lower that floor: the semantic index's last persisted generation and the generation of each continuous materialised view. Feed segments roll at the WAL segment size (`EngineOptions.walSegmentBytes`, 64 MiB). Must be ≥ 1. |

### Resource budgets
Expand Down
2 changes: 1 addition & 1 deletion docs/operations/docker.md
Original file line number Diff line number Diff line change
Expand Up @@ -243,7 +243,7 @@ skipped because `FORMAT` exists. To verify a backup, run
`docker run --rm -v restored:/var/lib/hstore/data hstore check /var/lib/hstore/data`.

To upgrade, stop the container, back up the volume, and start the new image on the same volume. `FORMAT`
records the on-disk format version (`format=2`, packed segment extents) and the page size. Both are checked at
records the on-disk format version (`format=3`, sized page references) and the page size. Both are checked at
open. A directory written in another format is refused with
`database … uses storage format N; this build reads format 2 (packed segment extents); export and reload it`,
so the server never misreads it. Export the data with HQL against the old version and reload it into a fresh
Expand Down
2 changes: 1 addition & 1 deletion docs/operations/logging.md
Original file line number Diff line number Diff line change
Expand Up @@ -109,7 +109,7 @@ and in Docker it is the last line of `docker logs`. The storage-integrity failur

| Message | Meaning |
|---|---|
| `database <dir> uses storage format N; this build reads format 2 (packed segment extents); export and reload it` | `FORMAT` was written by an incompatible build. |
| `database <dir> uses storage format N; this build reads format 3 (sized page references); export and reload it` | `FORMAT` was written by an incompatible build. |
| `database uses N byte pages, options request M` | `page_size` differs from the value fixed at `init`. |
| `write-ahead log segment <file> is damaged (<reason>) but later segments exist; replay would silently drop committed transactions` | `CORRUPT_LOG`: damage in a non-tail WAL segment. A torn record at the end of the *last* segment is normal after a crash. Replay simply ends there. Damage before later segments means storage corruption, so recovery refuses to guess. Restore from backup. |

Expand Down
10 changes: 5 additions & 5 deletions docs/storage/maintenance.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ Sources are under `engine/src/main/java/io/hstore/engine/maintenance/`, plus `St
| `COMPACT;` (HQL), `StorageEngine.compact()` | compact, checkpoint, reclaim | `StorageEngine.compact` |
| `StorageEngine.close()` | final checkpoint, then shutdown | `close` |

Checkpoint and compaction are serialized by `StorageEngine.maintenance`, a `ReentrantLock`. Background compaction plays the role of PostgreSQL's autovacuum. Its trigger is write volume, `pages.bytesWritten()` advancing by more than `checkpoint_wal_mb` since the last pass, so an idle database never pays for a liveness walk. The pass only relocates segments whose live ratio is below `compaction_live_ratio`, and `COMPACT;` runs the same pass on demand.
Checkpoint and compaction are serialized by `StorageEngine.maintenance`, a `ReentrantLock`. Background compaction plays the role of PostgreSQL's autovacuum. Its trigger is write volume, `pages.bytesWritten()` advancing by more than `checkpoint_wal_mb` since the last pass, so an idle database never pays for a liveness walk. The background pass relocates at most one segment, the emptiest one whose live bytes are below `compaction_live_ratio` of its allocated bytes. `COMPACT;` relocates every such segment on demand.

Every checkpoint is logged, and then the checkpoint listeners run. The database registers `SemanticPlane.persist`, which saves the HNSW index alongside the catalog. A sample log line:

Expand Down Expand Up @@ -71,23 +71,23 @@ stateDiagram-v2

### Liveness

`Compactor.liveness()` returns, per segment, the number of distinct node images reachable from:
`Compactor.liveness()` returns, per segment, the bytes of distinct node images reachable from:

* **every retained generation** (`transactions.history()`) plus `current`;
* **every branch** in each of those generations;
* **every slot root** in each branch.

`TreeWalker.visit(root, schema, visited)` adds each `Ref.Stored` page id it meets to one shared `visited` set, so an image shared by several generations, branches or trees is counted once. It descends into branches and, for codecs that hold references (catalog edge records, promoted posting lists), into nested trees.
`TreeWalker.visit(root, schema, visited)` records each `Ref.Stored` page id it meets, with the image size the reference carries, in one shared `visited` map, so an image shared by several generations, branches or trees is counted once. The size comes from the reference itself, so no page has to be read to learn it. It descends into branches and, for codecs that hold references (catalog edge records, promoted posting lists), into nested trees.

A detail that keeps this cheap: at height 0 of a tree whose codec has no nested references, the walker records the leaf's id **without loading it**. The summaries in the parent already prove that the leaf exists. Liveness is therefore proportional to the number of internal nodes plus the reference-holding leaves, not to total data size.

Liveness is exposed as `StorageEngine.liveness()`. The Studio dashboard shows it per segment next to `SegmentInfo.pages` and `bytes()`.
Liveness is exposed as `StorageEngine.liveness()` (live bytes per segment). The Studio dashboard shows it per segment next to `SegmentInfo.pages` and `bytes()`.

### Compaction

`Compactor.compact(liveThreshold)` (default threshold `compaction_live_ratio = 0.5`):

1. **Choose victims.** A victim is a `SEALED` segment, not the active one, with `live < pages × threshold`. Here `pages` is the number of images ever allocated in the segment, so the ratio is *live node images / images written*. If there are no victims, it reports and returns.
1. **Choose victims.** A victim is a `SEALED` segment, not the active one, with `live bytes < allocated bytes × threshold`, so the ratio is *live bytes / bytes written*. Candidates are ordered from emptiest to fullest. The background pass takes only the first one (`compact(threshold, 1)`), which bounds how long a single relocation holds the commit lock; `COMPACT` takes all of them. If there are no victims, it reports and returns.
2. Mark the victims `COMPACTING`.
3. **Relocate.** For every **active** branch of the current generation, call `transactions.rewrite(branch, roots -> roots.map(relocate))`. `TreeWalker.relocate(ref, schema, moving, scope)` rebuilds as a fresh `Ref.Pending` copy:
* every stored node whose page id lies in a victim segment;
Expand Down
4 changes: 2 additions & 2 deletions docs/storage/pages.md
Original file line number Diff line number Diff line change
Expand Up @@ -82,7 +82,7 @@ INTERNAL: Ref × slotCount
| zigzag weightMax | zigzag timeMin | zigzag timeMax | i64le roleBits | i64le fingerprint
```

For an internal node, "keys" are the separators. `separators[0]` is the first child's minimum (set by `Branch.frozenWith`), so the delta stream is non-negative. Each child reference embeds the child's **complete** summary. This lets readers prune and count without loading children, and is why a branch entry costs up to `KEY_BYTES + 8 + 96` bytes in the fanout calculation (`Layout`).
For an internal node, "keys" are the separators. `separators[0]` is the first child's minimum (set by `Branch.frozenWith`), so the delta stream is non-negative. Each child reference embeds the child's image size in 64-byte units and its **complete** summary. This lets readers prune and count without loading children, lets compaction measure live bytes without reading pages, and is why a branch entry costs up to `KEY_BYTES + 8 + 3 + 96` bytes in the fanout calculation (`Layout`).

**Size invariant.** Splits are decided from each value codec's `maxSize`, summed per entry, plus the exact key-stream size. `IncidenceCodec.maxSize` charges 1 byte per entry for leaf-level overhead. The real per-leaf fixed overhead is 1 flags byte plus up to 6 column bitmaps of ⌈n/8⌉ bytes. For n ≥ 8 the per-entry charge covers this: 7·⌈n/8⌉ ≤ n + 7 and the charge is n. For small leaves the excess is at most 7 bytes. `Layout` reserves 32 bytes beyond `leafBudget` (`RESERVED = 32`), which covers it, so an encoded leaf never exceeds `pageSize`. `NodeCodec.encode` encodes into a cursor of exactly `pageSize` bytes and would fail loudly (`node of N entries overflows a P byte page`) if the invariant were broken.

Expand Down Expand Up @@ -252,7 +252,7 @@ Because images are never overwritten in place, a torn write can only damage an i

| Version | Where | Meaning |
|---|---|---|
| `format=2` in `<data>/FORMAT` | `StorageEngine.verifyFormat` | Packed 64-byte-unit extents. A directory with `format=1` (or no format key) was written with fixed page slots, and opening it fails: `uses storage format 1; this build reads format 2 (packed segment extents); export and reload it`. |
| `format=3` in `<data>/FORMAT` | `StorageEngine.verifyFormat` | Packed 64-byte-unit extents, with image sizes in every stored reference. Directories written by older builds (`format=1` fixed page slots, `format=2` references without sizes) fail to open: `uses storage format 2; this build reads format 3 (sized page references); export and reload it`. |
| `page-size=N` in `<data>/FORMAT` | same | Fixed at creation. A mismatching `--page_size` is rejected. |
| header byte 4 = `1` | `PageHeader.FORMAT` | Node image layout; unchanged by packing. |
| `SegmentInfo` in catalog images | `CatalogImage` | Now `(id, u8 state, pages, units, retiredAt)` |
Expand Down
2 changes: 1 addition & 1 deletion docs/storage/persistent-tree.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@ Leaf Branch

```text
leafBudget = pageSize − 80 (PageHeader.SIZE) − 96 (Summary.MAX_ENCODED_BYTES) − 32 (reserve)
maxFanout = leafBudget / (KEY_BYTES 10 + Ref.maxEncodedSize() 8 + 96)
maxFanout = leafBudget / (KEY_BYTES 10 + Ref.maxEncodedSize() (8 + 3 + 96))
maxValueBytes = leafBudget / 2 − KEY_BYTES (largest inline value; larger values throw HStoreException.limit)
leaf underfull ⇔ leaf.bytes() < leafBudget / 4
branch underfull ⇔ branch.size() < max(2, maxFanout / 4)
Expand Down
2 changes: 1 addition & 1 deletion docs/transactions/wal-and-recovery.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@ All fixed-width integers are **little-endian** (`ByteCursor` uses `JAVA_INT_UNAL
| `svarlong` | zigzag (`(v << 1) ^ (v >> 63)`) then `varlong` |
| `blob` | `varint` length, then that many bytes |
| `string` | `blob` of UTF-8 |
| `Ref` | `i64 pageId`; `0` (`PageId.NONE`) means the empty tree and ends the encoding; otherwise followed by a `Summary` |
| `Ref` | `i64 pageId`; `0` (`PageId.NONE`) means the empty tree and ends the encoding; otherwise followed by `varlong units` (the image size in 64-byte units, at most 3 bytes) and a `Summary` |
| `Summary` | `varlong count`; if `count == 0` nothing else; otherwise `svarlong min`, `varlong (max - min)`, `svarlong weightSum`, `svarlong weightMin`, `svarlong weightMax`, `svarlong timeMin`, `svarlong timeMax`, `i64 roleBits`, `i64 fingerprint` (at most `Summary.MAX_ENCODED_BYTES = 96`) |

## 3. Segments and frames
Expand Down
6 changes: 3 additions & 3 deletions engine/src/main/java/io/hstore/engine/StorageEngine.java
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@
public final class StorageEngine implements AutoCloseable {

private static final System.Logger LOG = System.getLogger("hstore.engine");
private static final String FORMAT_VERSION = "2";
private static final String FORMAT_VERSION = "3";

private static final int MAX_ATTEMPTS = 8;
private static final Duration MAINTENANCE_INTERVAL = Duration.ofMillis(500);
Expand Down Expand Up @@ -116,7 +116,7 @@ private static void verifyFormat(Path directory, EngineOptions options) throws I
}
if (!format.getProperty("format", "1").equals(FORMAT_VERSION)) {
throw HStoreException.invalid("database " + directory + " uses storage format " + format.getProperty("format", "1")
+ "; this build reads format " + FORMAT_VERSION + " (packed segment extents); export and reload it");
+ "; this build reads format " + FORMAT_VERSION + " (sized page references); export and reload it");
}
int stored = Integer.parseInt(format.getProperty("page-size"));
if (stored != options.pageSize()) {
Expand Down Expand Up @@ -305,7 +305,7 @@ private void maintain() {
}
if (pages.bytesWritten() - writtenAtCompaction > options.checkpointWalBytes()) {
writtenAtCompaction = pages.bytesWritten();
Compactor.Report report = compactor.compact(options.compactionLiveRatio());
Compactor.Report report = compactor.compact(options.compactionLiveRatio(), 1);
if (!report.compacted().isEmpty()) {
timedCheckpoint();
LOG.log(System.Logger.Level.INFO, "background compaction relocated segments {0}", report.compacted());
Expand Down
18 changes: 13 additions & 5 deletions engine/src/main/java/io/hstore/engine/maintenance/Compactor.java
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,8 @@
import io.hstore.engine.tree.WriteScope;
import io.hstore.engine.txn.TransactionManager;

import java.util.HashSet;
import java.util.Comparator;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
import java.util.OptionalLong;
Expand All @@ -23,7 +24,7 @@

public final class Compactor {

public record Report(Map<Integer, Long> livePages, List<Integer> compacted, long generation) {
public record Report(Map<Integer, Long> liveBytes, List<Integer> compacted, long generation) {
}

private final TransactionManager transactions;
Expand All @@ -37,21 +38,28 @@ public Compactor(TransactionManager transactions, PageStore pages, SlotRegistry
}

public Map<Integer, Long> liveness() {
Set<Long> visited = new HashSet<>();
Map<Long, Integer> visited = new HashMap<>();
TreeWalker walker = new TreeWalker(transactions.source());
Stream.concat(transactions.history().stream(), Stream.of(transactions.current()))
.flatMap(generation -> generation.branches().values().stream())
.forEach(branch -> branch.roots().roots().forEach((slot, ref) ->
walker.visit(ref, slots.slot(slot).schema(), visited)));
return visited.stream().collect(Collectors.groupingBy(PageId::segmentOf, TreeMap::new, Collectors.counting()));
return visited.entrySet().stream().collect(Collectors.groupingBy(entry -> PageId.segmentOf(entry.getKey()), TreeMap::new,
Collectors.summingLong(entry -> (long) entry.getValue() * PageId.UNIT_BYTES)));
}

public Report compact(double liveThreshold) {
return compact(liveThreshold, Integer.MAX_VALUE);
}

public Report compact(double liveThreshold, int maxVictims) {
Map<Integer, Long> live = liveness();
int active = pages.activeSegment();
List<Integer> victims = pages.segments().stream()
.filter(segment -> segment.state() == SegmentState.SEALED && segment.id() != active)
.filter(segment -> live.getOrDefault(segment.id(), 0L) < segment.pages() * liveThreshold)
.filter(segment -> live.getOrDefault(segment.id(), 0L) < segment.bytes() * liveThreshold)
.sorted(Comparator.comparingDouble(segment -> (double) live.getOrDefault(segment.id(), 0L) / Math.max(1, segment.bytes())))
.limit(maxVictims)
.map(SegmentInfo::id)
.toList();
long generation = transactions.current().id();
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,7 @@ public Ref materialize(Ref ref, TreeSchema<?> schema) {
PageHeader.assign(image, pageId);
sink.accept(pageId, image, frozen);
pages++;
return new Ref.Stored(pageId, frozen.summary());
return new Ref.Stored(pageId, PageId.unitsFor(image.byteSize()), frozen.summary());
}

public long pagesWritten() {
Expand Down
Loading
Loading