Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,7 +68,7 @@ Multitask MUST for this product. Index: [`operations/README.md`](operations/READ
| [`operations/product-gateway-contract.md`](operations/product-gateway-contract.md) | Product gateway on memnet-mcp (MN-REQ-06.12): route by product, house, or session |
| [`operations/cluster-route-vs-slice-hand-carry.md`](operations/cluster-route-vs-slice-hand-carry.md) | Two named moves: ClusterRoute vs SliceHandCarry (#191 / #47 cousin; inventOnly) |
| [`operations/admin-usage-report.md`](operations/admin-usage-report.md) | Admin-only serve usage JSON for a product-gate admin MCP (opaque alias; not agent MCP) |
| [`operations/safe-upgrade.md`](operations/safe-upgrade.md) | Drain and restore a live serve without dropping sessions (MN-REQ-06.14) |
| [`operations/safe-upgrade.md`](operations/safe-upgrade.md) | Drain, restart, and restore a live serve inside memnet serve (MN-REQ-06.14) |
| [`operations/one-session-per-document.md`](operations/one-session-per-document.md) | One MemNet session per document over loopback serve (no MCP front) |

Product skill: [`.cursor/skills/memnet-reference/`](../.cursor/skills/memnet-reference/). SysML trail: MN-REQ-12 → [`sysml-models/outputs/multitask-case-study.md`](../sysml-models/outputs/multitask-case-study.md).
Expand Down
2 changes: 1 addition & 1 deletion docs/operations/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ Agent operating doctrine for this product (not domain recipes).
| [`product-gateway-contract.md`](product-gateway-contract.md) | Product gateway on memnet-mcp (MN-REQ-06.12): per-product route to one owning serve |
| [`cluster-route-vs-slice-hand-carry.md`](cluster-route-vs-slice-hand-carry.md) | Two named moves: ClusterRoute vs SliceHandCarry (#191 / #47 cousin; inventOnly) |
| [`admin-usage-report.md`](admin-usage-report.md) | Admin-only serve usage JSON (opaque alias; not agent MCP; MN-REQ-06.11) |
| [`safe-upgrade.md`](safe-upgrade.md) | Drain, restore, and `memnet-upgrade` so a serve swap keeps sessions (MN-REQ-06.14) |
| [`safe-upgrade.md`](safe-upgrade.md) | Drain and restore a live serve from inside memnet serve (MN-REQ-06.14) |
| [`one-session-per-document.md`](one-session-per-document.md) | Product gate: one serve session per document over loopback (no MCP); 0.19.18 probe |

Application pattern for `modelbasedPrj-*` / `SysMLEdgePrj-*`: [`../application-notes/system/llm-system-dev-multitask.md`](../application-notes/system/llm-system-dev-multitask.md). Index: [`../README.md`](../README.md).
106 changes: 18 additions & 88 deletions docs/operations/safe-upgrade.md
Original file line number Diff line number Diff line change
@@ -1,112 +1,42 @@
# Safe serve upgrade

Upgrade `memnet-serve` without silently dropping a loaded session (MN-REQ-06.14). The agent loop does not change: cue, then `pin_map`, then `mutate`. This is a patch-level cut on 0.19. Do not bump the package version from this procedure.
Upgrade `memnet-serve` without dropping a loaded session (MN-REQ-06.14). The procedure stays inside the serve process: save every session, restart on the new version, reload every session. The agent loop does not change: cue, then `pin_map`, then `mutate`.

The manual procedure this replaces is a side-by-side virtualenv: install the new version, save every session, stop the old serve, start the new serve on the same port, reload, and check counts. Anything that failed to save was easy to miss. The drain below names every session it could not snapshot and refuses a ready-to-stop result unless you pass an explicit override.
During the restart, clients see `@ERR: serve_draining|retry_after_s=<seconds>` or a connection error. There is no client retry window.

## Order
## Procedure

Deploy client tolerance **before** the serve swap.
1. Install the new version in a separate virtualenv. Leave the running venv in place.

1. Install the new `memnet-mcp` (and the droplet gateway, if this host is the gateway) so `serve_draining` and a brief connection refusal are retried. Leave the gateway's `pinned_version` on the version that is still running.
2. Confirm `MEMNET_UPGRADE_RETRY_S` is unset or about `30` on those client processes. `0` disables retry.
3. Only then run `memnet-upgrade` against the serve.

During the gap, new `session_open` calls receive `@ERR: serve_draining|retry_after_s=<seconds>` and clients back off. After the old process has exited and before the new one listens, connections are refused. That refusal is retried for the same window. A command the serve already accepted is not retried on timeout.

Downtime that remains: a few seconds while the port is closed, covered by that retry. Sessions keep their ids, CapsPolicy ACL bindings, TTL expiry clock, and house (the session tag map).

## What the serve does
```bash
python -m venv /opt/memnet/venv-new
/opt/memnet/venv-new/bin/pip install "memnet-llm==<new>"
```

Admin credential: `MEMNET_ADMIN_TOKEN` (same rule as the usage report). Unset is `@ERR: admin_unconfigured`. A mismatch is `@ERR: admin_denied`. This command is not an agent MCP tool.
2. Drain the running serve. `MEMNET_ADMIN_TOKEN` is required (unset is `@ERR: admin_unconfigured`; a mismatch is `@ERR: admin_denied`). This command is not an agent MCP tool.

```bash
memnet admin upgrade-prepare --state-dir "$MEMNET_STATE_DIR"
```

Envelope form: `{"upgrade_prepare": true, "admin_token": "<token>"}`.
3. Check ready-to-stop. Stop the process only when stderr contains `@STAT: upgrade_prepare|ready|`. If a session cannot be snapshotted, the command names it, exits non-zero, and does not write a ready manifest. Pass `--allow-unsaved` only when those named sessions may be dropped.

Drain behaviour:
4. Point the unit `ExecStart` at the new venv, on the same port and the same `MEMNET_STATE_DIR`, and restart.

- Refuse new `session_open` with `serve_draining`.
- Finish commands already in flight, then refuse further commands.
- Snapshot every loaded session with the lossless snapshot writer.
- Write `upgrade-manifest.json` and `upgrade-snapshots/` under `MEMNET_STATE_DIR` (default `~/.local/state/memnet`). The manifest records session ids, row and edge counts, sha256 checksums, the serve version, and snapshot format `1`.
- An ACL and clock passport is an extra `# upgrade-passport` line. Older v1 loaders skip `#` lines, so a rollback can still read the graph.
- If any session raises `snapshot_unsaveable` (or any other save error), the command exits non-zero, lists that session in `upgrade-blocked.json`, and does **not** write a ready manifest. The serve keeps accepting work. Pass `--allow-unsaved` only when you accept dropping those named sessions.
- Do not stop the process unless stderr contains `@STAT: upgrade_prepare|ready|`.

The new process reads the manifest on startup, reloads each snapshot, and checks counts and checksums. It prints:
5. Read the restore stat from the new process:

```text
@STAT: upgrade_restore|ok|<n>|failed|<m>
```

Snapshot files stay on disk until `memnet admin upgrade-retire` after a clean restore report. A format this process does not support, or a checksum or parse failure, exits `3` and does not delete or rewrite the files. Restarting before retire replays the upgrade snapshots (writes made after the restore are not in those files). Retire as soon as the stat line is clean.

## Helper

`memnet-upgrade` encodes the side-by-side venv procedure. It refuses to start unless `--clients-ready` is set.

```bash
memnet-upgrade \
--clients-ready \
--new-python /opt/memnet/venv-new/bin/python \
--new-exec "/opt/memnet/venv-new/bin/memnet serve --host 127.0.0.1 --port 18765" \
--state-dir /var/lib/memnet \
--unit /etc/systemd/system/memnet-serve.service \
--restart-unit memnet-serve
```

Steps:
A clean stat retires the manifest inside that process. A later restart does not replay the snapshots. The snapshot files stay on disk. A checksum failure, a parse failure, or an unsupported snapshot format exits `3`, leaves the files untouched, and does not retire.

1. Preflight: import `memnet` with the new interpreter and check `SUPPORTED_SNAPSHOT_FORMATS` against the current manifest (or format `1` when no manifest exists yet).
2. `admin upgrade-prepare` on the running serve.
3. Copy the unit to `memnet-serve.service.bak` and replace `ExecStart=`.
4. If `--gateway-config` and `--product` are set, copy the config to `*.bak` and set that product's `pinned_version` to the new version **after** the old serve is drained and **before** the new process listens.
5. Restart the unit. The new process restores and writes `upgrade-restore.json`.
6. If `failed` is not `0`, write the unit backup and the config backup back and restart again.

`--state-dir` must be the same directory as the unit's `Environment=MEMNET_STATE_DIR`. The serve reads that variable at startup; the prepare command also receives `--state-dir`.

Build the new venv beside the old one. Do not overwrite it until the restore has verified.

```bash
python -m venv /opt/memnet/venv-new
/opt/memnet/venv-new/bin/pip install "memnet-llm==<new>"
```

Use the index you trust. Do not put tokens or session ids in the unit file. Keep `MEMNET_ADMIN_TOKEN` in a root-only `EnvironmentFile`.

## rpi5-syson

Local engine on the Pi:

| Process | Port | Unit (adjust to the host) |
|---------|------|---------------------------|
| `memnet serve` | `18765` | `memnet-serve.service` |
| `memnet-mcp` streamable-http | `18766` | `memnet-mcp-http.service` |

Upgrade the MCP unit first (retry code, same serve version pin). Then `memnet-upgrade` the serve unit on `18765` with `MEMNET_STATE_DIR` on the Pi disk. Leave `18766` up so clients retry against the serve gap.

## Endleaf engine serve

The Endleaf product backend is the serve named in the droplet gateway registry (host and port live in that config, not in this repo). Use that unit's port and state directory. The drain and restore are the same command. Do not point the gateway at a second backend to "move" live sessions.

## Droplet gateway

The gateway process is `memnet-mcp --transport gateway` with `MEMNET_GATEWAY_CONFIG`. Ship the retry-capable build first, with `pinned_version` still equal to the running engine. Then, in the same `memnet-upgrade` invocation that swaps the engine, pass:

```bash
--gateway-config /etc/memnet/gateway.json \
--product endleaf \
--restart-unit memnet-gateway
```
## Rollback

The pin changes only after drain succeeds, and it is restored from `gateway.json.bak` if the engine restore fails. Restart the gateway after the pin write so it stops caching the previous version (`version_cache_s`). No plaintext credentials in git. Hashes stay in the config file on the droplet.
Point the unit `ExecStart` back at the old venv and start it when the restore stat is not clean. After a clean stat the manifest is retired, so a later start does not reload those snapshots.

## Rollback
## Hosts

A failed verification restores the previous `ExecStart=` and the previous pin, then restarts. Snapshot files are still in the state directory. The old venv loads format v1 snapshots (`#` passport lines are ignored). That reload keeps the session id and the graph. The old loader rebases the TTL clock from `ttl_minutes` and does not read ACL bindings from the passport; re-grant those on the old serve if you stay there. The new serve's restore path keeps the id, the ACL bindings, the expiry instant, and the house.
On rpi5-syson the serve listens on `18765` (`memnet-serve.service`). `memnet-mcp` on `18766` is not part of this procedure. The Endleaf engine serve is the backend named in the gateway registry; use that unit's port and state directory. The droplet gateway is unchanged: do not edit its version pin as part of this restart.

Neighbourhood reserves are not part of the snapshot. Take a fresh reserve after the new serve is up.
Do not put tokens or session ids in the unit file. Keep `MEMNET_ADMIN_TOKEN` in a root-only `EnvironmentFile`.
24 changes: 0 additions & 24 deletions parts/common/memnet/memnet/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -569,30 +569,6 @@ def admin_upgrade_prepare(
raise typer.Exit(result.exit_code)


@admin_app.command("upgrade-retire")
def admin_upgrade_retire(
token: Annotated[
str | None,
typer.Option("--token", help="Admin token presented by the caller (do not log)."),
] = None,
state_dir: Annotated[str | None, typer.Option("--state-dir")] = None,
) -> None:
"""Stop replaying an upgrade manifest after a verified restore. Keeps the files."""
from pathlib import Path

from memnet.admin_usage import authenticate, caller_token
from memnet.upgrade import retire_manifest

presented = token if token else caller_token()
try:
authenticate(presented)
retire_manifest(Path(state_dir) if state_dir else None)
except MemNetError as exc:
_handle_error(exc)
emit_stdout('{"ok": true, "retired": true}')
emit_stderr("@STAT: upgrade_retire|1|-")


@session_app.command("expire-status")
def session_expire_status() -> None:
"""Booleans for expire-save (serve_status). Path redacted; no sids."""
Expand Down
4 changes: 0 additions & 4 deletions parts/common/memnet/memnet/serve.py
Original file line number Diff line number Diff line change
Expand Up @@ -92,10 +92,6 @@ def _handle_request(payload: dict[str, Any]) -> dict[str, Any]:
from memnet.upgrade import upgrade_prepare_envelope

return upgrade_prepare_envelope(token_s, allow_unsaved=bool(payload.get("allow_unsaved")))
if payload.get("upgrade_retire") is True:
from memnet.upgrade import upgrade_retire_envelope

return upgrade_retire_envelope(token_s)

argv = payload.get("args", [])
if not isinstance(argv, list):
Expand Down
28 changes: 5 additions & 23 deletions parts/common/memnet/memnet/upgrade.py
Original file line number Diff line number Diff line change
@@ -1,8 +1,9 @@
"""Safe serve upgrade: drain, manifest, and startup restore (MN-REQ-06.14).

Admin-only. Not an agent MCP tool. A session that cannot be snapshotted
blocks ready-to-stop unless the operator passes allow-unsaved. Snapshot
files are not deleted when restore fails.
blocks ready-to-stop unless the operator passes allow-unsaved. A clean
restore retires the manifest in that same startup. Snapshot files are
not deleted, and a failed restore does not retire.
"""

from __future__ import annotations
Expand Down Expand Up @@ -80,10 +81,7 @@ def _drain_wait_s() -> float:
def _is_upgrade_argv(argv: list[Any]) -> bool:
if len(argv) < 2:
return False
return str(argv[0]) == "admin" and str(argv[1]) in {
"upgrade-prepare",
"upgrade-retire",
}
return str(argv[0]) == "admin" and str(argv[1]) == "upgrade-prepare"


class DrainGate:
Expand Down Expand Up @@ -581,6 +579,7 @@ def startup_restore_or_exit(directory: Path | None = None) -> RestoreReport:
sys.stderr.write(line)
if report.failed:
raise SystemExit(3)
retire_manifest(directory)
return report


Expand Down Expand Up @@ -627,20 +626,3 @@ def upgrade_prepare_envelope(
else:
stderr = result.stat + "\n"
return {"exit_code": result.exit_code, "stdout": result.stdout_json(), "stderr": stderr}


def upgrade_retire_envelope(token: str | None) -> dict[str, Any]:
try:
authenticate(token)
retire_manifest()
except MemNetError as exc:
return {
"exit_code": exc.exit_code,
"stdout": "",
"stderr": f"@ERR: {exc.code}|{exc.message}\n",
}
return {
"exit_code": 0,
"stdout": json.dumps({"ok": True, "retired": True}) + "\n",
"stderr": "@STAT: upgrade_retire|1|-\n",
}
67 changes: 0 additions & 67 deletions parts/common/memnet/memnet/upgrade_retry.py

This file was deleted.

Loading
Loading