Skip to content

Harness review: ~50 reproduced bugs fixed, security hardening, swarm-code trust - #1

Merged
skyblanket merged 44 commits into
mainfrom
claude/ecstatic-clarke-2x3202
Sep 24, 2026
Merged

skyblanket merged 44 commits into
mainfrom
claude/ecstatic-clarke-2x3202

Conversation

@skyblanket

Copy link
Copy Markdown
Owner

A full review of the harness. Three review agents plus end-to-end runs against a scripted OpenAI-SSE mock found about 50 reproducible bugs, and this PR fixes them. Each fix has a regression test that fails on the old code.

Depends on skyblanket/swarmrt#5 (merged). CI builds against swarmrt main, and this branch uses its new builtins (stdout_to_stderr, fd_write) and its runtime fixes.

What changes

Security

  • Untrusted project config. A repo's ./.swarm-code.json applies only safe keys (model, max_tokens, …) and can only tighten permissions. Its hooks, MCP servers, endpoint/api_key/providers/profiles and any looser permissions are ignored, with a notice, until you run swarm-code trust in that directory.
  • Trust command. swarm-code trust [DIR], untrust, trust --list, and /trust / /untrust in the REPL edit trusted_projects in ~/.swarm-code/settings.json. Other keys are kept, the file is written atomically and stays indented, and a file that doesn't parse is never overwritten.
  • Network gate. Endpoint URLs are parsed properly: userinfo (127.0.0.1@evil), uppercase schemes, prefix tricks, decimal IPs and fd… hostnames no longer pass. Every endpoint actually dialled is checked (primary, fallback, providers, /profile override).
    • Behavior change: an API key no longer bypasses the gate; remote endpoints need SWARM_CODE_ALLOW_REMOTE=1, as the README already said.
  • Control files. The model can't write swarm-code's own control files (hooks, schedule.json, settings, sessions); memory/ and skills/ stay writable. Path checks are case-insensitive, and browser_screenshot now goes through the same guard.
  • Headless approvals. Headless mode no longer auto-approves an "ask": dangerous commands, tools explicitly set to "ask", and MCP tools are denied unless SWARM_CODE_HEADLESS_APPROVE=1. Scheduled jobs also deny dangerous commands.
  • Command classifier. A token-aware shell classifier gates every command-running tool (bash, background, bg_server, run_tests). It catches respellings like rm -fr / or sh -c '…' and stops flagging words that only appear inside quotes, echo or grep patterns.
  • file_watch. Shell injection is fixed. Hooks with large payloads fail closed. Background tasks get a private directory per session, so one session can no longer kill another's task.
  • Redaction. Secrets are redacted before encoding, which keeps trajectory JSONL valid and covers AWS/PG/URL/YAML/PEM secrets. Skill slugs can't traverse paths, and the request-body debug dump only happens with SWARM_CODE_DEBUG=1 (mode 600).

Tools

  • read works on .js/.ts files (libmagic calls them application/javascript, so they were refused as binary). Binary now means a NUL byte in the first 8 KB, which doesn't depend on file being installed.
  • read reports a missing file as missing, reaches past the first 64 KB, and no longer shows a phantom last line.
  • edit works on CRLF files. It refuses NUL-containing and >1 MB files instead of truncating them, and creates parent directories.
  • bash survives a trailing # comment or heredoc, and shell syntax errors now reach the model. grep never reads stdin (it swallowed the MCP server's next request) and reports regex errors.
  • run_tests quotes its path and times out cleanly. ~ expands in paths.

Agent loop

  • Headless stdout carries only the --json line, or just the answer when stdout is piped, so swarm -p --json … | jq works.
  • A headless run never reports a previous run's answer as its own.
  • -p parses its prompt whatever the flag order, including prompts that start with -.
  • Tool calls cut off by the length limit or by Esc are never executed. Esc on a tool ends the turn instead of calling the model again.
  • The context budget scales with the window: it was negative below 68K.
  • Compaction keeps the live request and never loses history on a failed summary. A fatal 4xx keeps completed tool pairs, and a context overflow retries once.
  • A text rewrite that corrupted < into \< is removed. Profile overrides no longer override your environment variables.
  • The working directory comes from the real cwd, not a stale $PWD.

Multi-agent, MCP, scheduler

  • MCP client: a dead server is detected and reconnected (and no longer kills the process, a runtime fix), server-to-client requests aren't mistaken for replies, and health comes from call status rather than output text.
  • MCP server: follows the JSON-RPC spec for ping, unknown-tool error codes, and id/jsonrpc validation.
  • Subagents: they have their own guardrails and enforced type allow-lists, and return bounded, non-lossy results. A hung LLM call is reported with its reason.
  • /flows and scheduler: a malformed schedule.json or flow definition no longer crashes the session. /flows fan-out is capped, and schedule parsing is strict.
  • Session search reindexes journals that changed.

Testing

make check passes: unit 245/245 (was 160), smoke, integration 43/43 (was 8), module checks 8/8. The key scenarios were also re-run by hand on the final binary:

  • --json | jq;
  • reading JS/TS files;
  • café 日本 coming back intact from bash;
  • Esc during a tool call;
  • an interactive session with a live MCP server, which used to hang.

Notes for reviewers

  • After merging, run swarm-code trust in any repo whose .swarm-code.json you rely on. Otherwise its hooks and endpoint are ignored, with a notice.
  • Remote endpoints, including the default Moonshot setup, need SWARM_CODE_ALLOW_REMOTE=1.
  • Cron and flows that relied on headless auto-approval need SWARM_CODE_HEADLESS_APPROVE=1.

🤖 Generated with Claude Code

https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE


Generated by Claude Code

…rrectness

- Headless stdout carries only the result: the --json line, or the final
  answer when stdout is piped; the transcript (tool calls, streamed
  tokens) goes to stderr. `swarm -p --json ... | jq` used to fail on
  every run. Uses swarmrt's new stdout_to_stderr/fd_write.
- The working directory comes from /bin/sh, not a possibly stale $PWD:
  a launcher that chdir'd without updating PWD made the system prompt
  name another directory.
- ESC/Ctrl-C on a tool ends the turn: the rest of the batch is recorded
  as skipped and the model is not called again.
- read: .js/.ts files are text (libmagic calls them
  application/javascript and they were refused as binary); binary now
  means a NUL byte in the first 8KB, which needs no `file` binary; a
  missing file says "file not found"; an empty file says so; a trailing
  newline no longer adds a phantom last line.
- edit/multi_edit: an LF old_string matches a CRLF file and the edit
  keeps CRLF endings.
- bg_kill test counts live processes only (container inits don't reap
  zombies).

make check: unit 166/166, smoke, integration 10/10 (T9 clean stdout,
T10 stale PWD are new and fail on the previous binary), module 8/8.
Requires swarmrt with eprint/stdout_to_stderr/fd_write.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
bash wrapped the model's command as `( export …; CMD ) </dev/null 2>&1`.
A trailing `# comment` commented out the closing paren and a final heredoc
terminator became `EOF ) </dev/null…`, so `echo hi # say hi` or a
`cat > f <<'EOF' … EOF` write returned `[exit 2]` with empty output — and
sh's syntax error went to the agent's stderr, never to the model.

Util.noninteractive_wrap now puts `exec </dev/null 2>&1; export CI=1 …` on
line 1 (run before sh parses the user's lines, so a syntax error is
captured) and the command on its own lines, newline-terminated. The same
wrapped script is used for the auto-background path and the `background` /
`bg_server` tools (Background.launch_cmd records the raw command for
display), so a command behaves the same however it runs. Util.no_stdin is
the stdin-only variant for the harness's own helper pipelines.

Regression tests (fail before, pass after): t_bash_trailing_comment,
t_bash_heredoc_last, t_bash_syntax_error_reaches_model,
t_bg_trailing_comment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
Config.load merged ./.swarm-code.json over the user's settings with full
authority, so opening swarm-code inside a cloned repo let that repo run
SessionStart/PreToolUse hooks, start mcpServers, point endpoint/api_key/
providers/profiles/fallback_profile at an attacker, and loosen
permissions (an api_key even skipped the network gate).

Project scope is now an allow-list (model, max_tokens, llm_timeout_ms,
chat_template_kwargs, vision); `permissions` entries apply only when
stricter than the user's effective decision (allow < ask < deny), and
merge into the user's map instead of replacing it. Everything else is
dropped and main prints one notice naming the ignored keys (stderr in
--mcp-server mode). Opt-in: list the directory under "trusted_projects"
in ~/.swarm-code/settings.json (exact match on pwd / pwd -P) to apply
the file in full as before. load_one also rejects a non-object JSON top
level instead of crashing map_merge.

Tests:
- integration T11 (fails before: "project SessionStart hook executed";
  passes after, incl. tighten-only perms, user endpoint used, notice
  shown, and the trusted_projects opt-in restoring the hook)
- unit t_project_scope_strips_untrusted, t_project_scope_trusted_applies,
  t_project_ignored_keys, t_project_trusted_dir_match (new API)
- run.sh: run_swarm clears exported opt-in knobs and gains RUN_ENDPOINT
  and RUN_ENV for the security cases

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
grep/glob omitted the path when it was "." so rg would print clean
relative paths — but rg with no path searches its STDIN when stdin isn't a
tty. In --mcp-server mode that is the JSON-RPC stream: grep blocked until
the client closed it (30s timeout otherwise) and swallowed the next
request. Both now always pass the root (`.` by default) and strip the
`./` prefix with sed; anchored globs still match.

An invalid regex (`foo(`) returned "(no matches)": the `2>&1` sat after
`| head`, so rg's error went to the user's terminal and the exit code was
lost. The search's stderr now goes to a private mkstemp file and is
returned as `error: grep failed: <rg message>` (glob likewise). The grep
fallback uses -E so its dialect matches rg's.

Every other helper subprocess in tools.sw (read probes, git_*,
code_search, log_wait, file_watch, sw_check, web_search, web_fetch) now
runs through run_sh, which points stdin at /dev/null.

Regression tests (fail before, pass after): integration T11 (MCP grep with
no path + a second request), t_grep_invalid_regex_surfaces,
t_grep_glob_default_path. run.sh gains INTEG_ONLY="t4 t11" to run a subset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
The local-endpoint check in main.sw prefix-matched the raw string, so
http://127.0.0.1@0.0.0.0:P and http://localhost:1@0.0.0.0:P (userinfo)
and HTTP://0.0.0.0:P (uppercase scheme) reached a non-local host, and
http://127.0.0.1.x.invalid, http://fd-anything.invalid and
http://3221225985:9/ (decimal IPv4) passed as local. It only ran once
at startup on the primary endpoint: SWARM_CODE_FALLBACK_ENDPOINT /
fallback_profile, SWARM_CODE_PROVIDERS_JSON / providers and the
~/.swarm-code/.profile_override endpoint were never checked. And any
api_key skipped the gate, contradicting README/SECURITY.md.

Config.is_local_endpoint now parses the URL the way curl reads it:
http(s) only (any case), any '@' in the authority refused, host
lowercased with the port stripped and chars limited to [a-z0-9._-],
IPv6 only bracketed (::1, fc00::/7, fe80::/10 with a full first group),
a numeric / 0x last label means an IPv4 literal that must be a strict
dotted quad (no octal, no short forms); names are localhost, *.local,
*.ts.net or a bare dot-less name. main's startup gate uses it and no
longer exempts an api_key — SWARM_CODE_ALLOW_REMOTE=1 is the only
opt-in. llm.sw re-checks at the point of dial (stream_call for
chat_native/chat_inband, chat_silent, chat_for_subagent; Plan.generate
too); a refused URL is a fatal non-retried 403, so the fallback / next
provider still gets its own gated attempt, and the reason goes to
stderr (diag is silent headless).

Tests:
- integration T12 (fails before: "startup gate let
  http://127.0.0.1@0.0.0.0:P through (rc 0, 1 requests)"; after: the
  userinfo/uppercase/prefix/api-key cases exit 1 with no request, a
  non-local .profile_override and providers[0] are never dialed, and
  an ALLOW_REMOTE=1 control proves the host was reachable)
- unit t_endpoint_gate_bypasses_refused, t_endpoint_gate_locals_allowed,
  t_endpoint_host_parse, t_endpoint_refusal_reason

Behavior change: a remote endpoint with only an API key now needs
SWARM_CODE_ALLOW_REMOTE=1, as documented; scheme-less endpoints
("host:port") are refused.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
A turn cut off at the output-token limit (finish_reason=length) still
dispatched its tool calls, and the runtime's lenient json_decode turned
arguments cut mid-string into a plausible call: a `write` of half a
config.py ending in 'postgres://prod' reported "ok: wrote 59 bytes". ESC
while tool-call arguments were streaming (sync stream path, stdout piped)
ran the truncated command the same way.

- run_turn: a truncated or user-interrupted turn answers every tool call
  with a not-run result instead of dispatching it (each tool_call id still
  gets its tool message). Truncated -> the F4 recovery (raised max_tokens,
  then the smaller-edits nudge) asks the model to reissue the calls;
  interrupted -> the turn ends like ESC on a running tool. Subagent loop
  gets the same guard.
- llm.sw flags `interrupted` (runtime "[Request interrupted by user]"
  marker) and `truncated` on the RAW content, before inband parsing cuts
  the prose at the first call marker.
- Util.json_args_well_formed: strict structural check (balanced, matched
  braces/brackets outside strings, terminated strings, only whitespace
  after the top-level value), split on quotes so a 200KB write costs ms.
  Used by execute_all's F5 guard, subagent_exec_all and
  tcs_args_malformed, so a cut call is caught even without finish_reason.
- History stores "{}" for malformed arguments (sanitize_tool_calls):
  servers that parse them (vLLM chat templates) would reject every later
  request.
- Integration mock: raw-string arguments, "finish" override, HTTP error
  responses, a "silent" queue for non-streaming requests, raw SSE lines;
  run.sh takes an optional list of test names.

Tests: unit t_json_args_well_formed, t_args_malformed_despite_lenient_decode,
t_cut_turn_reason, t_cut_turn_calls_refused, t_sanitize_cut_tool_calls;
integration T11 (length-truncated write never runs), T12 (cut args without
finish_reason), T13 (interrupt-marked stream) — all three fail on the
previous binary. Real ESC mid-stream checked by hand through a pty.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
… reply matching

Three MCP-client bugs reproduced by review:

1. A server that exits mid-call surfaced as "did not respond" and only
   counted as a timeout strike (status stayed "ok"); the next call then
   wrote into the dead pipe. subprocess_recv_line returns nil EARLY only
   on EOF, so mcp_read_reply now classifies an early nil as 'lost', and
   mcp_note_result marks the server failed at once so the lazy-reconnect
   path runs.
2. A server->client request whose id collides with our in-flight call
   (e.g. {"id":100,"method":"roots/list"}) was taken as the response
   ("MCP response carried no result"). A message carrying `method` is a
   request/notification, never our reply (mcp_msg_kind). Server `ping`
   is answered with {} per spec, every other server request with
   -32601, and we keep reading for the real response.
3. Health checks substring-matched the tool's OUTPUT: a successful
   result containing "connection lost" forced a reconnect (state lost)
   and then "not running" for 60s; "did not respond" in output was a
   timeout strike. mcp_do_call now returns {'ok'|'error'|'timeout'|
   'lost', text} and bookkeeping keys on the status only.

Tests: unit t_mcp_msg_kind_collision, t_mcp_health_structured;
integration A1 (EOF -> lost + reconnect), A2 (colliding id + ping),
A3 (result text never touches health; --json stdout stays one line),
driven by the new tests/integration/fake_mcp.py. A1-A3 fail on the
pre-fix binary (A1 dies with SIGPIPE, exit 141).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…itively

validate_write only protected ~/.swarm-code/settings.json, so with bash
denied the model could still `write` ~/.swarm-code/hooks/pre_tool.sh —
which Hooks.run_pre_tool executes via shell() on the very next tool
call — or schedule.json (recurring headless runs), .profile_override
(LLM endpoint/key redirect) or a sessions/ journal (replayed as history).
All checks compared case-sensitively, so ~/.SSH/ (the same directory on
a case-insensitive macOS filesystem) was writable.

Writes are now refused anywhere under a .swarm-code/ directory except
its memory/ and skills/ data dirs (the only state the agent manages; the
remember / learn_skill tools write there directly, and nothing else
routes a legitimate write through write/edit), plus the directory itself
and any project-scope .swarm-code.json. An unresolved ".." (realpath -m
is missing on older macOS) can't climb out of a data dir. Every
PathGuard comparison (write, read, is_sensitive) runs on the lowercased
resolved path. SWARM_CODE_UNSAFE_WRITES=1 still lifts the write guards.

Tests:
- integration T13 (fails before: "BREACH: a model-written pre_tool hook
  executed"; after: hooks/, schedule.json, .profile_override (via edit)
  and ~/.SSH writes are refused, memory/ still writable)
- unit t_pathguard_control_files_blocked, t_pathguard_data_dirs_writable,
  t_pathguard_case_insensitive (the first and last fail before)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…ation

`swarm --mcp-server` deviated from the MCP / JSON-RPC 2.0 specs:
  * ping returned -32601; spec requires an empty result {}.
  * tools/call with an unknown tool name returned -32601 (method not
    found); spec: -32602 Invalid params. Non-object params/arguments
    and non-string names are -32602 too.
  * "jsonrpc" was never checked; anything but exactly "2.0" is now
    -32600 Invalid Request.
  * object / array / boolean / null ids were accepted and echoed back;
    ids must be a string or number -> -32600 with id null. Presence of
    `id` is checked by key (map_has_key is false for a null value), so
    `"id": null` is no longer mistaken for a notification.
  * an id-bearing "notifications/initialized" got no reply at all; it
    is now acknowledged (every request is owed a response).

Tests: unit t_mcp_server_spec_envelope; integration A4 drives the real
binary over stdio and validates every response with python; smoke.sh
gains a ping -> {} check. All fail on the pre-fix binary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
resolve_permission turned every 'ask' into 'allow' when headless, so
{"permissions":{"bash":"ask"}} + `swarm-code -p` ran `rm -rf ~/victim`
unprompted, the dangerous-bash gate was a no-op outside /flows (which
sets SWARM_CODE_DENY_DANGEROUS=1), and MCP tools (default 'ask', being
external and unvetted) auto-ran.

An 'ask' only arises for a dangerous bash command, a user-configured
"ask", or an MCP tool — default-allowed tools never reach it — so
headless now DENIES it unless the user exported
SWARM_CODE_HEADLESS_APPROVE=1 (restoring the old auto-approval; /flows
children still hard-deny dangerous commands via DENY_DANGEROUS). The
model-facing denial (permission_denial) names the opt-in, or the
per-tool "allow" setting when that would help; hardline / configured
denies keep the plain message. Documented in README, SECURITY.md and
--help; stale comments in config.sw / Flows.sw updated.

Tests:
- integration T14 (fails before: "headless auto-approved an explicit
  \"ask\" permission"; after: the ask and a dangerous rm -rf ~ are
  denied with the opt-in named, and HEADLESS_APPROVE=1 restores it)
- unit t_headless_ask_denied (fails before), and
  t_headless_default_allowed_still_run

Behavior change: headless runs that relied on auto-approving an "ask"
tool, an MCP tool, or a dangerous command must now set
SWARM_CODE_HEADLESS_APPROVE=1 (or "allow" that tool in settings.json).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
Resume is the headless default, and run_headless took the last assistant
message ANYWHERE in the (resumed) history as the result. A run whose LLM
call failed — HTTP 400, server down — therefore printed the previous
run's reply as {"status":"ok","summary":"FIRST_RUN_ANSWER"} and exited 0.

run_headless now gives the run a uuid; run_turn stamps the assistant
messages it appends with it (in memory only — the journal drops the
field, and compaction can't shift it the way a length/index comparison
would). The result is the history's FINAL message when it is a reply
stamped by this run with no pending tool calls; otherwise status error /
exit 1 (LLM failure, max steps, nothing appended). A slash-command prompt
(/compact, /profile …) counts as success with an empty summary instead of
reporting a stale answer.

Tests: unit t_headless_answer_this_run_only; integration T14 (run 1 ok,
run 2 same HOME against an HTTP 400 and run 3 with no server must both be
status error / exit 1) — fails on the previous binary with
{"status":"ok","summary":"FIRST_RUN_ANSWER_T14"}.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
chat_inband wrote every request body — the whole conversation, prompts,
tool output and any secrets in them — to /tmp/swarm-code-last-body.json:
a fixed, shared path created 0644 (world-readable, and pre-creatable by
another local user). chat_native wrote the same body to
~/.swarm-code/last-body.json, also 0644, on every call. Nothing in the
code reads either file back.

Both now call dump_last_body, which writes only when SWARM_CODE_DEBUG=1
(the existing debug knob), to ~/.swarm-code/last-body.json via
file_temp (mkstemp, 0600) + rename, so the body is never world-readable
even briefly. Noted in SECURITY.md.

Tests:
- integration T15 (fails before: "inband: request body written to
  /tmp/swarm-code-last-body.json"; after: neither inband nor native
  writes a dump by default, and SWARM_CODE_DEBUG=1 writes
  ~/.swarm-code/last-body.json with mode 0600)

Behavior change: ~/.swarm-code/last-body.json is no longer refreshed on
every native call; set SWARM_CODE_DEBUG=1 to get it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
file_watch spliced the model's path into `p="…"` via shell_inner_quote,
which only escaped `"` — `$(touch PWNED)` or backticks in the path ran,
even with bash denied. It now uses Util.shell_q, and shell_inner_quote
(its only user) is removed.

It also used BSD-only `stat -f %m`. On GNU, `-f` is --file-system, so it
printed free-space counters that change constantly and reported
"ok: changed" in 0.5s for an untouched file. The poll loop now detects the
stat flavour once and compares an mtime+size signature (GNU `%y %s`, BSD
`%m %z`), so two writes in the same second still register.

log_wait/file_watch `timeout_sec` is clamped to [1, 600] (default 60,
junk → 60): 0 used to reach shell_managed as "no timeout" (600s headless,
unbounded on a TTY). Schema text now says 60, not 30.

Regression tests (fail before, pass after): t_file_watch_no_injection,
t_file_watch_portable_mtime, t_wait_timeout_clamped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
CHANGELOG [Unreleased] gains a Security section covering the five fixes
and their user-visible changes (API key no longer lifts the network
gate, headless denies 'ask' without SWARM_CODE_HEADLESS_APPROVE=1,
untrusted project config, protected control files, debug-only body
dump). `swarm doctor` now warns when the resolved endpoint would be
refused at startup, naming SWARM_CODE_ALLOW_REMOTE=1 — the most likely
surprise for users who relied on the API-key exemption.

No new test: doctor output is advisory (smoke already runs doctor).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…dless asks

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…ous bash, strict exprs

Four scheduler bugs reproduced by review:

5. A wrong-shape schedule.json crashed every interactive session ~2s
   after launch, inside main's heartbeat handler: {"jobs":[]} → hd() on
   a map in prune_jobs_loop; "last_run":"yesterday" → "str" + int in
   compute_next_fire. read_state_at now requires a JSON array;
   normalize_job validates id / expr / prompt and coerces numeric
   fields (numeric strings, floats); invalid entries are skipped (left
   untouched on disk) with ONE warning per distinct problem set,
   printed with print_above so the pinned prompt survives.
6. /schedule on a corrupt schedule.json treated it as [] and replaced
   every job. The runtime json_decode guesses through damage ("[{…} ,,,
   oops" decodes to a nil-padded list), so corruption is now decided by
   the new strict RFC 8259 checker JsonCheck.valid; add refuses to write
   with a clear message, and appends to the RAW entries so entries it
   can't validate survive. All writes use file_atomic_write.
7. Dispatched jobs ran headless children that auto-approve 'ask', so a
   job whose model ran `rm -rf ~/victim` deleted it. dispatch_cmd now
   prefixes SWARM_CODE_DENY_DANGEROUS=1, same as /flows.
8. Expression parsing read leading digits and ignored the rest: 1.5h
   ran hourly, 10x5m every 10m, "daily :" at 00:00, "daily 9:5x" at
   09:05. parse_interval / daily_time_ms are strict (expr_error gives
   the reason, printed by add); `hourly` now fires at the top of the
   hour (UTC) as documented, 1h stays relative. README example fixed
   to the syntax /schedule actually accepts.

Note: agent.sw's /schedule handler (not in this change's scope) still
prints its generic "invalid EXPR" hint after the specific reason;
switching it to Scheduler.add_checked would drop that second line.

Tests: unit t_json_check_strict, t_sched_wrong_shape_never_panics,
t_sched_corrupt_refuses_write, t_sched_strict_exprs,
t_sched_hourly_on_the_hour, t_sched_dispatch_denies_dangerous;
integration A5 (session survives + stays responsive, one warning),
A6 (corrupt file byte-identical after /schedule), A7 (due job's child
is denied `rm -rf ~/victim`) — interactive, via the new
tests/integration/pty_run.py. A5-A7 fail on the pre-fix binary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
The hardline and dangerous-command gates only looked at the `bash` tool
and matched raw substrings:

* `background`, `bg_server` and `run_tests.command` ran shell commands
  with no gate at all (a hardline command via `background` executed).
* Trivially bypassed: `rm -r -f /`, `rm -fr /`, `rm -Rf /*`,
  `rm -rf  /*`, `rm -rf --no-preserve-root /`, `rm -rf "$HOME"`,
  `:(){ :|:& };:`, `chmod -R 000 /`, `dd of=/dev/sda if=…`, `sudo<TAB>ls`.
* False positives with no override: `grep -r shutdown src/`,
  `echo reboot required`, `git commit -m 'halt the build'`, or a heredoc
  writing `def shutdown(` were hard-denied as a bare "permission denied".

New src/CommandGuard.sw parses the command like sh (quotes, escapes,
comments, separators, redirections, heredoc bodies as data, $(…)/`…`/<(…)
parsed recursively), then judges each simple command by its command word
and flags after unwrapping assignments, sudo/env/nohup/timeout/xargs/…,
`sh -c SCRIPT` and `eval`. Hardline: rm -r on / or /* or with
--no-preserve-root, mkfs*/mkswap, dd or redirection onto a raw disk,
shutdown/reboot/halt/poweroff/telinit/init 0|6/systemctl poweroff,
chmod/chown/chgrp -R on /, fork bombs. Dangerous (ask): sudo/doas,
rm -r on ~/$HOME, rm -rf on home/system paths, dd to other devices.

Config.check_permission runs it for every tool in Config.command_of; a
configured "deny" now stays deny for a dangerous command (it used to
become "ask", i.e. allowed headless). Denials name the matched pattern and
the offending simple command (Config.denial_message, used by
ToolExecutor and the agent loop). The tool-level sudo refusal is
token-aware and covers background/bg_server/run_tests too.

Regression tests (fail before, pass after): t_classifier_catches_bypasses,
t_classifier_no_false_positives, t_command_tools_gated,
t_denial_names_reason, t_sudo_tab_blocked, integration T12.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
context_budget_tokens was window - 16384 - 52000 with no clamp: below a
68,385-token window it went negative (SWARM_CODE_MAX_TOKENS=32768 gave
-35,616), so every step printed "context at ~1197 / -35616 tokens,
compacting" and ran the compaction path; llm.sw clamped only its own copy
for the context meter (which then read x/1 tok).

One definition now, LLM.context_budget_tokens (agent.sw delegates; llm.sw
can't import agent.sw): reserve and buffer are each capped at window/4
(explicit env values too), so the budget is always >= window/2 with a
floor of 1 — 8K->4096, 32K->16384, 128K->81920, 262K->193760 (unchanged).
The context meter uses the same function.

Tests: unit t_context_budget_scales_with_window; integration T15
(SWARM_CODE_MAX_TOKENS=32768: no "compacting" on a two-step turn, meter
reads /16k) — fails on the previous binary at "-35616 tokens, compacting".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
compact_history kept "system + last 16" and summarized the rest, which:
- with 10-17 messages summarized an EMPTY transcript and prepended one
  more summary per call (summaries stacked, history grew);
- on a failed summarizer (503) still deleted the old messages, replacing
  them with "[compaction failed, messages elided]" — journaled;
- mid-turn (a long tool loop), summarized the live user request away, so
  the following requests carried no user message at all.

Now the verbatim tail is pulled back to start at or before the most
recent user message and never on a tool result cut off from its
assistant (compact_split / pair_start). Nothing old enough -> history
unchanged with no LLM call; empty or failed summary -> history unchanged;
an earlier summary is passed to the summarizer to fold in and replaced,
not stacked. /compact only reports a count when something changed.

Tests: unit t_compact_split_keeps_live_user, t_compact_split_pair_boundary,
t_compact_nothing_old_is_noop, t_compact_failed_summary_keeps_history;
integration T16 (/compact on seeded journals: 12 messages -> no summarizer
call; an old summary is merged, one summary left; a 503 summarizer keeps
all 30 messages) and T17 (12 tool rounds cross a 32K budget mid-turn: every
request still carries the live request, only pre-turn history is
summarized). T16 and T17 fail on the previous binary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
Background task ids restart at bg-0 in every session and the task files
lived at fixed /tmp/swarm-code-bg-N.{log,pid,exit}. Two sessions (or two
worktrees running the test suite) collided: B's launch `rm -f`'d A's
files, B's pid file overwrote A's so `bg_kill bg-0` in A SIGTERMed B's
task, and a stale exit file could mark a fresh task done.

Each Background table now gets a private directory on first launch
(`mktemp -d` under $TMPDIR or /tmp, mode 0700, stored as 'dir' in the
table); log/pid/exit files live there. log_path_for / pid_file_for /
exit_file_for take the table; tools.sw (bash auto-bg, bg_server,
log_wait task_id) and Flows (task status + log stats) use them instead of
building /tmp paths. A launch that can't create the directory returns an
error, which background/bg_server now surface instead of formatting it
as a task id.

Regression test: t_bg_sessions_isolated (two tables, same bg-0 id:
separate 0700 dirs, A's kill leaves B's task running). On the old code
A's kill_task reported B's pid ("pid <B>: killed") and A kept running.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…d non-lossy results

Three subagent bugs reproduced by review (agent.sw subagent block only):

9. sub_opts inherited the parent's guardrails_table, so a subagent's 8
   failing reads set the PARENT's halt_reason: the parent printed
   [guardrail halt], never saw the result, and exited {"status":"error"}.
   Each subagent now gets a fresh ToolGuardrails table (dropped after
   use — ETS tables are finite); run_subagent_loop checks the
   subagent's own halt after each tool batch and returns its partial
   work plus the halt reason.
10. explore/bash restrictions were prompt-only (an explore subagent ran
    bash and write). New ToolRegistry contexts subagent_explore
    (read-only inspection: read/glob/grep/git_status/git_diff/
    code_search) and subagent_bash (bash) are enforced twice: in
    subagent_exec_all with a clear refusal naming what is allowed, and
    by ToolExecutor via the subagent's execution_context. Unknown/odd
    subagent_type values (null, 5) normalise to general.
11. Results were lossy/unbounded: max steps returned only a notice
    (work discarded) — subagent_partial now returns the latest notes and
    a digest of the most recent tool calls/results with every abnormal
    stop; a 320KB answer reached the parent uncapped — capped at 24000
    like bash/MCP (head+tail, marker, UTF-8-safe cuts); the LLM worker
    was unlinked with a fixed 300s wait — it is now spawn_monitor'ed
    (a crash surfaces its reason), bounded by Config.llm_timeout_ms
    (SWARM_CODE_LLM_TIMEOUT_MS / settings), reports why it failed
    (timeout / rejected / retries exhausted), and an abandoned worker is
    killed so it cannot retry the hung request or hold headless exit.

Test infra: mock_llm.py gains "delay" (hung endpoint), "chunk" (many
small SSE deltas, as real servers send — the runtime cuts one huge
delta) and a per-request arrival time "t".

Tests: unit t_subagent_type_allowlists, t_subagent_result_capped,
t_subagent_partial_keeps_work; integration A8 (subagent streak never
halts the parent), A9 (explore can't bash/write), A10 (max steps keeps
FINDING_*), A11 (320KB answer capped), A12 (hung endpoint → reason
after the 3s LLM timeout, parent unblocked). A8-A12 fail on the
pre-fix binary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…once

On a fatal request error run_turn dropped the turn back to its user
message (drop_to_last_clean_user) and a poisoned session was cut back the
same way on resume: two writes landed on disk, a 400 followed, and the
journal — and the model after resume — had no record of them.

- Fatal and transient failures now both keep every completed
  assistant+tool pair and trim only the dangling tail (trim_incomplete).
  The old reason to drop — re-sending cut-off tool arguments — is gone
  (they are stored as "{}"). A poisoned session resumes the same way; the
  poison flag now only changes the resume notice.
- A rejection that says the context is too long ("maximum context
  length", "context_length_exceeded", "too many tokens", "prompt is too
  long", llama.cpp's "exceeds the available context size") is
  recoverable: llm.sw records the fatal message, run_turn shrinks the
  history to half its size (mechanical trim first, then summarize) and
  retries ONCE per turn — only if that actually made it smaller.
- Mechanical trim protects the most recent user message (mid-turn it is
  no longer the last message and a pasted 10KB request got stubbed); the
  overflow retry may also stub the last result, which the server already
  refused.

Tests: unit t_context_overflow_detection, t_mech_trim_protects_live_user;
integration T18 (a: overflow after a 12KB result is trimmed and retried,
the turn succeeds; b: a 400 after two writes keeps both results in the
journal and a resumed run sends them; c: repeated overflow retries once).
T18 fails on the previous binary; its journal kept only the user message.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
Both hook systems spliced the args JSON into the shell command line and an
env var. A write/edit over ~128KB hit E2BIG: system() failed, shell() then
polled 120s for an exit file that never came, and
  * a vetoing ~/.swarm-code/hooks/pre_tool.sh was SKIPPED (call allowed),
  * a settings.json PreToolUse hook reported `block` after 120s even when
    it would have allowed the call.
The post_tool/pre_llm/post_llm hooks stalled the same way.

Hooks.run_with_payload writes the payload to a private mkstemp file
(0600), makes it the hook's stdin, names it in $SWARM_HOOK_DATA_FILE /
$SWARM_CODE_ARGS_FILE, and still exports the inline env var (read from the
file inside the script, never the command line) when it is under 100KB
(else *_OMITTED=1). Hooks run under shell_managed (process-group kill on
timeout) instead of shell() + a perl alarm: 5s for filesystem hooks, 60s
for settings.json hooks.

Fail closed: if the payload file can't be created/written or the hook
can't be launched, pre_tool.sh vetoes and PreToolUse blocks, and the
error now says why (hook, exit code / timeout, start of its output) via
Config.run_hooks_verdict, instead of a bare "blocked by PreToolUse hook".

Regression tests: t_pre_tool_hook_big_payload,
t_configured_hook_big_payload, t_pre_tool_hook_fails_closed. On the old
code a probe showed veto=false after 120104ms and PreToolUse => block
after 120115ms for a hook that exits 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
12. `/flows` with a malformed workflow killed the interactive session:
    {"phases":"oops"} reached init_phases, which hd()'d a string. The
    workflow file is now checked with the strict JsonCheck.valid (the
    runtime decoder happily "parses" a truncated file into a partial
    workflow) and validate_workflow checks every shape the run walks —
    object root, phases/tasks arrays, task objects with a non-empty
    string prompt, string model/label — before the alt-screen opens or
    anything launches. A bad file prints one error line and returns.

    Fan-out was unbounded (12 tasks -> 12 children in 0.15s). A phase's
    tasks are now queued and launched at most flows_max_parallel() at a
    time (SWARM_CODE_FLOWS_MAX_PARALLEL, default 4); every render tick
    refills free slots (fill_phase_slots / launch_quota). A task whose
    Background.launch fails is finished as 'error' instead of carrying a
    bogus id that read 'pending' forever and hung the flow.

Tests: unit t_flows_validate_shapes, t_flows_launch_quota; integration
A13 (three malformed workflows: error shown, session survives and
answers, nothing launched) and A14 (6 tasks, cap 2: peak 2 in flight,
all 6 complete), interactive via pty_run.py. Both fail on the pre-fix
binary (A13: session dies; A14: 6 children at once).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
Config.matches treated the matcher as a raw substring in either
direction, so "edit|write" never fired for multi_edit and a "bash" guard
hook never saw background / bg_server / run_tests / file_watch — the
other tools that run shell commands — and Claude-Code-style "Bash" or
"Edit|Write" never matched at all.

A matcher is now `*` or "|"/","-separated alternatives, compared
case-insensitively (underscores ignored). An alternative matches a tool
whose name contains it, or a tool in its family: bash/shell → bash,
background, bg_server, run_tests, file_watch, log_wait; edit → edit,
multi_edit; multiedit → multi_edit. The spurious reverse match (matcher
"multi_edit" firing for "edit") is gone. Documented in config.sw and a
new README "Hooks" section (which also covers the stdin / *_FILE payload
contract from the previous commit).

Regression test: t_hook_matcher_families (fails before, passes after).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
fix_json_unicode_escapes replaced u0026/u003c/u003e/u0027/u0022/u002f/
u003d/u005c with the characters they name, on native prose and on the
raw inband content before tool-call parsing — a workaround for an old
runtime bug that no longer exists (the runtime decodes \uXXXX). It only
corrupted real text: in inband mode (the default for local endpoints) a
write of JS source `js = "<p>";` landed as `js = "\<p\>";`, and
prose that mentions < showed "\<". Removed with its three call sites
(chat_native, chat_inband, chat_for_subagent); content is used verbatim.

Tests: integration T19 — "<div>", a bare "<", a JSON-escaped <
inside the arguments and a literal < in the file round-trip byte
for byte, in the written file and in the prose, for native and inband.
Fails on the previous binary (native prose "\<div\>", inband file
"\<p\>"). run_swarm now lets RUN_ENV override the harness defaults and
RUN_UNSET remove them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…N/YAML/PEM secrets

13. Trajectory export (and Log.event) ran Log.redact over the ENCODED
    JSON line. The long-blob layer treated the `n` of a `\n` escape as
    the start of a run, producing `\[REDACTED]` — an invalid JSON escape,
    so exports were not valid JSONL. And several common secret shapes
    leaked: AWS_SECRET_ACCESS_KEY=wJalr…/K7MDENG/… (the '/' path
    exemption), PGPASSWORD=…, postgres://admin:pw@host, {"password": "…"}
    / "api_key": "…" (space after the colon), YAML `password: …`, PEM
    lines containing '/'.

    * Log.redact_value walks decoded values (maps/lists, keys kept) and
      redacts each string; event() and Trajectory.export_* encode AFTER
      redacting, so escapes are never touched.
    * New layers in Log.redact: PEM blocks masked wholesale (BEGIN/END
      lines kept; unterminated → to end), URL userinfo passwords
      (scheme://user:[REDACTED]@host), and a case-insensitive key/value
      scanner for identifiers ending in a secret word (password, passwd,
      passphrase, secret, api_key/apikey/api-key, access_key,
      private_key, authorization; token/_key/-key/credential(s) keep the
      old >= 8 threshold) with optional quotes/whitespace around
      = : => (not == / ::); YAML/header `:` values run to end of line for
      the unambiguous words; bare true/false/null and already-masked
      values are left alone.
    * The blob layer treats a backslash AND the byte it escapes as a run
      boundary; `;` and `)` end unquoted values.

Tests: unit t_trajectory_redacts_valid_jsonl (every exported line is
strict JSON via JsonCheck.valid and contains none of 10 planted
secrets) and t_log_redact_value_shapes (nested redaction, structure
kept, benign text like `max_token: 4096` / paths untouched). The export
logic fails on the pre-fix code (invalid `\[` escape, 3+ leaks).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
read loaded only `head -c 65000` of a larger file and then applied
offset/limit, so `offset: 15000` on a 20,000-line file returned an empty
string with no marker, and nothing past the first 64KB was reachable.

Now the window is chosen by line number from the whole file: files within
file_read's 1MB cap are split in-process (so UTF-8 is preserved); bigger
files are windowed with `sed -n 'A,Bp;Bq' | head -c` (bounded, streaming —
still no multi-GB slurp) plus a `wc -l` line count. A NUL byte past the
first 8KB (file_read stops there) falls back to the sed path with a note.

The OUTPUT is capped at 65000 bytes on a line boundary (a single giant
line is cut and flagged) and every cut says where to continue:
  [output capped at 65000 bytes — lines 1-1639 of 20000 shown; continue with offset=1640]
  [lines 1-2000 of 20000 shown — continue with offset=2001]
An offset past EOF is an explicit error naming the line count. Junk
offset/limit values fall back to the defaults (the to_int builtin returned
nil and leaked "nil" into the line numbers). Util.join_all does the
O(n log n) join the old per-line `acc ++` did quadratically.

Regression tests (fail before, pass after): t_read_offset_past_64k,
t_read_huge_file_window.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
… them

edit/multi_edit load the file with file_read, which is C-string based with
a 1MB cap:
  * a file with a NUL byte (`HEADER abc\0\1\2 tail`) was read up to the
    NUL, edited, written back as `HDR abc` — silent truncation, "ok".
  * a file over 1MB made file_read return nil, which edit took for a
    MISSING file, so `old_string: ""` overwrote it with new_string.

load_for_edit compares the loaded length with the on-disk size
(file_stat) — the whole file, not just a prefix — and refuses both cases
with a clear error (the file is left untouched; use bash/python). A
directory is refused too. write's overwrite-diff capture skips a file
whose prior content doesn't match its size, so no bogus diff is rendered.

Regression test (fails before, passes after): t_edit_refuses_nul_and_huge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
14. recall/forget spliced the model-supplied slug straight into
    skills_dir()/<slug>/SKILL.md, so forget_skill {"slug":
    "../../../work/proj"} deleted work/proj/SKILL.md and recall_skill
    read it. Skills.valid_slug now requires ONE safe path component —
    non-empty, <= 128 bytes, no '/', '\', '..', leading '.' or control
    characters — and recall/forget return a clear error otherwise. save
    (which slugifies to [a-z0-9_]) additionally refuses an empty slug,
    which used to write skills_dir()//SKILL.md.

Test: unit t_skill_slug_traversal_blocked plants a SKILL.md outside
skills_dir, aims a ../-slug at it, and requires recall to not return it
and forget to leave it on disk (both succeeded pre-fix), plus
valid/invalid slug cases.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…indexing

15. reindex only checked whether a journal was PRESENT in meta. A
    journal that another instance was still writing when this one booted
    got indexed mid-session, marked done, and was never refreshed — its
    later turns were unsearchable forever (unless it happened to be this
    instance's own .active journal).

    meta now records size + mtime (file_stat: an in-process stat(2) —
    the old "a stat costs a 1s shell() poll" reason for skipping the
    check no longer holds) and a journal is reindexed when either
    differs. The sig is taken BEFORE reading, so a journal that grows
    during ingestion is reindexed again next boot, never missed. Old
    index.db files are migrated with ALTER TABLE (NULL sizes → one
    reindex). Unchanged journals are still skipped. init_at / search_at
    take an explicit directory (init / search unchanged for callers).

Test: unit t_session_search_reindexes_changed — index a journal,
append a turn, re-init: the appended term must be found (it was not
pre-fix) and nothing is duplicated on a further re-init.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
TestRunner ran `cd <repo_path> && <cmd> 2>&1` through shell():
  * repo_path was unquoted — a path with a space failed with exit 2, and
    `/tmp; touch PWNED` ran the injected command;
  * shell() has no timeout or ESC, so a suite over 120s came back as
    "Exit code: -1" with no output and was left running, orphaned;
  * stdin was not explicitly /dev/null.

run_tests_timed now runs Util.noninteractive_wrap("cd '<repo>' || exit 2
<cmd>") under shell_managed (own process group, killed on timeout/ESC),
default 300s, overridable with the new `timeout_ms` arg (1s..600s). On
timeout the partial output is kept with a banner saying so. do_run_tests
also shows the output tail whenever the exit code is non-zero or the run
timed out, not only when the parser counted failures (a compile error
used to show just "Exit code: 2").

Regression test (fails before, passes after):
t_run_tests_quoting_timeout_output.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…he profile

The ~/.swarm-code/.profile_override file is global and persistent:
- A leftover override beat env vars: SWARM_CODE_MODEL=test sent the model
  of a /profile run days earlier (and its endpoint beat
  SWARM_CODE_ENDPOINT). Overrides are now stamped with a per-launch
  session_id (main.sw): this session's own /profile or /model still beats
  env, but one from an earlier session is only a default — an env var
  that is set wins, per field. /profile shows fields the environment
  shadows as "not applied".
- profile_to_override omitted chat_template_kwargs and apply_override
  then cleared the launch value, so `/profile qwen2` dropped
  enable_thinking:false. The profile's kwargs are written and applied; a
  /model-only override leaves them alone.
- /model X overwrote the file with {model: X}, silently reverting a
  /profile's endpoint/api_key while printing "unchanged". It now carries
  forward whatever part of the override is in effect
  (LLM.effective_override) and changes only the model.
- The system prompt is built once for the launch tool_format; after an
  override to inband the request had no tools array while the prompt said
  tools were in it. build_request_body now swaps the prompt's tool-use
  sections (Prompts.tool_sections, pure text) to match the format each
  request is sent in — overrides, fallback endpoints and provider chains
  alike; the history's system message is untouched.
- "/profile clear" (documented alias) was swallowed by the /profile NAME
  branch as a profile named "clear".

Tests: unit t_override_env_beats_stale_override,
t_profile_override_keeps_chat_template_kwargs,
t_model_override_keeps_active_profile, t_system_prompt_follows_wire_format;
integration T20 (settings profile; `/profile qwen2`, then a new run with
SWARM_CODE_MODEL/ENDPOINT set sends model "test" to the env endpoint with
enable_thinking:false; `/model other-model` keeps the kwargs) — fails on
the previous binary (the stale override's endpoint beat the env one).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…reate

t_skill_slug_traversal_blocked and t_session_search_reindexes_changed
emptied their mkstemp-derived directories but left the directories
themselves in /tmp on every run (file_delete is unlink(2); there is no
rmdir builtin). Remove them with exec_argv("rmdir", …).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…ema text

Low-severity items from the tool-layer review:

* `~` / `~/…` were never expanded (tool paths don't pass through a shell
  unquoted): `read ~/.bashrc` said "file not found" and `write ~/x` made a
  literal `./~/` directory. resolve_path_arg (read/write/edit/multi_edit)
  and the path args of glob/grep/code_search/log_wait/file_watch/run_tests
  now expand them to $HOME. write's overwrite-diff key stays the raw arg,
  which is what the agent looks it up by.
* edit with old_string="" on a missing file didn't create parent
  directories (write does), so a new file in a new directory failed.
* The 8-consecutive-failures guardrail only counted results starting
  "error:"; bash reports failure as "[exit N]", so it never fired for bash.
  A non-zero "[exit N]" or a "[timed out" banner now counts as a failure.
* Schema text disagreed with the code: grep said "Returns file paths by
  default" but returns matching lines; glob claimed mtime ordering it
  doesn't do. (log_wait/file_watch "default 30" was fixed with the clamp.)

Regression tests (fail before, pass after): t_edit_create_makes_parents,
t_tilde_paths_expand, t_guardrail_counts_bash_exit_codes,
t_schema_text_matches_code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…with "-"

get_print_arg treated ANY argument starting with "-" after -p as "read
the prompt from stdin":
- README's `swarm-code -p --json "list the test files"` read an empty
  stdin, sent 0 requests and exited 1;
- a prompt that itself starts with "-" (`-p "-- summarize the diff"`,
  `-p "-x ..."`) was dropped the same way — Scheduler and Flows pass user
  prompts as the argument after -p, so such jobs silently did nothing.

Now only an exact "-" means stdin; a KNOWN flag right after -p (--json,
--no-resume, --profile NAME, …) is skipped and the first positional
argument after -p is the prompt (stdin when there is none); anything
else after -p is the prompt, whatever it starts with. Positional-profile
detection skips that same argument, so `-p --json gemma` is a prompt, not
a profile selection. Usage text mentions `-p -`.

Tests: integration T21 — `-p --json "..."`, `--json -p "..."`,
`-p "-- summarize the diff ..."`, `-p "-x ..."`, `-p - <stdin`, and
`-p --json <stdin` each send exactly their prompt and exit 0. Fails on the
previous binary (the README form exits 1 with no request).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…oks, background isolation

Conflicts resolved against the security and agents merges: headless
denials keep the HEADLESS_APPROVE guidance and add the classifier's
matched pattern; the tools branch's integration T11/T12 are renamed
T16/T17 (the security branch already used T11/T12) and every test runs
through the INTEG_ONLY loop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…ls merge)

The merge resolution kept both appended test blocks but dropped the
closing brace between them, so test_runner.sw did not parse. This is the
tree the merged make check ran on (unit 224/224, integration 31/31).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…on, 4xx recovery, profiles, -p parsing

Conflicts resolved against the security, agents and tools merges:
- the loop branch's integration T11-T21 are renumbered T18-T28;
  run_swarm takes RUN_ENV as an array or a string, plus RUN_UNSET /
  RUN_STDIN / RUN_ENDPOINT; tests run by name (run.sh t4 t11) or
  INTEG_ONLY.
- mock_llm.py keeps the loop rewrite (raw/http/silent queues, finish
  overrides) and the agents additions (delay, chunk, per-request t).
- T12's .profile_override case drops the harness's endpoint/model env
  (a stale override now loses to env) and requires the gate's refusal
  notice, so it still exercises the dial-time check.
- the subagent loop refuses cut-off calls AND stops on its own guardrail.
- t_args_malformed_despite_lenient_decode no longer assumes json_decode
  accepts truncated input (current swarmrt rejects it).

make check: unit 241/241, smoke, integration 42/42, module 8/8.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
… one error

- browser_screenshot wrote PNG bytes wherever the model pointed it,
  around PathGuard; it now refuses the same paths write/edit do
  (credential dirs, swarm-code's control files). Test:
  t_browser_screenshot_path_guard (fails on the previous browser.sw).
- /schedule printed the scheduler's specific reason and then a generic
  "invalid EXPR" line; it uses Scheduler.add_checked and prints one.

make check: unit 242/242, smoke, integration 42/42, module 8/8.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
A repo's ./.swarm-code.json only applies its safe keys until its
directory is in "trusted_projects"; opting in meant hand-editing
~/.swarm-code/settings.json. `swarm-code trust [DIR]` now does it: the
directory is resolved to an absolute path and added once, every other
setting is kept, the file is written atomically and stays indented
(Util.json_pretty), and an unparseable settings.json is never
overwritten. It prints what the repo's file will now apply (hooks,
endpoint, MCP servers...). `untrust` removes it, `trust --list` shows
the list, /trust and /untrust do the same from the REPL (effective next
launch). The ignored-settings notice now names the command.

Tests: t_trust_edits_settings, t_trust_refuses_corrupt_settings,
t_json_pretty_round_trips, integration T29 (untrusted hook doesn't run
and the notice names the command; after `swarm-code trust` it runs).
make check: unit 245/245, smoke, integration 43/43, module 8/8.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
@coderabbitai

coderabbitai Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 4bc07eae-dd2b-4384-8696-0a2cd46e6110


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…EMOTE

`swarm-code doctor` sent GET /v1/models, with the Authorization header,
to whatever endpoint was configured, remote or not. It now applies the
same Config.endpoint_refusal gate as the LLM layer and prints "skipped"
for a refused endpoint; local endpoints and opted-in remote ones are
still probed.

SECURITY.md: protected paths are enforced by the file tools, not the
shell — say so under known limitations.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
@skyblanket
skyblanket merged commit a275556 into main Sep 24, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants