Repository navigation
Harness review: ~50 reproduced bugs fixed, security hardening, swarm-code trust - #1
Merged
Merged
Conversation
…rrectness - Headless stdout carries only the result: the --json line, or the final answer when stdout is piped; the transcript (tool calls, streamed tokens) goes to stderr. `swarm -p --json ... | jq` used to fail on every run. Uses swarmrt's new stdout_to_stderr/fd_write. - The working directory comes from /bin/sh, not a possibly stale $PWD: a launcher that chdir'd without updating PWD made the system prompt name another directory. - ESC/Ctrl-C on a tool ends the turn: the rest of the batch is recorded as skipped and the model is not called again. - read: .js/.ts files are text (libmagic calls them application/javascript and they were refused as binary); binary now means a NUL byte in the first 8KB, which needs no `file` binary; a missing file says "file not found"; an empty file says so; a trailing newline no longer adds a phantom last line. - edit/multi_edit: an LF old_string matches a CRLF file and the edit keeps CRLF endings. - bg_kill test counts live processes only (container inits don't reap zombies). make check: unit 166/166, smoke, integration 10/10 (T9 clean stdout, T10 stale PWD are new and fail on the previous binary), module 8/8. Requires swarmrt with eprint/stdout_to_stderr/fd_write. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
bash wrapped the model's command as `( export …; CMD ) </dev/null 2>&1`. A trailing `# comment` commented out the closing paren and a final heredoc terminator became `EOF ) </dev/null…`, so `echo hi # say hi` or a `cat > f <<'EOF' … EOF` write returned `[exit 2]` with empty output — and sh's syntax error went to the agent's stderr, never to the model. Util.noninteractive_wrap now puts `exec </dev/null 2>&1; export CI=1 …` on line 1 (run before sh parses the user's lines, so a syntax error is captured) and the command on its own lines, newline-terminated. The same wrapped script is used for the auto-background path and the `background` / `bg_server` tools (Background.launch_cmd records the raw command for display), so a command behaves the same however it runs. Util.no_stdin is the stdin-only variant for the harness's own helper pipelines. Regression tests (fail before, pass after): t_bash_trailing_comment, t_bash_heredoc_last, t_bash_syntax_error_reaches_model, t_bg_trailing_comment. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
Config.load merged ./.swarm-code.json over the user's settings with full authority, so opening swarm-code inside a cloned repo let that repo run SessionStart/PreToolUse hooks, start mcpServers, point endpoint/api_key/ providers/profiles/fallback_profile at an attacker, and loosen permissions (an api_key even skipped the network gate). Project scope is now an allow-list (model, max_tokens, llm_timeout_ms, chat_template_kwargs, vision); `permissions` entries apply only when stricter than the user's effective decision (allow < ask < deny), and merge into the user's map instead of replacing it. Everything else is dropped and main prints one notice naming the ignored keys (stderr in --mcp-server mode). Opt-in: list the directory under "trusted_projects" in ~/.swarm-code/settings.json (exact match on pwd / pwd -P) to apply the file in full as before. load_one also rejects a non-object JSON top level instead of crashing map_merge. Tests: - integration T11 (fails before: "project SessionStart hook executed"; passes after, incl. tighten-only perms, user endpoint used, notice shown, and the trusted_projects opt-in restoring the hook) - unit t_project_scope_strips_untrusted, t_project_scope_trusted_applies, t_project_ignored_keys, t_project_trusted_dir_match (new API) - run.sh: run_swarm clears exported opt-in knobs and gains RUN_ENDPOINT and RUN_ENV for the security cases Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
grep/glob omitted the path when it was "." so rg would print clean relative paths — but rg with no path searches its STDIN when stdin isn't a tty. In --mcp-server mode that is the JSON-RPC stream: grep blocked until the client closed it (30s timeout otherwise) and swallowed the next request. Both now always pass the root (`.` by default) and strip the `./` prefix with sed; anchored globs still match. An invalid regex (`foo(`) returned "(no matches)": the `2>&1` sat after `| head`, so rg's error went to the user's terminal and the exit code was lost. The search's stderr now goes to a private mkstemp file and is returned as `error: grep failed: <rg message>` (glob likewise). The grep fallback uses -E so its dialect matches rg's. Every other helper subprocess in tools.sw (read probes, git_*, code_search, log_wait, file_watch, sw_check, web_search, web_fetch) now runs through run_sh, which points stdin at /dev/null. Regression tests (fail before, pass after): integration T11 (MCP grep with no path + a second request), t_grep_invalid_regex_surfaces, t_grep_glob_default_path. run.sh gains INTEG_ONLY="t4 t11" to run a subset. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
The local-endpoint check in main.sw prefix-matched the raw string, so http://127.0.0.1@0.0.0.0:P and http://localhost:1@0.0.0.0:P (userinfo) and HTTP://0.0.0.0:P (uppercase scheme) reached a non-local host, and http://127.0.0.1.x.invalid, http://fd-anything.invalid and http://3221225985:9/ (decimal IPv4) passed as local. It only ran once at startup on the primary endpoint: SWARM_CODE_FALLBACK_ENDPOINT / fallback_profile, SWARM_CODE_PROVIDERS_JSON / providers and the ~/.swarm-code/.profile_override endpoint were never checked. And any api_key skipped the gate, contradicting README/SECURITY.md. Config.is_local_endpoint now parses the URL the way curl reads it: http(s) only (any case), any '@' in the authority refused, host lowercased with the port stripped and chars limited to [a-z0-9._-], IPv6 only bracketed (::1, fc00::/7, fe80::/10 with a full first group), a numeric / 0x last label means an IPv4 literal that must be a strict dotted quad (no octal, no short forms); names are localhost, *.local, *.ts.net or a bare dot-less name. main's startup gate uses it and no longer exempts an api_key — SWARM_CODE_ALLOW_REMOTE=1 is the only opt-in. llm.sw re-checks at the point of dial (stream_call for chat_native/chat_inband, chat_silent, chat_for_subagent; Plan.generate too); a refused URL is a fatal non-retried 403, so the fallback / next provider still gets its own gated attempt, and the reason goes to stderr (diag is silent headless). Tests: - integration T12 (fails before: "startup gate let http://127.0.0.1@0.0.0.0:P through (rc 0, 1 requests)"; after: the userinfo/uppercase/prefix/api-key cases exit 1 with no request, a non-local .profile_override and providers[0] are never dialed, and an ALLOW_REMOTE=1 control proves the host was reachable) - unit t_endpoint_gate_bypasses_refused, t_endpoint_gate_locals_allowed, t_endpoint_host_parse, t_endpoint_refusal_reason Behavior change: a remote endpoint with only an API key now needs SWARM_CODE_ALLOW_REMOTE=1, as documented; scheme-less endpoints ("host:port") are refused. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
A turn cut off at the output-token limit (finish_reason=length) still
dispatched its tool calls, and the runtime's lenient json_decode turned
arguments cut mid-string into a plausible call: a `write` of half a
config.py ending in 'postgres://prod' reported "ok: wrote 59 bytes". ESC
while tool-call arguments were streaming (sync stream path, stdout piped)
ran the truncated command the same way.
- run_turn: a truncated or user-interrupted turn answers every tool call
with a not-run result instead of dispatching it (each tool_call id still
gets its tool message). Truncated -> the F4 recovery (raised max_tokens,
then the smaller-edits nudge) asks the model to reissue the calls;
interrupted -> the turn ends like ESC on a running tool. Subagent loop
gets the same guard.
- llm.sw flags `interrupted` (runtime "[Request interrupted by user]"
marker) and `truncated` on the RAW content, before inband parsing cuts
the prose at the first call marker.
- Util.json_args_well_formed: strict structural check (balanced, matched
braces/brackets outside strings, terminated strings, only whitespace
after the top-level value), split on quotes so a 200KB write costs ms.
Used by execute_all's F5 guard, subagent_exec_all and
tcs_args_malformed, so a cut call is caught even without finish_reason.
- History stores "{}" for malformed arguments (sanitize_tool_calls):
servers that parse them (vLLM chat templates) would reject every later
request.
- Integration mock: raw-string arguments, "finish" override, HTTP error
responses, a "silent" queue for non-streaming requests, raw SSE lines;
run.sh takes an optional list of test names.
Tests: unit t_json_args_well_formed, t_args_malformed_despite_lenient_decode,
t_cut_turn_reason, t_cut_turn_calls_refused, t_sanitize_cut_tool_calls;
integration T11 (length-truncated write never runs), T12 (cut args without
finish_reason), T13 (interrupt-marked stream) — all three fail on the
previous binary. Real ESC mid-stream checked by hand through a pty.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
… reply matching
Three MCP-client bugs reproduced by review:
1. A server that exits mid-call surfaced as "did not respond" and only
counted as a timeout strike (status stayed "ok"); the next call then
wrote into the dead pipe. subprocess_recv_line returns nil EARLY only
on EOF, so mcp_read_reply now classifies an early nil as 'lost', and
mcp_note_result marks the server failed at once so the lazy-reconnect
path runs.
2. A server->client request whose id collides with our in-flight call
(e.g. {"id":100,"method":"roots/list"}) was taken as the response
("MCP response carried no result"). A message carrying `method` is a
request/notification, never our reply (mcp_msg_kind). Server `ping`
is answered with {} per spec, every other server request with
-32601, and we keep reading for the real response.
3. Health checks substring-matched the tool's OUTPUT: a successful
result containing "connection lost" forced a reconnect (state lost)
and then "not running" for 60s; "did not respond" in output was a
timeout strike. mcp_do_call now returns {'ok'|'error'|'timeout'|
'lost', text} and bookkeeping keys on the status only.
Tests: unit t_mcp_msg_kind_collision, t_mcp_health_structured;
integration A1 (EOF -> lost + reconnect), A2 (colliding id + ping),
A3 (result text never touches health; --json stdout stays one line),
driven by the new tests/integration/fake_mcp.py. A1-A3 fail on the
pre-fix binary (A1 dies with SIGPIPE, exit 141).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…itively validate_write only protected ~/.swarm-code/settings.json, so with bash denied the model could still `write` ~/.swarm-code/hooks/pre_tool.sh — which Hooks.run_pre_tool executes via shell() on the very next tool call — or schedule.json (recurring headless runs), .profile_override (LLM endpoint/key redirect) or a sessions/ journal (replayed as history). All checks compared case-sensitively, so ~/.SSH/ (the same directory on a case-insensitive macOS filesystem) was writable. Writes are now refused anywhere under a .swarm-code/ directory except its memory/ and skills/ data dirs (the only state the agent manages; the remember / learn_skill tools write there directly, and nothing else routes a legitimate write through write/edit), plus the directory itself and any project-scope .swarm-code.json. An unresolved ".." (realpath -m is missing on older macOS) can't climb out of a data dir. Every PathGuard comparison (write, read, is_sensitive) runs on the lowercased resolved path. SWARM_CODE_UNSAFE_WRITES=1 still lifts the write guards. Tests: - integration T13 (fails before: "BREACH: a model-written pre_tool hook executed"; after: hooks/, schedule.json, .profile_override (via edit) and ~/.SSH writes are refused, memory/ still writable) - unit t_pathguard_control_files_blocked, t_pathguard_data_dirs_writable, t_pathguard_case_insensitive (the first and last fail before) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…ation
`swarm --mcp-server` deviated from the MCP / JSON-RPC 2.0 specs:
* ping returned -32601; spec requires an empty result {}.
* tools/call with an unknown tool name returned -32601 (method not
found); spec: -32602 Invalid params. Non-object params/arguments
and non-string names are -32602 too.
* "jsonrpc" was never checked; anything but exactly "2.0" is now
-32600 Invalid Request.
* object / array / boolean / null ids were accepted and echoed back;
ids must be a string or number -> -32600 with id null. Presence of
`id` is checked by key (map_has_key is false for a null value), so
`"id": null` is no longer mistaken for a notification.
* an id-bearing "notifications/initialized" got no reply at all; it
is now acknowledged (every request is owed a response).
Tests: unit t_mcp_server_spec_envelope; integration A4 drives the real
binary over stdio and validates every response with python; smoke.sh
gains a ping -> {} check. All fail on the pre-fix binary.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
resolve_permission turned every 'ask' into 'allow' when headless, so
{"permissions":{"bash":"ask"}} + `swarm-code -p` ran `rm -rf ~/victim`
unprompted, the dangerous-bash gate was a no-op outside /flows (which
sets SWARM_CODE_DENY_DANGEROUS=1), and MCP tools (default 'ask', being
external and unvetted) auto-ran.
An 'ask' only arises for a dangerous bash command, a user-configured
"ask", or an MCP tool — default-allowed tools never reach it — so
headless now DENIES it unless the user exported
SWARM_CODE_HEADLESS_APPROVE=1 (restoring the old auto-approval; /flows
children still hard-deny dangerous commands via DENY_DANGEROUS). The
model-facing denial (permission_denial) names the opt-in, or the
per-tool "allow" setting when that would help; hardline / configured
denies keep the plain message. Documented in README, SECURITY.md and
--help; stale comments in config.sw / Flows.sw updated.
Tests:
- integration T14 (fails before: "headless auto-approved an explicit
\"ask\" permission"; after: the ask and a dangerous rm -rf ~ are
denied with the opt-in named, and HEADLESS_APPROVE=1 restores it)
- unit t_headless_ask_denied (fails before), and
t_headless_default_allowed_still_run
Behavior change: headless runs that relied on auto-approving an "ask"
tool, an MCP tool, or a dangerous command must now set
SWARM_CODE_HEADLESS_APPROVE=1 (or "allow" that tool in settings.json).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
Resume is the headless default, and run_headless took the last assistant
message ANYWHERE in the (resumed) history as the result. A run whose LLM
call failed — HTTP 400, server down — therefore printed the previous
run's reply as {"status":"ok","summary":"FIRST_RUN_ANSWER"} and exited 0.
run_headless now gives the run a uuid; run_turn stamps the assistant
messages it appends with it (in memory only — the journal drops the
field, and compaction can't shift it the way a length/index comparison
would). The result is the history's FINAL message when it is a reply
stamped by this run with no pending tool calls; otherwise status error /
exit 1 (LLM failure, max steps, nothing appended). A slash-command prompt
(/compact, /profile …) counts as success with an empty summary instead of
reporting a stale answer.
Tests: unit t_headless_answer_this_run_only; integration T14 (run 1 ok,
run 2 same HOME against an HTTP 400 and run 3 with no server must both be
status error / exit 1) — fails on the previous binary with
{"status":"ok","summary":"FIRST_RUN_ANSWER_T14"}.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
chat_inband wrote every request body — the whole conversation, prompts, tool output and any secrets in them — to /tmp/swarm-code-last-body.json: a fixed, shared path created 0644 (world-readable, and pre-creatable by another local user). chat_native wrote the same body to ~/.swarm-code/last-body.json, also 0644, on every call. Nothing in the code reads either file back. Both now call dump_last_body, which writes only when SWARM_CODE_DEBUG=1 (the existing debug knob), to ~/.swarm-code/last-body.json via file_temp (mkstemp, 0600) + rename, so the body is never world-readable even briefly. Noted in SECURITY.md. Tests: - integration T15 (fails before: "inband: request body written to /tmp/swarm-code-last-body.json"; after: neither inband nor native writes a dump by default, and SWARM_CODE_DEBUG=1 writes ~/.swarm-code/last-body.json with mode 0600) Behavior change: ~/.swarm-code/last-body.json is no longer refreshed on every native call; set SWARM_CODE_DEBUG=1 to get it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
file_watch spliced the model's path into `p="…"` via shell_inner_quote, which only escaped `"` — `$(touch PWNED)` or backticks in the path ran, even with bash denied. It now uses Util.shell_q, and shell_inner_quote (its only user) is removed. It also used BSD-only `stat -f %m`. On GNU, `-f` is --file-system, so it printed free-space counters that change constantly and reported "ok: changed" in 0.5s for an untouched file. The poll loop now detects the stat flavour once and compares an mtime+size signature (GNU `%y %s`, BSD `%m %z`), so two writes in the same second still register. log_wait/file_watch `timeout_sec` is clamped to [1, 600] (default 60, junk → 60): 0 used to reach shell_managed as "no timeout" (600s headless, unbounded on a TTY). Schema text now says 60, not 30. Regression tests (fail before, pass after): t_file_watch_no_injection, t_file_watch_portable_mtime, t_wait_timeout_clamped. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
CHANGELOG [Unreleased] gains a Security section covering the five fixes and their user-visible changes (API key no longer lifts the network gate, headless denies 'ask' without SWARM_CODE_HEADLESS_APPROVE=1, untrusted project config, protected control files, debug-only body dump). `swarm doctor` now warns when the resolved endpoint would be refused at startup, naming SWARM_CODE_ALLOW_REMOTE=1 — the most likely surprise for users who relied on the API-key exemption. No new test: doctor output is advisory (smoke already runs doctor). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…dless asks Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…ous bash, strict exprs
Four scheduler bugs reproduced by review:
5. A wrong-shape schedule.json crashed every interactive session ~2s
after launch, inside main's heartbeat handler: {"jobs":[]} → hd() on
a map in prune_jobs_loop; "last_run":"yesterday" → "str" + int in
compute_next_fire. read_state_at now requires a JSON array;
normalize_job validates id / expr / prompt and coerces numeric
fields (numeric strings, floats); invalid entries are skipped (left
untouched on disk) with ONE warning per distinct problem set,
printed with print_above so the pinned prompt survives.
6. /schedule on a corrupt schedule.json treated it as [] and replaced
every job. The runtime json_decode guesses through damage ("[{…} ,,,
oops" decodes to a nil-padded list), so corruption is now decided by
the new strict RFC 8259 checker JsonCheck.valid; add refuses to write
with a clear message, and appends to the RAW entries so entries it
can't validate survive. All writes use file_atomic_write.
7. Dispatched jobs ran headless children that auto-approve 'ask', so a
job whose model ran `rm -rf ~/victim` deleted it. dispatch_cmd now
prefixes SWARM_CODE_DENY_DANGEROUS=1, same as /flows.
8. Expression parsing read leading digits and ignored the rest: 1.5h
ran hourly, 10x5m every 10m, "daily :" at 00:00, "daily 9:5x" at
09:05. parse_interval / daily_time_ms are strict (expr_error gives
the reason, printed by add); `hourly` now fires at the top of the
hour (UTC) as documented, 1h stays relative. README example fixed
to the syntax /schedule actually accepts.
Note: agent.sw's /schedule handler (not in this change's scope) still
prints its generic "invalid EXPR" hint after the specific reason;
switching it to Scheduler.add_checked would drop that second line.
Tests: unit t_json_check_strict, t_sched_wrong_shape_never_panics,
t_sched_corrupt_refuses_write, t_sched_strict_exprs,
t_sched_hourly_on_the_hour, t_sched_dispatch_denies_dangerous;
integration A5 (session survives + stays responsive, one warning),
A6 (corrupt file byte-identical after /schedule), A7 (due job's child
is denied `rm -rf ~/victim`) — interactive, via the new
tests/integration/pty_run.py. A5-A7 fail on the pre-fix binary.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
The hardline and dangerous-command gates only looked at the `bash` tool
and matched raw substrings:
* `background`, `bg_server` and `run_tests.command` ran shell commands
with no gate at all (a hardline command via `background` executed).
* Trivially bypassed: `rm -r -f /`, `rm -fr /`, `rm -Rf /*`,
`rm -rf /*`, `rm -rf --no-preserve-root /`, `rm -rf "$HOME"`,
`:(){ :|:& };:`, `chmod -R 000 /`, `dd of=/dev/sda if=…`, `sudo<TAB>ls`.
* False positives with no override: `grep -r shutdown src/`,
`echo reboot required`, `git commit -m 'halt the build'`, or a heredoc
writing `def shutdown(` were hard-denied as a bare "permission denied".
New src/CommandGuard.sw parses the command like sh (quotes, escapes,
comments, separators, redirections, heredoc bodies as data, $(…)/`…`/<(…)
parsed recursively), then judges each simple command by its command word
and flags after unwrapping assignments, sudo/env/nohup/timeout/xargs/…,
`sh -c SCRIPT` and `eval`. Hardline: rm -r on / or /* or with
--no-preserve-root, mkfs*/mkswap, dd or redirection onto a raw disk,
shutdown/reboot/halt/poweroff/telinit/init 0|6/systemctl poweroff,
chmod/chown/chgrp -R on /, fork bombs. Dangerous (ask): sudo/doas,
rm -r on ~/$HOME, rm -rf on home/system paths, dd to other devices.
Config.check_permission runs it for every tool in Config.command_of; a
configured "deny" now stays deny for a dangerous command (it used to
become "ask", i.e. allowed headless). Denials name the matched pattern and
the offending simple command (Config.denial_message, used by
ToolExecutor and the agent loop). The tool-level sudo refusal is
token-aware and covers background/bg_server/run_tests too.
Regression tests (fail before, pass after): t_classifier_catches_bypasses,
t_classifier_no_false_positives, t_command_tools_gated,
t_denial_names_reason, t_sudo_tab_blocked, integration T12.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
context_budget_tokens was window - 16384 - 52000 with no clamp: below a 68,385-token window it went negative (SWARM_CODE_MAX_TOKENS=32768 gave -35,616), so every step printed "context at ~1197 / -35616 tokens, compacting" and ran the compaction path; llm.sw clamped only its own copy for the context meter (which then read x/1 tok). One definition now, LLM.context_budget_tokens (agent.sw delegates; llm.sw can't import agent.sw): reserve and buffer are each capped at window/4 (explicit env values too), so the budget is always >= window/2 with a floor of 1 — 8K->4096, 32K->16384, 128K->81920, 262K->193760 (unchanged). The context meter uses the same function. Tests: unit t_context_budget_scales_with_window; integration T15 (SWARM_CODE_MAX_TOKENS=32768: no "compacting" on a two-step turn, meter reads /16k) — fails on the previous binary at "-35616 tokens, compacting". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
compact_history kept "system + last 16" and summarized the rest, which: - with 10-17 messages summarized an EMPTY transcript and prepended one more summary per call (summaries stacked, history grew); - on a failed summarizer (503) still deleted the old messages, replacing them with "[compaction failed, messages elided]" — journaled; - mid-turn (a long tool loop), summarized the live user request away, so the following requests carried no user message at all. Now the verbatim tail is pulled back to start at or before the most recent user message and never on a tool result cut off from its assistant (compact_split / pair_start). Nothing old enough -> history unchanged with no LLM call; empty or failed summary -> history unchanged; an earlier summary is passed to the summarizer to fold in and replaced, not stacked. /compact only reports a count when something changed. Tests: unit t_compact_split_keeps_live_user, t_compact_split_pair_boundary, t_compact_nothing_old_is_noop, t_compact_failed_summary_keeps_history; integration T16 (/compact on seeded journals: 12 messages -> no summarizer call; an old summary is merged, one summary left; a 503 summarizer keeps all 30 messages) and T17 (12 tool rounds cross a 32K budget mid-turn: every request still carries the live request, only pre-turn history is summarized). T16 and T17 fail on the previous binary. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
Background task ids restart at bg-0 in every session and the task files
lived at fixed /tmp/swarm-code-bg-N.{log,pid,exit}. Two sessions (or two
worktrees running the test suite) collided: B's launch `rm -f`'d A's
files, B's pid file overwrote A's so `bg_kill bg-0` in A SIGTERMed B's
task, and a stale exit file could mark a fresh task done.
Each Background table now gets a private directory on first launch
(`mktemp -d` under $TMPDIR or /tmp, mode 0700, stored as 'dir' in the
table); log/pid/exit files live there. log_path_for / pid_file_for /
exit_file_for take the table; tools.sw (bash auto-bg, bg_server,
log_wait task_id) and Flows (task status + log stats) use them instead of
building /tmp paths. A launch that can't create the directory returns an
error, which background/bg_server now surface instead of formatting it
as a task id.
Regression test: t_bg_sessions_isolated (two tables, same bg-0 id:
separate 0700 dirs, A's kill leaves B's task running). On the old code
A's kill_task reported B's pid ("pid <B>: killed") and A kept running.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…d non-lossy results
Three subagent bugs reproduced by review (agent.sw subagent block only):
9. sub_opts inherited the parent's guardrails_table, so a subagent's 8
failing reads set the PARENT's halt_reason: the parent printed
[guardrail halt], never saw the result, and exited {"status":"error"}.
Each subagent now gets a fresh ToolGuardrails table (dropped after
use — ETS tables are finite); run_subagent_loop checks the
subagent's own halt after each tool batch and returns its partial
work plus the halt reason.
10. explore/bash restrictions were prompt-only (an explore subagent ran
bash and write). New ToolRegistry contexts subagent_explore
(read-only inspection: read/glob/grep/git_status/git_diff/
code_search) and subagent_bash (bash) are enforced twice: in
subagent_exec_all with a clear refusal naming what is allowed, and
by ToolExecutor via the subagent's execution_context. Unknown/odd
subagent_type values (null, 5) normalise to general.
11. Results were lossy/unbounded: max steps returned only a notice
(work discarded) — subagent_partial now returns the latest notes and
a digest of the most recent tool calls/results with every abnormal
stop; a 320KB answer reached the parent uncapped — capped at 24000
like bash/MCP (head+tail, marker, UTF-8-safe cuts); the LLM worker
was unlinked with a fixed 300s wait — it is now spawn_monitor'ed
(a crash surfaces its reason), bounded by Config.llm_timeout_ms
(SWARM_CODE_LLM_TIMEOUT_MS / settings), reports why it failed
(timeout / rejected / retries exhausted), and an abandoned worker is
killed so it cannot retry the hung request or hold headless exit.
Test infra: mock_llm.py gains "delay" (hung endpoint), "chunk" (many
small SSE deltas, as real servers send — the runtime cuts one huge
delta) and a per-request arrival time "t".
Tests: unit t_subagent_type_allowlists, t_subagent_result_capped,
t_subagent_partial_keeps_work; integration A8 (subagent streak never
halts the parent), A9 (explore can't bash/write), A10 (max steps keeps
FINDING_*), A11 (320KB answer capped), A12 (hung endpoint → reason
after the 3s LLM timeout, parent unblocked). A8-A12 fail on the
pre-fix binary.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…once
On a fatal request error run_turn dropped the turn back to its user
message (drop_to_last_clean_user) and a poisoned session was cut back the
same way on resume: two writes landed on disk, a 400 followed, and the
journal — and the model after resume — had no record of them.
- Fatal and transient failures now both keep every completed
assistant+tool pair and trim only the dangling tail (trim_incomplete).
The old reason to drop — re-sending cut-off tool arguments — is gone
(they are stored as "{}"). A poisoned session resumes the same way; the
poison flag now only changes the resume notice.
- A rejection that says the context is too long ("maximum context
length", "context_length_exceeded", "too many tokens", "prompt is too
long", llama.cpp's "exceeds the available context size") is
recoverable: llm.sw records the fatal message, run_turn shrinks the
history to half its size (mechanical trim first, then summarize) and
retries ONCE per turn — only if that actually made it smaller.
- Mechanical trim protects the most recent user message (mid-turn it is
no longer the last message and a pasted 10KB request got stubbed); the
overflow retry may also stub the last result, which the server already
refused.
Tests: unit t_context_overflow_detection, t_mech_trim_protects_live_user;
integration T18 (a: overflow after a 12KB result is trimmed and retried,
the turn succeeds; b: a 400 after two writes keeps both results in the
journal and a resumed run sends them; c: repeated overflow retries once).
T18 fails on the previous binary; its journal kept only the user message.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
Both hook systems spliced the args JSON into the shell command line and an
env var. A write/edit over ~128KB hit E2BIG: system() failed, shell() then
polled 120s for an exit file that never came, and
* a vetoing ~/.swarm-code/hooks/pre_tool.sh was SKIPPED (call allowed),
* a settings.json PreToolUse hook reported `block` after 120s even when
it would have allowed the call.
The post_tool/pre_llm/post_llm hooks stalled the same way.
Hooks.run_with_payload writes the payload to a private mkstemp file
(0600), makes it the hook's stdin, names it in $SWARM_HOOK_DATA_FILE /
$SWARM_CODE_ARGS_FILE, and still exports the inline env var (read from the
file inside the script, never the command line) when it is under 100KB
(else *_OMITTED=1). Hooks run under shell_managed (process-group kill on
timeout) instead of shell() + a perl alarm: 5s for filesystem hooks, 60s
for settings.json hooks.
Fail closed: if the payload file can't be created/written or the hook
can't be launched, pre_tool.sh vetoes and PreToolUse blocks, and the
error now says why (hook, exit code / timeout, start of its output) via
Config.run_hooks_verdict, instead of a bare "blocked by PreToolUse hook".
Regression tests: t_pre_tool_hook_big_payload,
t_configured_hook_big_payload, t_pre_tool_hook_fails_closed. On the old
code a probe showed veto=false after 120104ms and PreToolUse => block
after 120115ms for a hook that exits 0.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
12. `/flows` with a malformed workflow killed the interactive session:
{"phases":"oops"} reached init_phases, which hd()'d a string. The
workflow file is now checked with the strict JsonCheck.valid (the
runtime decoder happily "parses" a truncated file into a partial
workflow) and validate_workflow checks every shape the run walks —
object root, phases/tasks arrays, task objects with a non-empty
string prompt, string model/label — before the alt-screen opens or
anything launches. A bad file prints one error line and returns.
Fan-out was unbounded (12 tasks -> 12 children in 0.15s). A phase's
tasks are now queued and launched at most flows_max_parallel() at a
time (SWARM_CODE_FLOWS_MAX_PARALLEL, default 4); every render tick
refills free slots (fill_phase_slots / launch_quota). A task whose
Background.launch fails is finished as 'error' instead of carrying a
bogus id that read 'pending' forever and hung the flow.
Tests: unit t_flows_validate_shapes, t_flows_launch_quota; integration
A13 (three malformed workflows: error shown, session survives and
answers, nothing launched) and A14 (6 tasks, cap 2: peak 2 in flight,
all 6 complete), interactive via pty_run.py. Both fail on the pre-fix
binary (A13: session dies; A14: 6 children at once).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
Config.matches treated the matcher as a raw substring in either direction, so "edit|write" never fired for multi_edit and a "bash" guard hook never saw background / bg_server / run_tests / file_watch — the other tools that run shell commands — and Claude-Code-style "Bash" or "Edit|Write" never matched at all. A matcher is now `*` or "|"/","-separated alternatives, compared case-insensitively (underscores ignored). An alternative matches a tool whose name contains it, or a tool in its family: bash/shell → bash, background, bg_server, run_tests, file_watch, log_wait; edit → edit, multi_edit; multiedit → multi_edit. The spurious reverse match (matcher "multi_edit" firing for "edit") is gone. Documented in config.sw and a new README "Hooks" section (which also covers the stdin / *_FILE payload contract from the previous commit). Regression test: t_hook_matcher_families (fails before, passes after). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
fix_json_unicode_escapes replaced u0026/u003c/u003e/u0027/u0022/u002f/ u003d/u005c with the characters they name, on native prose and on the raw inband content before tool-call parsing — a workaround for an old runtime bug that no longer exists (the runtime decodes \uXXXX). It only corrupted real text: in inband mode (the default for local endpoints) a write of JS source `js = "<p>";` landed as `js = "\<p\>";`, and prose that mentions < showed "\<". Removed with its three call sites (chat_native, chat_inband, chat_for_subagent); content is used verbatim. Tests: integration T19 — "<div>", a bare "<", a JSON-escaped < inside the arguments and a literal < in the file round-trip byte for byte, in the written file and in the prose, for native and inband. Fails on the previous binary (native prose "\<div\>", inband file "\<p\>"). run_swarm now lets RUN_ENV override the harness defaults and RUN_UNSET remove them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…N/YAML/PEM secrets
13. Trajectory export (and Log.event) ran Log.redact over the ENCODED
JSON line. The long-blob layer treated the `n` of a `\n` escape as
the start of a run, producing `\[REDACTED]` — an invalid JSON escape,
so exports were not valid JSONL. And several common secret shapes
leaked: AWS_SECRET_ACCESS_KEY=wJalr…/K7MDENG/… (the '/' path
exemption), PGPASSWORD=…, postgres://admin:pw@host, {"password": "…"}
/ "api_key": "…" (space after the colon), YAML `password: …`, PEM
lines containing '/'.
* Log.redact_value walks decoded values (maps/lists, keys kept) and
redacts each string; event() and Trajectory.export_* encode AFTER
redacting, so escapes are never touched.
* New layers in Log.redact: PEM blocks masked wholesale (BEGIN/END
lines kept; unterminated → to end), URL userinfo passwords
(scheme://user:[REDACTED]@host), and a case-insensitive key/value
scanner for identifiers ending in a secret word (password, passwd,
passphrase, secret, api_key/apikey/api-key, access_key,
private_key, authorization; token/_key/-key/credential(s) keep the
old >= 8 threshold) with optional quotes/whitespace around
= : => (not == / ::); YAML/header `:` values run to end of line for
the unambiguous words; bare true/false/null and already-masked
values are left alone.
* The blob layer treats a backslash AND the byte it escapes as a run
boundary; `;` and `)` end unquoted values.
Tests: unit t_trajectory_redacts_valid_jsonl (every exported line is
strict JSON via JsonCheck.valid and contains none of 10 planted
secrets) and t_log_redact_value_shapes (nested redaction, structure
kept, benign text like `max_token: 4096` / paths untouched). The export
logic fails on the pre-fix code (invalid `\[` escape, 3+ leaks).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
read loaded only `head -c 65000` of a larger file and then applied offset/limit, so `offset: 15000` on a 20,000-line file returned an empty string with no marker, and nothing past the first 64KB was reachable. Now the window is chosen by line number from the whole file: files within file_read's 1MB cap are split in-process (so UTF-8 is preserved); bigger files are windowed with `sed -n 'A,Bp;Bq' | head -c` (bounded, streaming — still no multi-GB slurp) plus a `wc -l` line count. A NUL byte past the first 8KB (file_read stops there) falls back to the sed path with a note. The OUTPUT is capped at 65000 bytes on a line boundary (a single giant line is cut and flagged) and every cut says where to continue: [output capped at 65000 bytes — lines 1-1639 of 20000 shown; continue with offset=1640] [lines 1-2000 of 20000 shown — continue with offset=2001] An offset past EOF is an explicit error naming the line count. Junk offset/limit values fall back to the defaults (the to_int builtin returned nil and leaked "nil" into the line numbers). Util.join_all does the O(n log n) join the old per-line `acc ++` did quadratically. Regression tests (fail before, pass after): t_read_offset_past_64k, t_read_huge_file_window. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
… them
edit/multi_edit load the file with file_read, which is C-string based with
a 1MB cap:
* a file with a NUL byte (`HEADER abc\0\1\2 tail`) was read up to the
NUL, edited, written back as `HDR abc` — silent truncation, "ok".
* a file over 1MB made file_read return nil, which edit took for a
MISSING file, so `old_string: ""` overwrote it with new_string.
load_for_edit compares the loaded length with the on-disk size
(file_stat) — the whole file, not just a prefix — and refuses both cases
with a clear error (the file is left untouched; use bash/python). A
directory is refused too. write's overwrite-diff capture skips a file
whose prior content doesn't match its size, so no bogus diff is rendered.
Regression test (fails before, passes after): t_edit_refuses_nul_and_huge.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
14. recall/forget spliced the model-supplied slug straight into
skills_dir()/<slug>/SKILL.md, so forget_skill {"slug":
"../../../work/proj"} deleted work/proj/SKILL.md and recall_skill
read it. Skills.valid_slug now requires ONE safe path component —
non-empty, <= 128 bytes, no '/', '\', '..', leading '.' or control
characters — and recall/forget return a clear error otherwise. save
(which slugifies to [a-z0-9_]) additionally refuses an empty slug,
which used to write skills_dir()//SKILL.md.
Test: unit t_skill_slug_traversal_blocked plants a SKILL.md outside
skills_dir, aims a ../-slug at it, and requires recall to not return it
and forget to leave it on disk (both succeeded pre-fix), plus
valid/invalid slug cases.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…indexing
15. reindex only checked whether a journal was PRESENT in meta. A
journal that another instance was still writing when this one booted
got indexed mid-session, marked done, and was never refreshed — its
later turns were unsearchable forever (unless it happened to be this
instance's own .active journal).
meta now records size + mtime (file_stat: an in-process stat(2) —
the old "a stat costs a 1s shell() poll" reason for skipping the
check no longer holds) and a journal is reindexed when either
differs. The sig is taken BEFORE reading, so a journal that grows
during ingestion is reindexed again next boot, never missed. Old
index.db files are migrated with ALTER TABLE (NULL sizes → one
reindex). Unchanged journals are still skipped. init_at / search_at
take an explicit directory (init / search unchanged for callers).
Test: unit t_session_search_reindexes_changed — index a journal,
append a turn, re-init: the appended term must be found (it was not
pre-fix) and nothing is duplicated on a further re-init.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
TestRunner ran `cd <repo_path> && <cmd> 2>&1` through shell():
* repo_path was unquoted — a path with a space failed with exit 2, and
`/tmp; touch PWNED` ran the injected command;
* shell() has no timeout or ESC, so a suite over 120s came back as
"Exit code: -1" with no output and was left running, orphaned;
* stdin was not explicitly /dev/null.
run_tests_timed now runs Util.noninteractive_wrap("cd '<repo>' || exit 2
<cmd>") under shell_managed (own process group, killed on timeout/ESC),
default 300s, overridable with the new `timeout_ms` arg (1s..600s). On
timeout the partial output is kept with a banner saying so. do_run_tests
also shows the output tail whenever the exit code is non-zero or the run
timed out, not only when the parser counted failures (a compile error
used to show just "Exit code: 2").
Regression test (fails before, passes after):
t_run_tests_quoting_timeout_output.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…he profile
The ~/.swarm-code/.profile_override file is global and persistent:
- A leftover override beat env vars: SWARM_CODE_MODEL=test sent the model
of a /profile run days earlier (and its endpoint beat
SWARM_CODE_ENDPOINT). Overrides are now stamped with a per-launch
session_id (main.sw): this session's own /profile or /model still beats
env, but one from an earlier session is only a default — an env var
that is set wins, per field. /profile shows fields the environment
shadows as "not applied".
- profile_to_override omitted chat_template_kwargs and apply_override
then cleared the launch value, so `/profile qwen2` dropped
enable_thinking:false. The profile's kwargs are written and applied; a
/model-only override leaves them alone.
- /model X overwrote the file with {model: X}, silently reverting a
/profile's endpoint/api_key while printing "unchanged". It now carries
forward whatever part of the override is in effect
(LLM.effective_override) and changes only the model.
- The system prompt is built once for the launch tool_format; after an
override to inband the request had no tools array while the prompt said
tools were in it. build_request_body now swaps the prompt's tool-use
sections (Prompts.tool_sections, pure text) to match the format each
request is sent in — overrides, fallback endpoints and provider chains
alike; the history's system message is untouched.
- "/profile clear" (documented alias) was swallowed by the /profile NAME
branch as a profile named "clear".
Tests: unit t_override_env_beats_stale_override,
t_profile_override_keeps_chat_template_kwargs,
t_model_override_keeps_active_profile, t_system_prompt_follows_wire_format;
integration T20 (settings profile; `/profile qwen2`, then a new run with
SWARM_CODE_MODEL/ENDPOINT set sends model "test" to the env endpoint with
enable_thinking:false; `/model other-model` keeps the kwargs) — fails on
the previous binary (the stale override's endpoint beat the env one).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…reate
t_skill_slug_traversal_blocked and t_session_search_reindexes_changed
emptied their mkstemp-derived directories but left the directories
themselves in /tmp on every run (file_delete is unlink(2); there is no
rmdir builtin). Remove them with exec_argv("rmdir", …).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…ema text Low-severity items from the tool-layer review: * `~` / `~/…` were never expanded (tool paths don't pass through a shell unquoted): `read ~/.bashrc` said "file not found" and `write ~/x` made a literal `./~/` directory. resolve_path_arg (read/write/edit/multi_edit) and the path args of glob/grep/code_search/log_wait/file_watch/run_tests now expand them to $HOME. write's overwrite-diff key stays the raw arg, which is what the agent looks it up by. * edit with old_string="" on a missing file didn't create parent directories (write does), so a new file in a new directory failed. * The 8-consecutive-failures guardrail only counted results starting "error:"; bash reports failure as "[exit N]", so it never fired for bash. A non-zero "[exit N]" or a "[timed out" banner now counts as a failure. * Schema text disagreed with the code: grep said "Returns file paths by default" but returns matching lines; glob claimed mtime ordering it doesn't do. (log_wait/file_watch "default 30" was fixed with the clamp.) Regression tests (fail before, pass after): t_edit_create_makes_parents, t_tilde_paths_expand, t_guardrail_counts_bash_exit_codes, t_schema_text_matches_code. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…action Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…with "-" get_print_arg treated ANY argument starting with "-" after -p as "read the prompt from stdin": - README's `swarm-code -p --json "list the test files"` read an empty stdin, sent 0 requests and exited 1; - a prompt that itself starts with "-" (`-p "-- summarize the diff"`, `-p "-x ..."`) was dropped the same way — Scheduler and Flows pass user prompts as the argument after -p, so such jobs silently did nothing. Now only an exact "-" means stdin; a KNOWN flag right after -p (--json, --no-resume, --profile NAME, …) is skipped and the first positional argument after -p is the prompt (stdin when there is none); anything else after -p is the prompt, whatever it starts with. Positional-profile detection skips that same argument, so `-p --json gemma` is a prompt, not a profile selection. Usage text mentions `-p -`. Tests: integration T21 — `-p --json "..."`, `--json -p "..."`, `-p "-- summarize the diff ..."`, `-p "-x ..."`, `-p - <stdin`, and `-p --json <stdin` each send exactly their prompt and exit 0. Fails on the previous binary (the README form exits 1 with no request). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…oks, background isolation Conflicts resolved against the security and agents merges: headless denials keep the HEADLESS_APPROVE guidance and add the classifier's matched pattern; the tools branch's integration T11/T12 are renamed T16/T17 (the security branch already used T11/T12) and every test runs through the INTEG_ONLY loop. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…ls merge) The merge resolution kept both appended test blocks but dropped the closing brace between them, so test_runner.sw did not parse. This is the tree the merged make check ran on (unit 224/224, integration 31/31). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
…on, 4xx recovery, profiles, -p parsing Conflicts resolved against the security, agents and tools merges: - the loop branch's integration T11-T21 are renumbered T18-T28; run_swarm takes RUN_ENV as an array or a string, plus RUN_UNSET / RUN_STDIN / RUN_ENDPOINT; tests run by name (run.sh t4 t11) or INTEG_ONLY. - mock_llm.py keeps the loop rewrite (raw/http/silent queues, finish overrides) and the agents additions (delay, chunk, per-request t). - T12's .profile_override case drops the harness's endpoint/model env (a stale override now loses to env) and requires the gate's refusal notice, so it still exercises the dial-time check. - the subagent loop refuses cut-off calls AND stops on its own guardrail. - t_args_malformed_despite_lenient_decode no longer assumes json_decode accepts truncated input (current swarmrt rejects it). make check: unit 241/241, smoke, integration 42/42, module 8/8. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
… one error - browser_screenshot wrote PNG bytes wherever the model pointed it, around PathGuard; it now refuses the same paths write/edit do (credential dirs, swarm-code's control files). Test: t_browser_screenshot_path_guard (fails on the previous browser.sw). - /schedule printed the scheduler's specific reason and then a generic "invalid EXPR" line; it uses Scheduler.add_checked and prints one. make check: unit 242/242, smoke, integration 42/42, module 8/8. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
A repo's ./.swarm-code.json only applies its safe keys until its directory is in "trusted_projects"; opting in meant hand-editing ~/.swarm-code/settings.json. `swarm-code trust [DIR]` now does it: the directory is resolved to an absolute path and added once, every other setting is kept, the file is written atomically and stays indented (Util.json_pretty), and an unparseable settings.json is never overwritten. It prints what the repo's file will now apply (hooks, endpoint, MCP servers...). `untrust` removes it, `trust --list` shows the list, /trust and /untrust do the same from the REPL (effective next launch). The ignored-settings notice now names the command. Tests: t_trust_edits_settings, t_trust_refuses_corrupt_settings, t_json_pretty_round_trips, integration T29 (untrusted hook doesn't run and the notice names the command; after `swarm-code trust` it runs). make check: unit 245/245, smoke, integration 43/43, module 8/8. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…EMOTE `swarm-code doctor` sent GET /v1/models, with the Authorization header, to whatever endpoint was configured, remote or not. It now applies the same Config.endpoint_refusal gate as the LLM layer and prints "skipped" for a refused endpoint; local endpoints and opted-in remote ones are still probed. SECURITY.md: protected paths are enforced by the file tools, not the shell — say so under known limitations. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A full review of the harness. Three review agents plus end-to-end runs against a scripted OpenAI-SSE mock found about 50 reproducible bugs, and this PR fixes them. Each fix has a regression test that fails on the old code.
Depends on skyblanket/swarmrt#5 (merged). CI builds against swarmrt
main, and this branch uses its new builtins (stdout_to_stderr,fd_write) and its runtime fixes.What changes
Security
./.swarm-code.jsonapplies only safe keys (model,max_tokens, …) and can only tighten permissions. Its hooks, MCP servers, endpoint/api_key/providers/profiles and any looser permissions are ignored, with a notice, until you runswarm-code trustin that directory.swarm-code trust [DIR],untrust,trust --list, and/trust//untrustin the REPL edittrusted_projectsin~/.swarm-code/settings.json. Other keys are kept, the file is written atomically and stays indented, and a file that doesn't parse is never overwritten.127.0.0.1@evil), uppercase schemes, prefix tricks, decimal IPs andfd…hostnames no longer pass. Every endpoint actually dialled is checked (primary, fallback, providers,/profileoverride).SWARM_CODE_ALLOW_REMOTE=1, as the README already said.schedule.json, settings, sessions);memory/andskills/stay writable. Path checks are case-insensitive, andbrowser_screenshotnow goes through the same guard.SWARM_CODE_HEADLESS_APPROVE=1. Scheduled jobs also deny dangerous commands.bash,background,bg_server,run_tests). It catches respellings likerm -fr /orsh -c '…'and stops flagging words that only appear inside quotes,echoorgreppatterns.file_watch. Shell injection is fixed. Hooks with large payloads fail closed. Background tasks get a private directory per session, so one session can no longer kill another's task.SWARM_CODE_DEBUG=1(mode 600).Tools
readworks on.js/.tsfiles (libmagic calls themapplication/javascript, so they were refused as binary). Binary now means a NUL byte in the first 8 KB, which doesn't depend onfilebeing installed.readreports a missing file as missing, reaches past the first 64 KB, and no longer shows a phantom last line.editworks on CRLF files. It refuses NUL-containing and >1 MB files instead of truncating them, and creates parent directories.bashsurvives a trailing# commentor heredoc, and shell syntax errors now reach the model.grepnever reads stdin (it swallowed the MCP server's next request) and reports regex errors.run_testsquotes its path and times out cleanly.~expands in paths.Agent loop
--jsonline, or just the answer when stdout is piped, soswarm -p --json … | jqworks.-pparses its prompt whatever the flag order, including prompts that start with-.<into\<is removed. Profile overrides no longer override your environment variables.$PWD.Multi-agent, MCP, scheduler
ping, unknown-tool error codes, and id/jsonrpcvalidation./flowsand scheduler: a malformedschedule.jsonor flow definition no longer crashes the session./flowsfan-out is capped, and schedule parsing is strict.Testing
make checkpasses: unit 245/245 (was 160), smoke, integration 43/43 (was 8), module checks 8/8. The key scenarios were also re-run by hand on the final binary:--json | jq;café 日本coming back intact from bash;Notes for reviewers
swarm-code trustin any repo whose.swarm-code.jsonyou rely on. Otherwise its hooks and endpoint are ignored, with a notice.SWARM_CODE_ALLOW_REMOTE=1.SWARM_CODE_HEADLESS_APPROVE=1.🤖 Generated with Claude Code
https://claude.ai/code/session_016jS65kn5WxDsbVU2teBZZE
Generated by Claude Code