diff --git a/CHANGELOG.md b/CHANGELOG.md index a96548b..93c123e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,32 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] +### Security + +- **A cloned repo's `./.swarm-code.json` is untrusted.** It may set only + `model`, `max_tokens`, `vision`, `chat_template_kwargs`, `llm_timeout_ms` + and *tighten* `permissions`; hooks, `mcpServers`, `endpoint`, `api_key`, + `providers`, `profiles` and `fallback_profile` in it are ignored with a + one-line notice. Opt a repo in with `swarm-code trust` (or `/trust`), which + adds it to `"trusted_projects"` in `~/.swarm-code/settings.json`; + `swarm-code untrust` and `swarm-code trust --list` manage the list. +- **Network gate parses URLs properly and covers every LLM dial.** Userinfo + (`http://127.0.0.1@host`), uppercase schemes, name-prefix and numeric-IP + tricks no longer pass, and fallback / `providers` / `/profile` override + endpoints are checked at the point of dial. **Changed:** an API key no + longer bypasses the gate — remote endpoints need `SWARM_CODE_ALLOW_REMOTE=1` + as documented; scheme-less endpoints (`host:port`) are refused. +- **The model can't write swarm-code's control files** (`~/.swarm-code/` + hooks, `schedule.json`, `.profile_override`, sessions, `settings.json`, + `.swarm-code.json`); `memory/` and `skills/` stay writable. Sensitive-path + checks are case-insensitive (`~/.SSH`). +- **Headless no longer auto-approves an 'ask'.** A dangerous command, an + explicit `"ask"` permission or an MCP tool is denied in `-p` / cron / + `/flows` runs unless `SWARM_CODE_HEADLESS_APPROVE=1` is set. +- **The raw request body is no longer written to + `/tmp/swarm-code-last-body.json`** (or refreshed in `~/.swarm-code/` on every + call); `SWARM_CODE_DEBUG=1` writes `~/.swarm-code/last-body.json`, mode 600. + ## [1.1.0] - 2026-07-15 The responsiveness release: the agent never blocks the terminal and the diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index ca022c8..bb34bdc 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -7,6 +7,7 @@ | `main.sw` | CLI flags, env config, headless mode, the main loop | | `agent.sw` | Prompt assembly, tool-call parsing/serialising, turn loop | | `ToolExecutor.sw` | Shared context, hook, guardrail, and permission boundary | +| `CommandGuard.sw` | Token-aware shell-command risk classifier (hardline / dangerous) used by the permission gate for every command-running tool | | `tools.sw` | Raw tool handlers (`do_bash`, `do_read`, …) and `exec_raw()` dispatch | | `ToolRegistry.sw` | Tool identity plus execution-context allow/block policy | | `ToolSchemas.sw` | OpenAI-compatible function schemas for native tool calling | diff --git a/README.md b/README.md index faf7ee1..bc14662 100644 --- a/README.md +++ b/README.md @@ -30,8 +30,8 @@ search: docker compose (3 hits) > recall_skill deploy-mally-otp # pull full playbook from ~/.swarm-code/skills/ SKILL.md loaded. Building burrito binary … -> /schedule add 1h "review open PRs" # heartbeat-driven cron -job 3 added (every 1h) +> /schedule "1h" "review open PRs" # heartbeat-driven cron (or "hourly", "daily 09:00") +✓ scheduled job 3 (1h): review open PRs ``` ## Quickstart @@ -78,7 +78,21 @@ Point it at any OpenAI-compatible endpoint via `~/.swarm-code/settings.json`. Pr } ``` -Remote endpoints are opt-in — set `SWARM_CODE_ALLOW_REMOTE=1` (local-network-only by default). Optional semantic memory recall uses `SWARM_CODE_EMBED_ENDPOINT`. +Remote endpoints are opt-in — set `SWARM_CODE_ALLOW_REMOTE=1` (local-network-only by default; an API key alone is not an opt-in). The check applies to every URL the LLM layer dials — primary, fallback, `providers`, and `/profile` switches. Optional semantic memory recall uses `SWARM_CODE_EMBED_ENDPOINT`. + +A repo-local `./.swarm-code.json` is **untrusted** (it ships with whatever you cloned): it may set `model`, `max_tokens`, `vision`, `chat_template_kwargs` and `llm_timeout_ms`, and may only *tighten* `permissions`. Its hooks, MCP servers, endpoints, API keys, providers and profiles are ignored, with a one-line notice. To let a repo you trust apply its file in full, run `swarm-code trust` in it (or `/trust` in the REPL; `swarm-code untrust` reverses it, `swarm-code trust --list` shows the list). That adds the directory to `"trusted_projects"` in `~/.swarm-code/settings.json`, which you can also edit by hand. + +### Hooks + +`settings.json` can run shell commands around tool calls; a `PreToolUse` hook that exits non-zero (or times out after 60s, or cannot be started) blocks the call, and the hook's output is shown to the model: + +```json +{ "hooks": { + "PreToolUse": [ { "matcher": "bash", "command": "./scripts/check-cmd.sh" } ], + "PostToolUse": [ { "matcher": "edit|write", "command": "make fmt" } ] } } +``` + +A matcher is `*` or `|`-separated, case-insensitive alternatives, each matching a tool whose name contains it — plus tool families: `bash` also fires for every other tool that runs a shell command (`background`, `bg_server`, `run_tests`, `file_watch`, `log_wait`), and `edit` also for `multi_edit`. The tool arguments arrive as JSON on the hook's **stdin** and in the private file `$SWARM_CODE_ARGS_FILE`; `$SWARM_CODE_ARGS` carries them inline only when under 100KB (else `$SWARM_CODE_ARGS_OMITTED=1`), so a hook that must see every call should read stdin. `$SWARM_CODE_EVENT` / `$SWARM_CODE_TOOL` name the event and tool. Executable scripts in `~/.swarm-code/hooks/` (`pre_tool.sh`, `post_tool.sh`, `pre_llm.sh`, `post_llm.sh`) get the same treatment via stdin / `$SWARM_HOOK_DATA_FILE` — see `src/Hooks.sw`. ## Features @@ -119,8 +133,11 @@ Panel agents run under the fail-closed `council_panel` context: they may inspect swarm-code runs shell commands, reads and writes files, and can reach the network — so it is built fail-closed: - **Local-network-only by default**; remote endpoints require an explicit `SWARM_CODE_ALLOW_REMOTE=1`. +- A cloned repo's `./.swarm-code.json` cannot run hooks, start MCP servers, redirect the endpoint/key, or loosen permissions unless you trust the directory (`swarm-code trust`). +- The `write`/`edit` tools refuse swarm-code's own control files (`~/.swarm-code/` hooks, schedule, settings, sessions, profile override; `.swarm-code.json`) — only `~/.swarm-code/memory/` and `skills/` are writable — and credential dirs (`.ssh`, `.aws`, `.gnupg`, …) case-insensitively. - Every tool runs through one **`ToolExecutor` policy boundary** — context allow-lists, argument-rewriting hooks, guardrails, and permissions — *before* any raw handler executes, and **fails closed** on a missing or unknown execution context. -- A **hardline command blocklist** (`rm -rf /`, `mkfs`, `dd`, fork bombs, …) cannot be bypassed by environment overrides. +- Headless runs (`-p`, cron jobs, `/flows` children) never auto-approve a call that needs permission — a dangerous command, an explicit `"ask"` setting, or an MCP tool — unless you set `SWARM_CODE_HEADLESS_APPROVE=1`. +- A **hardline command blocklist** (`rm -rf /`, `mkfs`, `dd` to a disk, halt/reboot, fork bombs, …) cannot be bypassed by environment overrides. It covers every tool that runs a shell command (`bash`, `background`, `bg_server`, `run_tests`), and commands are parsed like `sh` does — respellings such as `rm -fr /`, `dd of=/dev/sda if=…` or `sh -c '…'` are caught, while words inside quotes, `echo`/`grep` arguments or heredoc bodies are not flagged. Denials name the matched pattern. - Subagents, MCP, and council contexts run under restricted (often read-only) policies. - Secrets are redacted from session logs and trajectory exports. diff --git a/SECURITY.md b/SECURITY.md index 273335d..090a775 100644 --- a/SECURITY.md +++ b/SECURITY.md @@ -23,22 +23,53 @@ backported. ## Security model - **Network isolation by default.** Only local-network endpoints are allowed - unless you explicitly set `SWARM_CODE_ALLOW_REMOTE=1`. + unless you explicitly set `SWARM_CODE_ALLOW_REMOTE=1` (an API key is not an + opt-in). The check runs on every URL the LLM layer dials — primary, + fallback, `providers`, and the `/profile` override — and parses it the way + curl does: http(s) only, no `user@` credentials, IP literals only as strict + dotted quads or bracketed IPv6. +- **Untrusted project config.** A repository's `./.swarm-code.json` can only + set harmless keys (`model`, `max_tokens`, …) and *tighten* permissions. Its + hooks, MCP servers, endpoint / API key / providers / profiles, and any + permission loosening are ignored (with a notice) unless the directory is + listed under `trusted_projects` in `~/.swarm-code/settings.json` — + `swarm-code trust` (or `/trust`) adds the current directory for you. + Trust only repos you control: a trusted repo's hooks and MCP servers run + commands on your machine. - **Single policy boundary.** Every tool call — from the main agent, subagents, the council, and the MCP server — passes through `ToolExecutor`: context allow-lists, argument-rewriting hooks, guardrails, and permissions are applied *before* any raw handler runs. Execution **fails closed** when the execution context is missing or unknown. +- **No unattended approvals.** In headless mode (`-p`, scheduled jobs, `/flows` + children) there is nobody to answer a permission prompt, so a call that + needs one — a dangerous command, a tool you set to `"ask"`, an MCP tool — + is denied. `SWARM_CODE_HEADLESS_APPROVE=1` opts back into auto-approval + (`/flows` children still hard-deny dangerous commands). - **Hardline command blocklist.** Destructive commands (`rm -rf /`, `mkfs`, `dd` to a device, fork bombs, and similar) are blocked and **cannot be bypassed by - environment overrides**. + environment overrides**. The check covers every tool that runs a shell + command (`bash`, `background`, `bg_server`, `run_tests`) and parses the + command like `sh` (quotes, separators, `$(…)`, `sh -c`, wrappers such as + `sudo`/`env`/`xargs`), so respellings are caught; it is a floor against + accidents, not a sandbox — a command assembled at runtime can't be judged. +- **Protected paths.** The `write` / `edit` / `multi_edit` tools refuse + credential locations (`.ssh`, `.gnupg`, `.aws`, `/etc`, …) and swarm-code's + own control files — everything under `~/.swarm-code/` except the `memory/` + and `skills/` data dirs (hooks there run on every tool call), plus + `.swarm-code.json`. Matching is case-insensitive (macOS filesystems are). + `SWARM_CODE_UNSAFE_WRITES=1` lifts these guards. - **Restricted contexts.** Subagents, MCP-server, and council-panel contexts run under narrowed, often read-only, tool policies. - **Secret redaction.** Known secret patterns are redacted from session logs and - trajectory exports. + trajectory exports. The raw LLM request body is written to disk only with + `SWARM_CODE_DEBUG=1`, to `~/.swarm-code/last-body.json` with mode 600. ## Known limitations +- Protected paths are enforced by the file tools (`write`, `edit`, + `multi_edit`, `browser_screenshot`), not by the shell: a `bash` command can + still write anywhere your user can, so review the commands you approve. - The council panel's read-only isolation is **tool-level**, not yet a filesystem sandbox — panelists inherit the read tool's filesystem visibility. - swarm-code runs the commands you (or a model you configured) direct it to. Run diff --git a/src/CommandGuard.sw b/src/CommandGuard.sw new file mode 100644 index 0000000..d4df1ff --- /dev/null +++ b/src/CommandGuard.sw @@ -0,0 +1,813 @@ +module CommandGuard + +# ============================================================ +# CommandGuard — token-aware risk classifier for shell commands +# ============================================================ +# +# Config.check_permission asks this module how risky a model-supplied +# shell command is, for EVERY tool that runs one (bash, background, +# bg_server, run_tests.command — see Config.command_of). The old gate was +# raw substring matching on the bash tool only, which was wrong both ways: +# +# * trivially bypassed — `rm -r -f /`, `rm -fr /`, `rm -rf /*` (two +# spaces), `dd of=/dev/sda if=…`, `:(){ :|:& };:`, `sudols`, … +# * false positives with no override — `grep -r shutdown src/`, +# `echo reboot required`, `git commit -m 'halt the build'`, a heredoc +# writing `def shutdown(` were all hard-denied. +# +# So the command is parsed the way sh would split it, then each SIMPLE +# COMMAND is judged by its command word and flags: +# +# parse — a small sh lexer: quotes ('…', "…", $'…'), backslash escapes, +# `# comments`, separators (; & && | || newline ( ) ), redirections +# (targets kept aside), heredoc bodies (skipped as data; an +# unquoted body's $(…) is still parsed), and command substitution +# ($(…), `…`, <(…), >(…)) — whose inner commands are parsed too. +# judge — strip leading VAR=val assignments and reserved words, unwrap +# sudo/env/nohup/timeout/xargs/… , recurse into `sh -c SCRIPT` and +# `eval`, then match the command word: rm (recursive/force flags in +# any spelling + target), dd (of=), mkfs*/mkswap, shutdown/reboot/ +# halt/poweroff/telinit/init 0|6/systemctl poweroff, chmod/chown/ +# chgrp -R on /, sudo/doas. Redirections onto a raw disk and fork +# bombs are checked on the whole command. +# +# This is a safety FLOOR against accidents, not a sandbox: a command built +# at runtime (`$cmd`, `x=re; ${x}boot`) can't be judged statically. +# +# Findings are {level, code, reason} with level 'hardline' (never allowed) +# or 'dangerous' (ask; denied when SWARM_CODE_DENY_DANGEROUS=1). The reason +# names the pattern AND the offending simple command, and is surfaced in the +# denial message so the model can adapt instead of retrying blindly. + +export [classify, risk_of, uses_sudo, commands_of] + +# ------------------------------------------------------------ +# Public API +# ------------------------------------------------------------ + +# Most severe finding: {'hardline', reason} | {'dangerous', reason} | {'ok', ""}. +fun risk_of(cmd) { + if (cmd == nil) { {'ok', ""} } + else { worst(classify(to_string(cmd)), {'ok', ""}) } +} + +# Every finding for the command (list of {level, code, reason}). +fun classify(cmd) { + classify_text(to_string(cmd), 0) +} + +# 'true' when any simple command (through wrappers, `sh -c`, `$(…)`) runs sudo/doas. +fun uses_sudo(cmd) { + has_code(classify(to_string(cmd)), 'sudo') +} + +# The parsed simple commands (list of word lists) — for tests and debugging. +fun commands_of(cmd) { + map_get(parse(to_string(cmd)), 'cs') +} + +fun worst(findings, best) { + if (length(findings) == 0) { best } + else { + f = hd(findings) + lvl = elem(f, 0) + if (lvl == 'hardline') { {'hardline', elem(f, 2)} } + else { + next = if (elem(best, 0) == 'ok') { {'dangerous', elem(f, 2)} } else { best } + worst(tl(findings), next) + } + } +} + +fun has_code(findings, code) { + if (length(findings) == 0) { 'false' } + else { if (elem(hd(findings), 1) == code) { 'true' } else { has_code(tl(findings), code) } } +} + +# Nesting guard for `sh -c "sh -c '…'"`, wrapper chains and $(…) inside $(…). +fun max_depth() { 8 } + +fun classify_text(s, depth) { + if (depth > max_depth()) { [] } + else { + st = parse(s) + judge_all(map_get(st, 'cs'), depth, []) ++ + redirect_findings(map_get(st, 'rt'), []) ++ + fork_bomb_findings(s) + } +} + +# ============================================================ +# Lexer — sh command string → simple commands +# ============================================================ +# State (a map threaded through self-tail-recursive loops): +# q quote mode: 0 none, 1 '…', 2 "…", 3 $'…', 4 unquoted heredoc body +# esc 'true' after a backslash +# wl chars of the word being built (in order) hw word in progress +# wq the word contained quoting (a quoted heredoc delimiter → literal body) +# ws words of the current simple command cs finished commands +# rd what the next word is: 'none' | 'out' | 'in' | 'hplain' | 'hdash' +# rt output-redirect targets (all commands) +# hd pending heredocs [%{d, dash, quoted}] body 'true' while skipping bodies +# cm capture mode nil | 'paren' | 'tick' for $(…) / `…` / <(…) +# cl captured chars cd paren depth cq capture quote cesc cret (q to resume) + +fun init_state() { + %{q: 0, esc: 'false', wl: [], hw: 'false', wq: 'false', ws: [], cs: [], rd: 'none', + rt: [], hd: [], body: 'false', cm: nil, cl: [], cd: 0, cq: 0, cesc: 'false', cret: 0} +} + +fun parse(s) { + finish_input(lex_lines(string_split(s, "\n"), init_state())) +} + +fun lex_lines(lines, st) { + if (length(lines) == 0) { st } + else { + line = hd(lines) + rest = tl(lines) + if (map_get(st, 'cm') == nil && map_get(st, 'body') == 'true') { + h = hd(map_get(st, 'hd')) + if (is_delim_line(line, h) == 'true') { + left = tl(map_get(st, 'hd')) + st2 = map_put(map_put(st, 'hd', left), 'body', if (length(left) > 0) { 'true' } else { 'false' }) + lex_lines(rest, st2) + } else { if (map_get(h, 'quoted') == 'true') { + lex_lines(rest, st) + } else { + # Unquoted body: data, but $(…) / `…` in it still execute. + lex_lines(rest, end_line(lex_chars(string_chars(line), map_put(st, 'q', 4)))) + }} + } else { + lex_lines(rest, end_line(lex_chars(string_chars(line), st))) + } + } +} + +fun is_delim_line(line, h) { + l0 = if (string_ends_with(line, "\r") == 'true') { string_sub(line, 0, string_length(line) - 1) } else { line } + l = if (map_get(h, 'dash') == 'true') { strip_leading_tabs(l0) } else { l0 } + if (l == map_get(h, 'd')) { 'true' } else { 'false' } +} + +fun strip_leading_tabs(s) { + if (string_starts_with(s, "\t") == 'true') { strip_leading_tabs(string_sub(s, 1, string_length(s) - 1)) } + else { s } +} + +# End of input: an unterminated $(…) is still judged; an open word/command closes. +fun finish_input(st) { + if (map_get(st, 'cm') != nil) { finish_cmd(end_capture(st)) } + else { finish_cmd(st) } +} + +# What a newline means depends on where we are. +fun end_line(st) { + q = map_get(st, 'q') + if (map_get(st, 'cm') != nil) { map_put(st, 'cl', list_append(map_get(st, 'cl'), "\n")) } + else { if (q == 1 || q == 2 || q == 3) { add_char(st, "\n") } + else { if (q == 4) { map_put(map_put(st, 'q', 0), 'esc', 'false') } + else { if (map_get(st, 'esc') == 'true') { map_put(st, 'esc', 'false') } # line continuation + else { + st2 = finish_cmd(st) + if (length(map_get(st2, 'hd')) > 0) { map_put(st2, 'body', 'true') } else { st2 } + }}}} +} + +fun lex_chars(chars, st) { + if (length(chars) == 0) { st } + else { if (map_get(st, 'cm') == nil && map_get(st, 'esc') == 'false') { + # Fast path: take a whole run of ordinary chars in one go (one state + # update per run, not per char) — a 75KB `python -c '…'` word + # otherwise costs a map update per byte. + q = map_get(st, 'q') + wl = map_get(st, 'wl') + n0 = length(wl) + # The run is appended straight onto the word (list_append is O(1) + # amortized; `wl ++ run` would re-copy the word on every line). + r = scan_plain(chars, q, if (q == 4) { [] } else { wl }) + grown = elem(r, 1) + st2 = if (q == 4 || length(grown) == n0) { st } + else { map_put(map_put(st, 'wl', grown), 'hw', 'true') } + rest = elem(r, 0) + if (length(rest) == 0) { st2 } + else { + r2 = step(hd(rest), tl(rest), st2) + lex_chars(elem(r2, 0), elem(r2, 1)) + } + } else { + r = step(hd(chars), tl(chars), st) + lex_chars(elem(r, 0), elem(r, 1)) + }} +} + +# Leading run of chars that are literal in quote mode q → {rest, run}. +fun scan_plain(chars, q, acc) { + if (length(chars) == 0) { {chars, acc} } + else { + c = hd(chars) + if (is_special(q, c) == 'true') { {chars, acc} } + else { scan_plain(tl(chars), q, list_append(acc, c)) } + } +} + +fun is_special(q, c) { + if (q == 1) { if (c == "'") { 'true' } else { 'false' } } + else { if (q == 2) { in_list(["\"", "\\", "$", "`"], c) } + else { if (q == 3) { if (c == "'" || c == "\\") { 'true' } else { 'false' } } + else { if (q == 4) { in_list(["\\", "$", "`"], c) } + else { in_list([" ", "\t", "\\", "'", "\"", "`", "$", "#", ";", "|", "(", ")", "&", ">", "<"], c) }}}} +} + +fun peek(rest) { if (length(rest) == 0) { "" } else { hd(rest) } } + +# One character → {remaining chars, new state}. +fun step(c, rest, st) { + if (map_get(st, 'cm') != nil) { cap_step(c, rest, st) } + else { + q = map_get(st, 'q') + if (map_get(st, 'esc') == 'true') { + st2 = map_put(st, 'esc', 'false') + if (q == 4) { {rest, st2} } + else { if (q == 3) { {rest, add_char(add_char(st2, "\\"), c)} } + else { {rest, add_char(st2, c)} } } + } else { if (q == 1) { + if (c == "'") { {rest, map_put(st, 'q', 0)} } else { {rest, add_char(st, c)} } + } else { if (q == 3) { + if (c == "\\") { {rest, map_put(st, 'esc', 'true')} } + else { if (c == "'") { {rest, map_put(st, 'q', 0)} } else { {rest, add_char(st, c)} } } + } else { if (q == 2) { + if (c == "\\") { {rest, map_put(st, 'esc', 'true')} } + else { if (c == "\"") { {rest, map_put(st, 'q', 0)} } + else { if (c == "`") { {rest, start_capture(st, 'tick', 2)} } + else { if (c == "$" && peek(rest) == "(") { {tl(rest), start_capture(st, 'paren', 2)} } + else { {rest, add_char(st, c)} } } } } + } else { if (q == 4) { + if (c == "\\") { {rest, map_put(st, 'esc', 'true')} } + else { if (c == "`") { {rest, start_capture(st, 'tick', 4)} } + else { if (c == "$" && peek(rest) == "(") { {tl(rest), start_capture(st, 'paren', 4)} } + else { {rest, st} } } } + } else { + step_unquoted(c, rest, st) + }}}}} + } +} + +fun step_unquoted(c, rest, st) { + nxt = peek(rest) + if (c == " " || c == "\t") { {rest, finish_word(st)} } + else { if (c == "\\") { {rest, map_put(mark_quoted(st), 'esc', 'true')} } + else { if (c == "'") { {rest, map_put(mark_quoted(st), 'q', 1)} } + else { if (c == "\"") { {rest, map_put(mark_quoted(st), 'q', 2)} } + else { if (c == "`") { {rest, start_capture(st, 'tick', 0)} } + else { if (c == "$" && nxt == "(") { {tl(rest), start_capture(st, 'paren', 0)} } + else { if (c == "$" && nxt == "'") { {tl(rest), map_put(mark_quoted(st), 'q', 3)} } + else { if (c == "#" && map_get(st, 'hw') == 'false') { {[], st} } # comment to end of line + else { if (c == ";" || c == "|" || c == "(" || c == ")") { {rest, finish_cmd(st)} } + else { if (c == "&") { + # `&>file` is a redirection; `&`, `&&`, `|&` end the command. + if (nxt == ">") { {rest, finish_word(st)} } else { {rest, finish_cmd(st)} } + } + else { if (c == ">" || c == "<") { redirect_step(c, rest, st) } + else { {rest, add_char(st, c)} }}}}}}}}}}} +} + +# `>`, `>>`, `>&`, `>|`, `<`, `<&`, `<>`, `<<`, `<<-`, `<<<`, `<(…)`, `>(…)`. +fun redirect_step(c, rest, st) { + st1 = finish_word(st) + nxt = peek(rest) + if (nxt == "(") { {tl(rest), start_capture(st1, 'paren', 0)} } # process substitution + else { if (c == ">") { + rest2 = if (nxt == ">" || nxt == "&" || nxt == "|") { tl(rest) } else { rest } + {rest2, map_put(st1, 'rd', 'out')} + } else { if (nxt == "<") { + r2 = tl(rest) + n2 = peek(r2) + if (n2 == "<") { {tl(r2), map_put(st1, 'rd', 'in')} } # here-string + else { if (n2 == "-") { {tl(r2), map_put(st1, 'rd', 'hdash')} } + else { {r2, map_put(st1, 'rd', 'hplain')} } } + } else { + rest2 = if (nxt == "&" || nxt == ">") { tl(rest) } else { rest } + {rest2, map_put(st1, 'rd', 'in')} + }}} +} + +fun add_char(st, c) { + map_put(map_put(st, 'wl', list_append(map_get(st, 'wl'), c)), 'hw', 'true') +} + +fun mark_quoted(st) { + map_put(map_put(st, 'wq', 'true'), 'hw', 'true') +} + +fun finish_word(st) { + if (map_get(st, 'hw') == 'false') { st } + else { + w = join_chars(map_get(st, 'wl')) + rd = map_get(st, 'rd') + st2 = if (rd == 'out') { map_put(st, 'rt', list_append(map_get(st, 'rt'), w)) } + else { if (rd == 'in') { st } + else { if (rd == 'hplain' || rd == 'hdash') { + h = %{d: w, dash: if (rd == 'hdash') { 'true' } else { 'false' }, + quoted: map_get(st, 'wq')} + map_put(st, 'hd', list_append(map_get(st, 'hd'), h)) + } else { map_put(st, 'ws', list_append(map_get(st, 'ws'), w)) }}} + map_put(map_put(map_put(map_put(st2, 'wl', []), 'hw', 'false'), 'wq', 'false'), 'rd', 'none') + } +} + +fun finish_cmd(st) { + st1 = finish_word(st) + ws = map_get(st1, 'ws') + st2 = if (length(ws) > 0) { map_put(st1, 'cs', list_append(map_get(st1, 'cs'), ws)) } else { st1 } + map_put(map_put(st2, 'ws', []), 'rd', 'none') +} + +# ---------- command substitution capture ---------- + +fun start_capture(st, mode, ret) { + map_put(map_put(map_put(map_put(map_put(map_put(st, 'cm', mode), 'cl', []), 'cd', 1), 'cq', 0), + 'cesc', 'false'), 'cret', ret) +} + +fun cap_add(st, c) { map_put(st, 'cl', list_append(map_get(st, 'cl'), c)) } + +fun cap_step(c, rest, st) { + if (map_get(st, 'cesc') == 'true') { {rest, map_put(cap_add(st, c), 'cesc', 'false')} } + else { if (map_get(st, 'cm') == 'tick') { cap_tick(c, rest, st) } + else { cap_paren(c, rest, st) }} +} + +fun cap_tick(c, rest, st) { + if (c == "\\") { {rest, map_put(cap_add(st, c), 'cesc', 'true')} } + else { if (c == "`") { {rest, end_capture(st)} } + else { {rest, cap_add(st, c)} }} +} + +# Inside $(…): track quotes so a `)` in a string doesn't close it, and +# nesting depth so $(a $(b)) closes at the right paren. +fun cap_paren(c, rest, st) { + cq = map_get(st, 'cq') + if (cq == 1) { + if (c == "'") { {rest, map_put(cap_add(st, c), 'cq', 0)} } else { {rest, cap_add(st, c)} } + } + else { if (c == "\\") { {rest, map_put(cap_add(st, c), 'cesc', 'true')} } + else { if (cq == 2) { + if (c == "\"") { {rest, map_put(cap_add(st, c), 'cq', 0)} } else { {rest, cap_add(st, c)} } + } + else { if (c == "'") { {rest, map_put(cap_add(st, c), 'cq', 1)} } + else { if (c == "\"") { {rest, map_put(cap_add(st, c), 'cq', 2)} } + else { if (c == "(") { {rest, map_put(cap_add(st, c), 'cd', map_get(st, 'cd') + 1)} } + else { if (c == ")") { cap_close(rest, st) } + else { {rest, cap_add(st, c)} }}}}}}} +} + +fun cap_close(rest, st) { + d = map_get(st, 'cd') + if (d <= 1) { {rest, end_capture(st)} } + else { {rest, map_put(cap_add(st, ")"), 'cd', d - 1)} } +} + +# The captured text is a full command line of its own: parse it and fold its +# commands and redirect targets into ours. The surrounding word gets a `$` +# placeholder (its runtime value is unknowable). +fun end_capture(st) { + inner = parse(join_chars(map_get(st, 'cl'))) + ret = map_get(st, 'cret') + st1 = map_put(map_put(map_put(st, 'cm', nil), 'cl', []), 'q', ret) + st2 = map_put(map_put(st1, 'cs', map_get(st1, 'cs') ++ map_get(inner, 'cs')), + 'rt', map_get(st1, 'rt') ++ map_get(inner, 'rt')) + if (ret == 4) { st2 } else { add_char(st2, "$") } +} + +# ---------- joining chars without quadratic copying ---------- +# `acc ++ c` per char is O(n²) on a long word (a big `python -c '…'`), so +# join in balanced halves: O(n log n). +fun join_chars(lst) { join_n(lst, length(lst)) } + +fun join_n(lst, n) { + if (n <= 32) { join_lin(lst, n, "") } + else { + h = n / 2 + join_n(take_n(lst, h, []), h) ++ join_n(drop_n(lst, h), n - h) + } +} + +fun join_lin(lst, n, acc) { + if (n <= 0 || length(lst) == 0) { acc } else { join_lin(tl(lst), n - 1, acc ++ hd(lst)) } +} + +fun take_n(lst, n, acc) { + if (n <= 0 || length(lst) == 0) { acc } else { take_n(tl(lst), n - 1, list_append(acc, hd(lst))) } +} + +fun drop_n(lst, n) { + if (n <= 0 || length(lst) == 0) { lst } else { drop_n(tl(lst), n - 1) } +} + +# ============================================================ +# Judge — one simple command (a word list) → findings +# ============================================================ + +fun judge_all(cmds, depth, acc) { + if (length(cmds) == 0) { acc } + else { judge_all(tl(cmds), depth, acc ++ judge(hd(cmds), depth)) } +} + +fun finding(level, code, reason, words) { + {level, code, reason ++ " — in `" ++ string_truncate(join_words(words, ""), 120) ++ "`"} +} + +fun judge(words, depth) { + ws = skip_prefix(words) + if (length(ws) == 0 || depth > max_depth()) { [] } + else { + name = basename(hd(ws)) + args = tl(ws) + if (name == "sudo" || name == "doas") { + [finding('dangerous', 'sudo', name ++ " (privilege escalation)", ws)] ++ + judge(skip_opts(args, ["-u", "-g", "-C", "-D", "-h", "-p", "-r", "-t", "-U", "-T"]), depth + 1) + } + else { if (name == "env") { judge(skip_env(args), depth + 1) } + else { if (is_plain_wrapper(name) == 'true') { judge(skip_opts(args, []), depth + 1) } + else { if (name == "nice") { judge(skip_opts(args, ["-n"]), depth + 1) } + else { if (name == "ionice") { judge(skip_opts(args, ["-c", "-n", "-p", "-P", "-u"]), depth + 1) } + else { if (name == "timeout") { judge(drop_n(skip_opts(args, ["-s", "-k"]), 1), depth + 1) } + else { if (name == "xargs") { + judge(skip_opts(args, ["-I", "-i", "-n", "-P", "-L", "-l", "-s", "-d", "-E", "-e", "-a"]), depth + 1) + } + else { if (name == "watch") { judge(skip_opts(args, ["-n", "-d"]), depth + 1) } + else { if (is_shell(name) == 'true') { + script = shell_c_script(args, 'false') + if (script == nil) { [] } else { classify_text(script, depth + 1) } + } + else { if (name == "eval") { classify_text(join_words(args, ""), depth + 1) } + else { judge_verb(name, args, ws) }}}}}}}}}} + } +} + +fun is_plain_wrapper(name) { + in_list(["nohup", "exec", "command", "builtin", "time", "setsid", "unbuffer", + "busybox", "stdbuf", "chrt", "taskset", "caffeinate"], name) +} + +fun is_shell(name) { + in_list(["sh", "bash", "zsh", "dash", "ksh", "mksh", "ash", "fish"], name) +} + +# The script argument of `sh -c SCRIPT` (also -lc / -ec / -xc bundles), or +# nil for `sh file.sh` (a file we can't see). `-o opt` / `+o opt` take a value. +fun shell_c_script(args, saw_c) { + if (length(args) == 0) { nil } + else { + a = hd(args) + if (a == "-o" || a == "+o" || a == "-O" || a == "+O") { shell_c_script(drop_n(args, 2), saw_c) } + else { if (a == "--") { shell_c_script(tl(args), saw_c) } + else { if (string_starts_with(a, "-") == 'true' || string_starts_with(a, "+") == 'true') { + has_c = if (string_starts_with(a, "--") == 'false' && string_contains(a, "c") == 'true') { 'true' } else { saw_c } + shell_c_script(tl(args), has_c) + } else { + if (saw_c == 'true') { a } else { nil } + }}} + } +} + +fun judge_verb(name, args, ws) { + if (name == "rm") { rm_findings(args, ws) } + else { if (name == "mkfs" || string_starts_with(name, "mkfs.") == 'true' || + name == "mke2fs" || name == "mkswap") { + [finding('hardline', 'mkfs', name ++ " formats a filesystem / swap device", ws)] + } + else { if (name == "dd") { dd_findings(args, ws, []) } + else { if (in_list(["shutdown", "reboot", "halt", "poweroff", "telinit"], name) == 'true') { + [finding('hardline', 'halt', name ++ " halts or reboots the machine", ws)] + } + else { if (name == "init" && length(args) > 0 && (hd(args) == "0" || hd(args) == "6")) { + [finding('hardline', 'halt', "init " ++ hd(args) ++ " halts or reboots the machine", ws)] + } + else { if (name == "systemctl") { + sub = first_non_flag(args) + if (sub != nil && in_list(["poweroff", "reboot", "halt", "kexec"], sub) == 'true') { + [finding('hardline', 'halt', "systemctl " ++ sub ++ " halts or reboots the machine", ws)] + } else { [] } + } + else { if (name == "chmod" || name == "chown" || name == "chgrp") { perm_findings(name, args, ws) } + else { [] }}}}}}} +} + +# ---------- rm ---------- +# Recursive = -r/-R/--recursive, force = -f/--force, in any order or bundle +# (-rf, -fr, -Rf, -r -f). Everything after `--` is a target. +fun rm_findings(args, ws) { + s = rm_scan(args, 'false', 'false', 'false', [], 'false') + rec = elem(s, 0) + force = elem(s, 1) + nopres = elem(s, 2) + targets = elem(s, 3) + if (rec == 'true' && (nopres == 'true' || any_target(targets, 'root') == 'true')) { + what = if (nopres == 'true') { "rm -r --no-preserve-root" } else { "rm -r on the filesystem root (/ or /*)" } + [finding('hardline', 'rm_root', what, ws)] + } + else { if (rec == 'true' && any_target(targets, 'home') == 'true') { + [finding('dangerous', 'rm_home', "rm -r on your home directory", ws)] + } + else { if (rec == 'true' && force == 'true' && + (any_target(targets, 'under_home') == 'true' || any_target(targets, 'system') == 'true')) { + [finding('dangerous', 'rm_outside', "rm -rf on a path outside the project (home or system directory)", ws)] + } + else { [] }}} +} + +fun rm_scan(args, rec, force, nopres, targets, after_dd) { + if (length(args) == 0) { {rec, force, nopres, targets} } + else { + a = hd(args) + rest = tl(args) + if (after_dd == 'true') { rm_scan(rest, rec, force, nopres, list_append(targets, a), after_dd) } + else { if (a == "--") { rm_scan(rest, rec, force, nopres, targets, 'true') } + else { if (a == "--recursive") { rm_scan(rest, 'true', force, nopres, targets, after_dd) } + else { if (a == "--force") { rm_scan(rest, rec, 'true', nopres, targets, after_dd) } + else { if (a == "--no-preserve-root") { rm_scan(rest, rec, force, 'true', targets, after_dd) } + else { if (string_starts_with(a, "--") == 'true') { rm_scan(rest, rec, force, nopres, targets, after_dd) } + else { if (string_starts_with(a, "-") == 'true' && string_length(a) > 1) { + r2 = if (string_contains(a, "r") == 'true' || string_contains(a, "R") == 'true') { 'true' } else { rec } + f2 = if (string_contains(a, "f") == 'true') { 'true' } else { force } + rm_scan(rest, r2, f2, nopres, targets, after_dd) + } + else { rm_scan(rest, rec, force, nopres, list_append(targets, a), after_dd) }}}}}}} + } +} + +fun any_target(targets, kind) { + if (length(targets) == 0) { 'false' } + else { + t = norm_path(hd(targets)) + hit = if (kind == 'root') { is_root_path(t) } + else { if (kind == 'home') { is_home_path(t) } + else { if (kind == 'under_home') { is_under_home(t) } + else { is_system_path(t) }}} + if (hit == 'true') { 'true' } else { any_target(tl(targets), kind) } + } +} + +# Collapse `//` and drop trailing slashes (keeping a lone `/`). +fun norm_path(p) { + c = collapse_slashes(p) + strip_trailing_slashes(c) +} + +fun collapse_slashes(p) { + if (string_contains(p, "//") == 'true') { collapse_slashes(string_replace(p, "//", "/")) } else { p } +} + +fun strip_trailing_slashes(p) { + if (string_length(p) > 1 && string_ends_with(p, "/") == 'true') { + strip_trailing_slashes(string_sub(p, 0, string_length(p) - 1)) + } else { p } +} + +fun is_root_path(t) { + in_list(["/", "/*", "/.", "/..", "/.*"], t) +} + +# ~, $HOME, ${HOME} (optionally /* or /.) — or the literal $HOME path. +fun is_home_path(t) { + rest = home_rest(t) + if (rest != nil && in_list(["", "/*", "/.", "/.*"], rest) == 'true') { 'true' } + else { + h = getenv("HOME") + if (h == nil) { 'false' } + else { + hn = norm_path(to_string(h)) + if (hn != "/" && (t == hn || t == hn ++ "/*" || t == hn ++ "/.")) { 'true' } else { 'false' } + } + } +} + +fun is_under_home(t) { + if (home_rest(t) != nil) { 'true' } else { 'false' } +} + +# What follows a leading ~ / $HOME / ${HOME}, or nil if the path doesn't start with one. +fun home_rest(t) { + if (string_starts_with(t, "${HOME}") == 'true') { string_sub(t, 7, string_length(t) - 7) } + else { if (string_starts_with(t, "$HOME") == 'true') { string_sub(t, 5, string_length(t) - 5) } + else { if (string_starts_with(t, "~") == 'true') { + r = string_sub(t, 1, string_length(t) - 1) + # ~user/… is another user's home: treat like ~/… + if (r == "" || string_starts_with(r, "/") == 'true') { r } else { "/" ++ r } + } else { nil }}} +} + +# An absolute path outside the usual scratch/project roots (the pre-existing +# rm -rf policy: /tmp, /var/…, /home/…, /Users/…, /opt/… are fine). +fun is_system_path(t) { + if (string_starts_with(t, "/") == 'false') { 'false' } + else { if (starts_with_any(t, ["/tmp", "/var/", "/Users/", "/home/", "/opt/", + "/private/tmp", "/private/var/"]) == 'true') { 'false' } + else { 'true' }} +} + +# ---------- dd ---------- +fun dd_findings(args, ws, acc) { + if (length(args) == 0) { acc } + else { + a = hd(args) + next = if (string_starts_with(a, "of=") == 'true') { + dev = norm_path(string_sub(a, 3, string_length(a) - 3)) + if (is_raw_disk(dev) == 'true') { + list_append(acc, finding('hardline', 'dd_disk', "dd writing to raw disk device " ++ dev, ws)) + } else { if (string_starts_with(dev, "/dev/") == 'true' && is_benign_dev(dev) == 'false') { + list_append(acc, finding('dangerous', 'dd_dev', "dd writing to device " ++ dev, ws)) + } else { acc }} + } else { acc } + dd_findings(tl(args), ws, next) + } +} + +fun is_raw_disk(dev) { + starts_with_any(dev, ["/dev/sd", "/dev/nvme", "/dev/disk", "/dev/rdisk", "/dev/hd", "/dev/vd", + "/dev/xvd", "/dev/mmcblk", "/dev/md", "/dev/dm-", "/dev/mapper/"]) +} + +fun is_benign_dev(dev) { + if (in_list(["/dev/null", "/dev/zero", "/dev/stdout", "/dev/stderr", "/dev/tty", + "/dev/full", "/dev/random", "/dev/urandom"], dev) == 'true') { 'true' } + else { starts_with_any(dev, ["/dev/fd/", "/dev/pts/"]) } +} + +fun redirect_findings(targets, acc) { + if (length(targets) == 0) { acc } + else { + t = norm_path(hd(targets)) + next = if (is_raw_disk(t) == 'true') { + list_append(acc, {'hardline', 'redirect_disk', "shell redirection onto raw disk device " ++ t}) + } else { acc } + redirect_findings(tl(targets), next) + } +} + +# ---------- chmod / chown / chgrp ---------- +# Recursive (-R, bundles, --recursive) on / or /*, or `chmod 000 /`. +fun perm_findings(name, args, ws) { + rec = perm_recursive(args) + rest = non_flags(args, []) + mode = if (length(rest) == 0) { "" } else { hd(rest) } + targets = if (length(rest) == 0) { [] } else { tl(rest) } + lockout = if (name == "chmod" && in_list(["000", "0000", "a-rwx", "ugo-rwx", "a=", "ugo="], mode) == 'true') { 'true' } else { 'false' } + if (any_target(targets, 'root') == 'true' && (rec == 'true' || lockout == 'true')) { + flag = if (rec == 'true') { " -R" } else { " " ++ mode } + [finding('hardline', 'perm_root', name ++ flag ++ " on the filesystem root", ws)] + } else { [] } +} + +fun perm_recursive(args) { + if (length(args) == 0) { 'false' } + else { + a = hd(args) + if (a == "--recursive" || (string_starts_with(a, "-") == 'true' && string_starts_with(a, "--") == 'false' && + string_contains(a, "R") == 'true')) { 'true' } + else { perm_recursive(tl(args)) } + } +} + +# ---------- fork bomb ---------- +# `NAME(){ NAME|NAME& };NAME` in any spacing: find each `(){` definition in +# the whitespace-stripped command and check whether its body pipes the +# function into itself in the background. +fun fork_bomb_findings(s) { + stripped = string_replace(string_replace(string_replace(s, " ", ""), "\t", ""), "\n", "") + if (string_contains(stripped, "(){") == 'false') { [] } + else { fb_scan(string_split(stripped, "(){"), stripped) } +} + +fun fb_scan(parts, stripped) { + if (length(parts) < 2) { [] } + else { + name = trailing_name(hd(parts)) + if (string_length(name) > 0 && string_contains(stripped, name ++ "|" ++ name ++ "&") == 'true') { + [{'hardline', 'fork_bomb', "fork bomb (" ++ name ++ "(){ " ++ name ++ "|" ++ name ++ "& })"}] + } else { fb_scan(tl(parts), stripped) } + } +} + +# The function name right before `(){`: the chars after the last separator. +fun trailing_name(part) { + tn_loop(string_chars(part), "") +} + +fun tn_loop(chars, cur) { + if (length(chars) == 0) { cur } + else { + c = hd(chars) + if (in_list([";", "&", "|", "{", "}", "(", ")"], c) == 'true') { tn_loop(tl(chars), "") } + else { tn_loop(tl(chars), cur ++ c) } + } +} + +# ============================================================ +# Word helpers +# ============================================================ + +# Drop leading VAR=val assignments and reserved words (`!`, `{`, `if`, `then`, +# `do`, …) so the real command word comes first. +fun skip_prefix(words) { + if (length(words) == 0) { words } + else { + w = hd(words) + if (is_assignment(w) == 'true' || is_reserved(w) == 'true') { skip_prefix(tl(words)) } + else { words } + } +} + +fun is_reserved(w) { + in_list(["!", "{", "}", "if", "then", "else", "elif", "fi", "do", "done", "while", + "until", "case", "esac", "in", "for", "select", "function", "coproc", "[[", "]]"], w) +} + +fun is_assignment(w) { + i = string_index_of(w, "=") + if (i <= 0) { 'false' } + else { is_ident(string_sub(w, 0, i)) } +} + +fun is_ident(s) { + cs = string_chars(s) + if (length(cs) == 0) { 'false' } + else { if (is_digit_char(hd(cs)) == 'true') { 'false' } else { all_ident(cs) } } +} + +fun all_ident(cs) { + if (length(cs) == 0) { 'true' } + else { + c = hd(cs) + if ((c >= "a" && c <= "z") || (c >= "A" && c <= "Z") || is_digit_char(c) == 'true' || c == "_") { + all_ident(tl(cs)) + } else { 'false' } + } +} + +fun is_digit_char(c) { if (c >= "0" && c <= "9" && string_length(c) == 1) { 'true' } else { 'false' } } + +# Skip leading options; options named in `with_arg` consume the next word. +# Stops at the first non-option (the wrapped command) or after `--`. +fun skip_opts(args, with_arg) { + if (length(args) == 0) { args } + else { + a = hd(args) + if (a == "--") { tl(args) } + else { if (string_starts_with(a, "-") == 'true' && string_length(a) > 1) { + if (in_list(with_arg, a) == 'true') { skip_opts(drop_n(args, 2), with_arg) } + else { skip_opts(tl(args), with_arg) } + } else { args }} + } +} + +# env [-i] [-u NAME] [-C DIR] [NAME=val ...] CMD — assignments are dropped by +# skip_prefix inside judge(). +fun skip_env(args) { + skip_opts(args, ["-u", "-C", "-S", "--unset", "--chdir"]) +} + +fun first_non_flag(args) { + if (length(args) == 0) { nil } + else { if (string_starts_with(hd(args), "-") == 'true') { first_non_flag(tl(args)) } else { hd(args) } } +} + +fun non_flags(args, acc) { + if (length(args) == 0) { acc } + else { + a = hd(args) + next = if (string_starts_with(a, "-") == 'true') { acc } else { list_append(acc, a) } + non_flags(tl(args), next) + } +} + +fun basename(w) { + parts = string_split(w, "/") + last = last_of(parts, w) + if (string_length(last) == 0) { w } else { last } +} + +fun last_of(lst, fallback) { + if (length(lst) == 0) { fallback } + else { if (length(lst) == 1) { hd(lst) } else { last_of(tl(lst), fallback) } } +} + +fun join_words(words, acc) { + if (length(words) == 0) { acc } + else { + sep = if (string_length(acc) == 0) { "" } else { " " } + join_words(tl(words), acc ++ sep ++ hd(words)) + } +} + +fun starts_with_any(s, prefixes) { + if (length(prefixes) == 0) { 'false' } + else { if (string_starts_with(s, hd(prefixes)) == 'true') { 'true' } else { starts_with_any(s, tl(prefixes)) } } +} + +fun in_list(lst, item) { + if (length(lst) == 0) { 'false' } + else { if (hd(lst) == item) { 'true' } else { in_list(tl(lst), item) } } +} diff --git a/src/Flows.sw b/src/Flows.sw index 30515c7..662dbe4 100644 --- a/src/Flows.sw +++ b/src/Flows.sw @@ -30,8 +30,20 @@ module Flows # } # ] # } +# +# Robustness: a workflow file is user input. Its SHAPE is validated +# (validate_workflow) before the alt-screen opens or anything launches — +# a malformed file ({"phases":"oops"}, a task that isn't an object, a +# missing prompt) prints one error line and returns to the prompt +# instead of panicking the interactive session. +# +# Concurrency: a phase's tasks are QUEUED and launched at most +# flows_max_parallel() at a time (SWARM_CODE_FLOWS_MAX_PARALLEL, default +# 4); each render tick refills free slots. Previously a 12-task phase +# forked 12 full swarm-code children within 0.15s. import Background +import JsonCheck import FlowsRender import Scheduler import Util @@ -82,7 +94,10 @@ fun run_from_json_file(path, opts) { if (content == nil) { print(UI.err_color() ++ "error: cannot read file: " ++ path ++ UI.reset()) } else { - decoded = json_decode(content) + # Strict check first: json_decode guesses through damage (a + # truncated file decodes to a partial workflow that would run). + decoded = if (JsonCheck.valid(string_trim(content)) == 'true') { json_decode(content) } + else { nil } if (decoded == nil) { print(UI.err_color() ++ "error: invalid JSON in: " ++ path ++ UI.reset()) } else { @@ -92,10 +107,83 @@ fun run_from_json_file(path, opts) { } fun run_json_workflow(workflow_map, opts) { - # Reuse the agent's Background table (carried in opts): a private - # table restarts the bg-N id counter at bg-0 and collides with the - # agent's own (never-cleaned) /tmp/swarm-code-bg-N.* files, so a - # stale exit file could instantly mark a fresh task as done. + v = validate_workflow(workflow_map) + if (elem(v, 0) != 'ok') { + print(UI.err_color() ++ "error: invalid workflow — " ++ to_string(elem(v, 1)) ++ UI.reset()) + print(UI.grey_text() ++ " expected { \"phases\": [ { \"name\": \"…\", \"tasks\": " ++ + "[ { \"label\": \"…\", \"prompt\": \"…\" } ] } ] }" ++ UI.reset()) + } else { + run_valid_workflow(workflow_map, opts) + } +} + +# validate_workflow(w) → {'ok'} | {'error', reason}. Every shape the +# rest of this module walks with hd/tl/map_get is checked here, so the +# run itself never meets a non-list or a non-map. +fun validate_workflow(w) { + if (is_map(w) == 'false') { {'error', "the workflow must be a JSON object"} } + else { + phases = map_get(w, 'phases') + if (phases == nil) { {'ok'} } + else { if (is_list(phases) == 'false') { + {'error', "\"phases\" must be an array (got " ++ typeof(phases) ++ ")"} + } else { validate_phases(phases, 0) } } + } +} + +fun validate_phases(phases, i) { + if (length(phases) == 0) { {'ok'} } + else { + p = hd(phases) + where = "phase " ++ to_string(i + 1) + if (is_map(p) == 'false') { {'error', where ++ " must be an object"} } + else { + tasks = map_get(p, 'tasks') + r = if (tasks == nil) { {'ok'} } + else { if (is_list(tasks) == 'false') { + {'error', where ++ ": \"tasks\" must be an array (got " ++ typeof(tasks) ++ ")"} + } else { validate_tasks(tasks, where, 0) } } + if (elem(r, 0) != 'ok') { r } else { validate_phases(tl(phases), i + 1) } + } + } +} + +fun validate_tasks(tasks, where, j) { + if (length(tasks) == 0) { {'ok'} } + else { + t = hd(tasks) + at = where ++ " task " ++ to_string(j + 1) + if (is_map(t) == 'false') { {'error', at ++ " must be an object"} } + else { + prompt = map_get(t, 'prompt') + model = map_get(t, 'model') + label = map_get(t, 'label') + if (prompt == nil || typeof(prompt) != "string" || string_length(string_trim(prompt)) == 0) { + {'error', at ++ ": \"prompt\" must be a non-empty string"} + } else { if (model != nil && typeof(model) != "string") { + {'error', at ++ ": \"model\" must be a string"} + } else { if (label != nil && is_map(label) == 'true') { + {'error', at ++ ": \"label\" must be a string"} + } else { if (label != nil && is_list(label) == 'true') { + {'error', at ++ ": \"label\" must be a string"} + } else { validate_tasks(tl(tasks), where, j + 1) }}}} + } + } +} + +# Max concurrently running tasks: SWARM_CODE_FLOWS_MAX_PARALLEL (a +# positive integer; anything else is ignored), default 4. +fun flows_max_parallel() { + env = getenv("SWARM_CODE_FLOWS_MAX_PARALLEL") + n = if (env == nil) { nil } else { to_int(string_trim(env)) } + if (n == nil || n < 1 || to_string(n) != string_trim(to_string(env))) { 4 } else { n } +} + +fun run_valid_workflow(workflow_map, opts) { + # Reuse the agent's Background table (carried in opts) so flow tasks + # share the session's bg-N id space and private task directory (a + # fresh table would get its own directory — safe, but the tasks would + # be invisible to /bg). existing = map_get(opts, 'bg_table') bg_table = if (existing == nil) { Background.init() } else { existing } state = init_state(workflow_map) @@ -186,35 +274,103 @@ fun init_tasks(tasks, idx, acc) { } } -# spawn_phase_tasks(phase_idx, state, bg_table, opts) — launch all tasks in a phase +# spawn_phase_tasks(phase_idx, state, bg_table, opts) — start a phase: +# mark it running and launch its first flows_max_parallel() tasks; the +# rest stay queued ('pending', no bg_task_id) for fill_phase_slots. fun spawn_phase_tasks(phase_idx, state, bg_table, opts) { + phases = map_get(state, 'phases') + phase = list_nth(phases, phase_idx) + if (phase == nil) { state } + else { + new_phase = map_put(phase, 'status', 'running') + state2 = map_put(state, 'phases', list_replace_nth(phases, phase_idx, new_phase)) + fill_phase_slots(phase_idx, map_put(state2, 'selected_phase', phase_idx), bg_table, opts) + } +} + +# Launch queued tasks of phase `phase_idx` into free slots (called on +# phase start and every render tick). A no-op once the phase is full or +# drained. +fun fill_phase_slots(phase_idx, state, bg_table, opts) { phases = map_get(state, 'phases') phase = list_nth(phases, phase_idx) if (phase == nil) { state } else { tasks = map_get(phase, 'tasks') - new_tasks = spawn_task_list(tasks, bg_table, opts, phase_idx, []) - new_phase = map_put(phase, 'tasks', new_tasks) - new_phase2 = map_put(new_phase, 'status', 'running') - new_phases = list_replace_nth(phases, phase_idx, new_phase2) - state2 = map_put(state, 'phases', new_phases) - map_put(state2, 'selected_phase', phase_idx) + quota = launch_quota(tasks, flows_max_parallel()) + if (quota <= 0) { state } + else { + new_tasks = launch_queued(tasks, quota, bg_table, []) + new_phase = map_put(phase, 'tasks', new_tasks) + map_put(state, 'phases', list_replace_nth(phases, phase_idx, new_phase)) + } + } +} + +# How many queued tasks may start now: free slots (cap − running), +# bounded by how many are still queued. Pure — unit-tested. +fun launch_quota(tasks, cap) { + running = count_running(tasks, 0) + queued = count_queued(tasks, 0) + free = cap - running + if (free <= 0) { 0 } else { if (queued < free) { queued } else { free } } +} + +# Launched and not finished (bg id set, status not terminal). +fun count_running(tasks, acc) { + if (length(tasks) == 0) { acc } + else { + t = hd(tasks) + live = map_get(t, 'bg_task_id') != nil && task_is_terminal(t) == 'false' + count_running(tl(tasks), (if (live) { acc + 1 } else { acc })) } } -fun spawn_task_list(tasks, bg_table, opts, phase_idx, acc) { +# Not launched yet (and not failed to launch). +fun count_queued(tasks, acc) { if (length(tasks) == 0) { acc } else { t = hd(tasks) - prompt = map_get(t, 'prompt') - label = map_get(t, 'label') - model = map_get(t, 'model') - cmd = build_task_cmd(to_string(prompt), model) - bg_task_id = Background.launch(bg_table, cmd, to_string(label)) + q = map_get(t, 'bg_task_id') == nil && task_is_terminal(t) == 'false' + count_queued(tl(tasks), (if (q) { acc + 1 } else { acc })) + } +} + +fun task_is_terminal(t) { + s = map_get(t, 'status') + s_str = if (s == nil) { "pending" } else { to_string(s) } + if (s_str == "done" || s_str == "error" || s_str == "killed") { 'true' } else { 'false' } +} + +# Launch the first `quota` queued tasks, in order; others unchanged. +fun launch_queued(tasks, quota, bg_table, acc) { + if (length(tasks) == 0) { acc } + else { + t = hd(tasks) + queued = map_get(t, 'bg_task_id') == nil && task_is_terminal(t) == 'false' + if (quota > 0 && queued) { + launch_queued(tl(tasks), quota - 1, bg_table, list_append(acc, launch_task(t, bg_table))) + } else { + launch_queued(tl(tasks), quota, bg_table, list_append(acc, t)) + } + } +} + +fun launch_task(t, bg_table) { + prompt = map_get(t, 'prompt') + label = map_get(t, 'label') + model = map_get(t, 'model') + cmd = build_task_cmd(to_string(prompt), model) + bg_task_id = Background.launch(bg_table, cmd, to_string(label)) + now = timestamp() + if (string_starts_with(to_string(bg_task_id), "error:") == 'true') { + # Launch failed: finish the task as an error instead of leaving a + # bogus id that would read 'pending' forever and hang the flow. + map_put(map_put(map_put(t, 'status', 'error'), 'started_ms', now), 'ended_ms', now) + } else { new_t = map_put(t, 'bg_task_id', bg_task_id) new_t2 = map_put(new_t, 'status', 'running') - new_t3 = map_put(new_t2, 'started_ms', timestamp()) - spawn_task_list(tl(tasks), bg_table, opts, phase_idx, list_append(acc, new_t3)) + map_put(new_t2, 'started_ms', now) } } @@ -223,7 +379,8 @@ fun spawn_task_list(tasks, bg_table, opts, phase_idx, acc) { # settings in load_opts) — the binary has no --model flag, and a bare # --model value would be scanned as a positional profile name. # SWARM_CODE_DENY_DANGEROUS=1 keeps the dangerous-bash gate a hard deny -# in the headless children, which otherwise auto-approve every 'ask'. +# in the headless children, even if SWARM_CODE_HEADLESS_APPROVE=1 is +# exported (which makes them auto-approve every other 'ask'). fun build_task_cmd(prompt, model) { binary = swarm_binary() base = "SWARM_CODE_DENY_DANGEROUS=1 " ++ Util.shell_q(binary) ++ @@ -259,9 +416,11 @@ fun render_loop(state, bg_table, opts) { w = UI.term_width() FlowsRender.render_frame(refreshed3, w, 24) - # Check if current phase is done, spawn next phase if needed + # Refill the current phase's free slots from its queue, then check + # whether it is done and the next phase should start. selected = map_get(refreshed3, 'selected_phase') - phases = map_get(refreshed3, 'phases') + refilled = fill_phase_slots(selected, refreshed3, bg_table, opts) + phases = map_get(refilled, 'phases') current_phase = list_nth(phases, selected) cur_done = if (current_phase == nil) { 'true' } else { all_agents_done(current_phase) } @@ -270,12 +429,12 @@ fun render_loop(state, bg_table, opts) { next_idx = selected + 1 if (next_idx < length(phases)) { # Spawn next phase - spawn_phase_tasks(next_idx, refreshed3, bg_table, opts) + spawn_phase_tasks(next_idx, refilled, bg_table, opts) } else { # All phases done - map_put(refreshed3, 'phase', 'done') + map_put(refilled, 'phase', 'done') } - } else { refreshed3 } + } else { refilled } # Check if all done or user requested stop via stop-file all_done = all_phases_done(state_after_phase) @@ -329,14 +488,14 @@ fun refresh_tasks_loop(tasks, bg_table, acc) { bg_id = map_get(t, 'bg_task_id') new_t = if (bg_id == nil) { t } else { - s = task_file_status(to_string(bg_id)) + s = task_file_status(bg_table, to_string(bg_id)) t2 = map_put(t, 'status', s) if (s == 'done' || s == 'error') { ended = map_get(t2, 'ended_ms') t3 = if (ended == nil) { map_put(t2, 'ended_ms', timestamp()) } else { t2 } - exit_file = "/tmp/swarm-code-" ++ to_string(bg_id) ++ ".exit" - exit_content = file_read(exit_file) + exit_file = Background.exit_file_for(bg_table, to_string(bg_id)) + exit_content = if (exit_file == nil) { nil } else { file_read(exit_file) } ec = if (exit_content == nil) { 0 } else { parse_exit_code(string_trim(exit_content)) } map_put(t3, 'exit_code', ec) @@ -348,15 +507,15 @@ fun refresh_tasks_loop(tasks, bg_table, acc) { # task_file_status — check bg task status via exit/pid files directly, # bypassing Background ETS (which requires poll_and_notify to update). -fun task_file_status(bg_id) { - exit_file = "/tmp/swarm-code-" ++ bg_id ++ ".exit" +fun task_file_status(bg_table, bg_id) { + exit_file = to_string(Background.exit_file_for(bg_table, bg_id)) if (file_exists(exit_file) == 'true') { exit_content = file_read(exit_file) ec = if (exit_content == nil) { 0 } else { parse_exit_code(string_trim(exit_content)) } if (ec == 0) { 'done' } else { 'error' } } else { - pid_file = "/tmp/swarm-code-" ++ bg_id ++ ".pid" + pid_file = to_string(Background.pid_file_for(bg_table, bg_id)) if (file_exists(pid_file) == 'true') { 'running' } else { 'pending' } } } @@ -364,35 +523,35 @@ fun task_file_status(bg_id) { # refresh_log_stats — read token/tool counts from log files fun refresh_log_stats(state) { phases = map_get(state, 'phases') - new_phases = refresh_log_phases(phases, []) + new_phases = refresh_log_phases(phases, map_get(state, 'bg_table'), []) map_put(state, 'phases', new_phases) } -fun refresh_log_phases(phases, acc) { +fun refresh_log_phases(phases, bg_table, acc) { if (length(phases) == 0) { acc } else { p = hd(phases) tasks = map_get(p, 'tasks') - new_tasks = refresh_log_tasks(tasks, []) + new_tasks = refresh_log_tasks(tasks, bg_table, []) new_p = map_put(p, 'tasks', new_tasks) - refresh_log_phases(tl(phases), list_append(acc, new_p)) + refresh_log_phases(tl(phases), bg_table, list_append(acc, new_p)) } } -fun refresh_log_tasks(tasks, acc) { +fun refresh_log_tasks(tasks, bg_table, acc) { if (length(tasks) == 0) { acc } else { t = hd(tasks) bg_id = map_get(t, 'bg_task_id') - new_t = if (bg_id == nil) { t } + new_t = if (bg_id == nil || bg_table == nil) { t } else { - log_file = "/tmp/swarm-code-" ++ to_string(bg_id) ++ ".log" + log_file = to_string(Background.log_path_for(bg_table, to_string(bg_id))) stats = parse_log_stats(log_file) t2 = map_put(t, 'tokens_in', map_get(stats, 'tokens_in')) t3 = map_put(t2, 'tokens_out', map_get(stats, 'tokens_out')) map_put(t3, 'tool_calls', map_get(stats, 'tools')) } - refresh_log_tasks(tl(tasks), list_append(acc, new_t)) + refresh_log_tasks(tl(tasks), bg_table, list_append(acc, new_t)) } } diff --git a/src/Hooks.sw b/src/Hooks.sw index 414f17c..58c25f3 100644 --- a/src/Hooks.sw +++ b/src/Hooks.sw @@ -7,9 +7,17 @@ import Util # ============================================================ # # Executable scripts in ~/.swarm-code/hooks/ are called at key points -# in the agent loop. Context is passed via the SWARM_HOOK_DATA env var -# as a JSON string. Hooks have a 5-second timeout; on failure or -# timeout the agent proceeds normally (best-effort). +# in the agent loop. Hooks have a 5-second timeout (enforced by +# shell_managed, which kills the hook's whole process group). +# +# Context is a JSON document, delivered three ways (see run_with_payload): +# * on the hook's STDIN — always; the recommended way to read it +# * in a private 0600 temp file named by $SWARM_HOOK_DATA_FILE — always +# * in $SWARM_HOOK_DATA — only when it is under inline_payload_cap() +# bytes; above that it is omitted and $SWARM_HOOK_DATA_OMITTED=1 +# The payload used to go ONLY into the env var + command line, so a big +# write/edit (>128KB) hit E2BIG: the hook never ran, shell() then polled +# 120s for it, and a vetoing pre_tool.sh was silently skipped. # # Hook scripts: # pre_tool.sh — before every tool dispatch @@ -22,11 +30,14 @@ import Util # Exit 0 + prints {"args": {...}} → use modified args for dispatch # Exit 0 + prints anything else / nil → proceed as normal # Exit non-zero or timeout → proceed as normal +# Hook could not be RUN at all (payload file or launch failed) +# → VETO (fail closed), with the reason # # All other hooks are fire-and-forget; their output and exit code are # ignored (except for logging). -export [run_pre_tool, run_post_tool, run_pre_llm, run_post_llm, hooks_dir] +export [run_pre_tool, run_pre_tool_at, run_post_tool, run_pre_llm, run_post_llm, hooks_dir, + run_with_payload, inline_payload_cap] fun hooks_dir() { getenv("HOME") ++ "/.swarm-code/hooks" @@ -38,53 +49,112 @@ fun hook_path(name) { fun hook_timeout_s() { 5 } -# Build a timeout-wrapped shell command that sets SWARM_HOOK_DATA and -# executes the hook script. The JSON data is single-quote-escaped so it -# is safe to splice into the shell command line. -fun hook_cmd(path, json_data) { - "export SWARM_HOOK_DATA=" ++ Util.shell_q(json_data) ++ "; " ++ - "perl -e 'alarm shift; exec @ARGV' " ++ to_string(hook_timeout_s()) ++ - " sh " ++ Util.shell_q(path) +# Largest payload also exported inline as an env var. One env string must +# stay under the kernel's per-string limit (MAX_ARG_STRLEN, 128KB) with room +# for the rest of the environment. +fun inline_payload_cap() { 100000 } + +fun tmp_prefix() { + t = getenv("TMPDIR") + base = if (t == nil || string_length(to_string(t)) == 0) { "/tmp" } else { to_string(t) } + base ++ "/swarm-code-hook-" +} + +# ============================================================ +# run_with_payload — run a hook shell snippet with a JSON payload +# ============================================================ +# The payload goes into a private temp file (mkstemp → 0600), which becomes +# the hook's stdin and is named by $_FILE; $ carries it +# inline too when it's small (set INSIDE the script from the file, so it +# never touches the command line). `exports` is a shell prefix of extra +# `export K=V;` lines. stderr is folded into stdout. +# +# Returns %{ran: 'true', code, out, interrupted} when the hook ran, or +# %{ran: 'false', error} when it could NOT be run (callers fail closed). +fun run_with_payload(body, payload, env_name, exports, timeout_ms, prefix) { + f = file_temp(prefix) + if (f == nil) { + %{ran: 'false', error: "could not create a temp file for the hook payload (" ++ prefix ++ "…)"} + } else { + data = to_string(payload) + wrote = file_write(f, data) + if (wrote != 'ok') { + file_delete(f) + %{ran: 'false', error: "could not write the hook payload to " ++ f} + } else { + file_var = env_name ++ "_FILE" + inline = if (string_length(data) <= inline_payload_cap()) { + env_name ++ "=$(cat \"$" ++ file_var ++ "\"); export " ++ env_name + } else { + "unset " ++ env_name ++ "; export " ++ env_name ++ "_OMITTED=1" + } + script = "export " ++ file_var ++ "=" ++ Util.shell_q(f) ++ "; " ++ exports ++ "\n" ++ + inline ++ "\n" ++ + "exec <\"$" ++ file_var ++ "\" 2>&1\n" ++ + to_string(body) ++ "\n" + r = shell_managed(script, timeout_ms) + file_delete(f) + code = elem(r, 0) + if (code < 0) { + %{ran: 'false', error: "hook failed to launch: " ++ to_string(elem(r, 1))} + } else { + %{ran: 'true', code: code, out: to_string(elem(r, 1)), interrupted: elem(r, 2)} + } + } + } +} + +fun run_script_hook(path, data_json) { + run_with_payload("sh " ++ Util.shell_q(path), data_json, "SWARM_HOOK_DATA", "", + hook_timeout_s() * 1000, tmp_prefix()) } # ============================================================ # run_pre_tool — called before every tool dispatch. # Returns %{veto: 'false', args: original_args} normally, or -# %{veto: 'true', args: original_args} when vetoed. +# %{veto: 'true', args: original_args, reason} when vetoed. # If the hook prints {"args": {...}}, the modified args map is returned. # ============================================================ fun run_pre_tool(tool_name, args_map, opts) { - path = hook_path("pre_tool.sh") + run_pre_tool_at(hook_path("pre_tool.sh"), tool_name, args_map, tmp_prefix()) +} + +fun run_pre_tool_at(path, tool_name, args_map, prefix) { if (file_exists(path) == 'false') { %{veto: 'false', args: args_map} } else { - args_json = json_encode(args_map) data_json = json_encode(%{tool: to_string(tool_name), args: args_map}) - cmd = hook_cmd(path, data_json) - result = shell(cmd) - code = elem(result, 0) - out = string_trim(elem(result, 1)) - if (code != 0) { - # Hook failed or timed out — proceed normally - %{veto: 'false', args: args_map} + res = run_with_payload("sh " ++ Util.shell_q(path), data_json, "SWARM_HOOK_DATA", "", + hook_timeout_s() * 1000, prefix) + if (map_get(res, 'ran') != 'true') { + # Fail CLOSED: a veto hook that never ran must not wave the call through. + %{veto: 'true', args: args_map, + reason: "pre_tool hook " ++ path ++ " could not run — " ++ to_string(map_get(res, 'error'))} } else { - if (string_length(out) == 0) { + code = map_get(res, 'code') + out = string_trim(map_get(res, 'out')) + if (code != 0) { + # Hook failed or timed out — proceed normally (documented contract) %{veto: 'false', args: args_map} } else { - parsed = json_decode(out) - if (parsed == nil) { + if (string_length(out) == 0) { %{veto: 'false', args: args_map} } else { - veto_val = map_get(parsed, 'veto') - new_args = map_get(parsed, 'args') - is_veto = veto_val == 'true' || veto_val == "true" || veto_val == true - if (is_veto == true) { - %{veto: 'true', args: args_map} + parsed = json_decode(out) + if (parsed == nil || is_map(parsed) == 'false') { + %{veto: 'false', args: args_map} } else { - if (new_args != nil) { - %{veto: 'false', args: new_args} + veto_val = map_get(parsed, 'veto') + new_args = map_get(parsed, 'args') + is_veto = veto_val == 'true' || veto_val == "true" || veto_val == true + if (is_veto == true) { + %{veto: 'true', args: args_map, reason: "vetoed by pre_tool hook " ++ path} } else { - %{veto: 'false', args: args_map} + if (new_args != nil) { + %{veto: 'false', args: new_args} + } else { + %{veto: 'false', args: args_map} + } } } } @@ -106,8 +176,7 @@ fun run_post_tool(tool_name, result_str, exit_code, opts) { result: to_string(result_str), exit_code: exit_code }) - cmd = hook_cmd(path, data_json) - shell(cmd) + run_script_hook(path, data_json) 'ok' } } @@ -124,8 +193,7 @@ fun run_pre_llm(model, n_messages, opts) { model: to_string(model), messages: n_messages }) - cmd = hook_cmd(path, data_json) - shell(cmd) + run_script_hook(path, data_json) 'ok' } } @@ -143,8 +211,7 @@ fun run_post_llm(model, tokens, latency_ms, opts) { tokens: tokens, latency_ms: latency_ms }) - cmd = hook_cmd(path, data_json) - shell(cmd) + run_script_hook(path, data_json) 'ok' } } diff --git a/src/JsonCheck.sw b/src/JsonCheck.sw new file mode 100644 index 0000000..4018e37 --- /dev/null +++ b/src/JsonCheck.sw @@ -0,0 +1,182 @@ +module JsonCheck + +# ============================================================ +# JsonCheck — strict JSON well-formedness (RFC 8259) +# ============================================================ +# +# The runtime's json_decode is deliberately forgiving: it never fails +# on a truncated or hand-mangled document, it guesses. Observed: +# "[1, 2" → [1, 2] (unterminated) +# "[1,,2]" → [1, nil, 2] +# "[tru]" → [nil, nil, nil] +# "{\"a\":} junk" → %{a: nil} +# That is fine for tolerant reads, but a WRITER must be able to tell +# "this file is valid JSON" from "this file is damaged": treating a +# damaged schedule.json as a parsed list and rewriting it silently +# destroys the user's data. valid(s) answers that question exactly — +# one JSON value, optionally surrounded by whitespace, nothing else. +# +# Byte-level scan via codepoint_at (sw has no regex). Structure is +# recursive descent with nesting capped at 512; every per-element loop +# (array items, object members, string bytes, digits) is SELF-tail- +# recursive, so long documents run in a flat stack. UTF-8 bytes >= +# 0x80 inside strings are accepted as-is (not re-validated). + +export [valid] + +fun max_depth() { 512 } + +# 'true' iff `s` is exactly one well-formed JSON value (+ whitespace). +fun valid(s) { + if (s == nil) { 'false' } + else { + str = to_string(s) + n = string_length(str) + i = jc_ws(str, 0, n) + j = jc_value(str, i, n, 0) + if (j < 0) { 'false' } + else { if (jc_ws(str, j, n) == n) { 'true' } else { 'false' } } + } +} + +# Skip JSON whitespace (space, tab, LF, CR). Returns the next index. +fun jc_ws(s, i, n) { + if (i >= n) { i } + else { + c = codepoint_at(s, i) + if (c == 32 || c == 9 || c == 10 || c == 13) { jc_ws(s, i + 1, n) } + else { i } + } +} + +# Parse one value starting at i. Returns the index just past it, or -1. +fun jc_value(s, i, n, depth) { + if (i >= n || depth > max_depth()) { 0 - 1 } + else { + c = codepoint_at(s, i) + if (c == 123) { jc_object(s, jc_ws(s, i + 1, n), n, depth + 1) } # { + else { if (c == 91) { jc_array(s, jc_ws(s, i + 1, n), n, depth + 1) } # [ + else { if (c == 34) { jc_string(s, i + 1, n) } # " + else { if (c == 116) { jc_lit(s, i, n, "true") } + else { if (c == 102) { jc_lit(s, i, n, "false") } + else { if (c == 110) { jc_lit(s, i, n, "null") } + else { if (c == 45 || jc_digit(c) == 'true') { jc_number(s, i, n) } + else { 0 - 1 }}}}}}} + } +} + +fun jc_lit(s, i, n, word) { + k = string_length(word) + if (i + k <= n && string_sub(s, i, k) == word) { i + k } else { 0 - 1 } +} + +fun jc_digit(c) { if (c >= 48 && c <= 57) { 'true' } else { 'false' } } + +fun jc_hex(c) { + if (jc_digit(c) == 'true' || (c >= 97 && c <= 102) || (c >= 65 && c <= 70)) { 'true' } + else { 'false' } +} + +# Array body; `i` is just past '[' and its whitespace. +fun jc_array(s, i, n, depth) { + if (i < n && codepoint_at(s, i) == 93) { i + 1 } # empty [] + else { jc_array_items(s, i, n, depth) } +} + +fun jc_array_items(s, i, n, depth) { + j = jc_value(s, i, n, depth) + if (j < 0) { 0 - 1 } + else { + k = jc_ws(s, j, n) + if (k >= n) { 0 - 1 } + else { + c = codepoint_at(s, k) + if (c == 44) { jc_array_items(s, jc_ws(s, k + 1, n), n, depth) } # , + else { if (c == 93) { k + 1 } else { 0 - 1 } } # ] + } + } +} + +# Object body; `i` is just past '{' and its whitespace. +fun jc_object(s, i, n, depth) { + if (i < n && codepoint_at(s, i) == 125) { i + 1 } # empty {} + else { jc_members(s, i, n, depth) } +} + +fun jc_members(s, i, n, depth) { + if (i >= n || codepoint_at(s, i) != 34) { 0 - 1 } # key must be a string + else { + kend = jc_string(s, i + 1, n) + if (kend < 0) { 0 - 1 } + else { + colon = jc_ws(s, kend, n) + if (colon >= n || codepoint_at(s, colon) != 58) { 0 - 1 } # : + else { + j = jc_value(s, jc_ws(s, colon + 1, n), n, depth) + if (j < 0) { 0 - 1 } + else { + k = jc_ws(s, j, n) + if (k >= n) { 0 - 1 } + else { + c = codepoint_at(s, k) + if (c == 44) { jc_members(s, jc_ws(s, k + 1, n), n, depth) } + else { if (c == 125) { k + 1 } else { 0 - 1 } } + } + } + } + } + } +} + +# String body; `i` is just past the opening quote. Control characters +# must be escaped; escapes are \" \\ \/ \b \f \n \r \t \uXXXX only. +fun jc_string(s, i, n) { + if (i >= n) { 0 - 1 } + else { + c = codepoint_at(s, i) + if (c == 34) { i + 1 } + else { if (c < 32) { 0 - 1 } + else { if (c == 92) { + if (i + 1 >= n) { 0 - 1 } + else { + e = codepoint_at(s, i + 1) + if (e == 34 || e == 92 || e == 47 || e == 98 || e == 102 || + e == 110 || e == 114 || e == 116) { jc_string(s, i + 2, n) } + else { if (e == 117) { + if (i + 5 < n && jc_hex(codepoint_at(s, i + 2)) == 'true' && + jc_hex(codepoint_at(s, i + 3)) == 'true' && + jc_hex(codepoint_at(s, i + 4)) == 'true' && + jc_hex(codepoint_at(s, i + 5)) == 'true') { jc_string(s, i + 6, n) } + else { 0 - 1 } + } else { 0 - 1 } } + } + } else { jc_string(s, i + 1, n) } } } + } +} + +# -?(0|[1-9][0-9]*)(\.[0-9]+)?([eE][+-]?[0-9]+)? +fun jc_number(s, i, n) { + a = if (codepoint_at(s, i) == 45) { i + 1 } else { i } + if (a >= n || jc_digit(codepoint_at(s, a)) == 'false') { 0 - 1 } + else { + b = if (codepoint_at(s, a) == 48) { a + 1 } else { jc_digits(s, a, n) } + c = if (b < n && codepoint_at(s, b) == 46) { # . + if (b + 1 < n && jc_digit(codepoint_at(s, b + 1)) == 'true') { jc_digits(s, b + 1, n) } + else { 0 - 1 } + } else { b } + if (c < 0) { 0 - 1 } + else { + if (c < n && (codepoint_at(s, c) == 101 || codepoint_at(s, c) == 69)) { # e E + d = if (c + 1 < n && (codepoint_at(s, c + 1) == 43 || codepoint_at(s, c + 1) == 45)) { c + 2 } + else { c + 1 } + if (d < n && jc_digit(codepoint_at(s, d)) == 'true') { jc_digits(s, d, n) } + else { 0 - 1 } + } else { c } + } + } +} + +fun jc_digits(s, i, n) { + if (i < n && jc_digit(codepoint_at(s, i)) == 'true') { jc_digits(s, i + 1, n) } + else { i } +} diff --git a/src/McpServer.sw b/src/McpServer.sw index 59c6d25..de37f42 100644 --- a/src/McpServer.sw +++ b/src/McpServer.sw @@ -13,9 +13,15 @@ module McpServer # # Handled methods: # initialize → server capabilities + info +# ping → {} (MCP basic/utilities/ping) # notifications/initialized → (notification, no reply needed) # tools/list → list available tools # tools/call → execute a tool, return text result +# (unknown tool → -32602 Invalid params) +# +# Envelope rules (JSON-RPC 2.0 + MCP basic/messages): "jsonrpc" must be +# exactly "2.0"; a request id must be a string or number (never null, +# an object, array or boolean) — violations get -32600 Invalid Request. # # Exposed tools (a practical subset that covers 95% of coding tasks): # bash, read, write, edit, glob, grep, web_fetch @@ -153,18 +159,26 @@ fun handle_tools_list(id) { ok_response(id, %{tools: mcp_tools}) } +# tools/call param errors are -32602 Invalid params — including an +# unknown tool name (MCP spec, server/tools "Error Handling": "Unknown +# tools" → -32602). -32601 is reserved for an unknown METHOD. fun handle_tools_call(id, params, opts) { - name = if (params == nil) { nil } else { map_get(params, 'name') } - args = if (params == nil) { %{} } else { - raw = map_get(params, 'arguments') - if (raw == nil) { %{} } else { raw } - } - if (name == nil) { + pmap = if (params == nil) { %{} } else { params } + name = if (is_map(pmap) == 'true') { map_get(pmap, 'name') } else { nil } + raw_args = if (is_map(pmap) == 'true') { map_get(pmap, 'arguments') } else { nil } + args = if (raw_args == nil) { %{} } else { raw_args } + if (is_map(pmap) == 'false') { + err_response(id, -32602, "invalid params: tools/call params must be an object") + } else { if (name == nil) { err_response(id, -32602, "missing tool name") + } else { if (typeof(name) != "string") { + err_response(id, -32602, "invalid params: tool name must be a string") + } else { if (is_map(args) == 'false') { + err_response(id, -32602, "invalid params: arguments must be an object") } else { name_s = to_string(name) if (list_member(exposed_tool_names(), name_s) == 'false') { - err_response(id, -32601, "tool not found: " ++ name_s) + err_response(id, -32602, "unknown tool: " ++ name_s) } else { # ToolRegistry.atom_for converts "bash" → 'bash' so the shared # execution boundary can reach the raw handler registry. @@ -180,40 +194,92 @@ fun handle_tools_call(id, params, opts) { tool_result_response(id, result_s) } } - } + }}}} } # ============================================================ # Dispatch a single parsed request # ============================================================ +# Key presence by name: json_decode keys are atoms and map_has_key +# reports 'false' for a key whose value is JSON null — but `"id": null` +# is a present-and-invalid id, not a Notification. +fun has_member(m, k) { member_scan(map_keys(m), k) } + +fun member_scan(keys, k) { + if (length(keys) == 0) { 'false' } + else { if (to_string(hd(keys)) == k) { 'true' } + else { member_scan(tl(keys), k) } } +} + +# MCP (basic/messages): a request id MUST be a string or integer and +# MUST NOT be null. Numbers with a fraction are tolerated (JSON-RPC +# 2.0 only says SHOULD NOT); objects / arrays / booleans are rejected. +fun valid_id(v) { + t = typeof(v) + if (v == nil) { 'false' } + else { if (t == "string" || t == "int" || t == "float") { 'true' } + else { 'false' } } +} + +# Validate the JSON-RPC 2.0 envelope. Returns nil when well-formed, or +# the -32600 Invalid Request error line. Per JSON-RPC §5 the error +# carries the request's id when that id was itself valid, else null. +fun envelope_error(msg) { + has_id = has_member(msg, "id") + id_raw = map_get(msg, 'id') + id_ok = if (has_id == 'false') { 'true' } else { valid_id(id_raw) } + reply_id = if (id_ok == 'true') { id_raw } else { nil } + jsonrpc = map_get(msg, 'jsonrpc') + method = map_get(msg, 'method') + params = map_get(msg, 'params') + if (jsonrpc != "2.0" || typeof(jsonrpc) != "string") { + err_response(reply_id, -32600, "invalid request: \"jsonrpc\" must be exactly \"2.0\"") + } else { if (id_ok == 'false') { + err_response(nil, -32600, "invalid request: id must be a string or number (not null, object, array or boolean)") + } else { if (method == nil) { + # No method member: an Invalid Request, not an unknown method + # (which would be -32601). + err_response(reply_id, -32600, "invalid request: missing method") + } else { if (typeof(method) != "string") { + err_response(reply_id, -32600, "invalid request: method must be a string") + } else { if (params != nil && is_map(params) == 'false' && is_list(params) == 'false') { + err_response(reply_id, -32600, "invalid request: params must be an object or array") + } else { nil }}}}} +} + fun dispatch(msg, opts) { + invalid = envelope_error(msg) id = map_get(msg, 'id') method = map_get(msg, 'method') params = map_get(msg, 'params') method_s = if (method == nil) { "" } else { to_string(method) } - # JSON-RPC 2.0 §4.1: a message without `id` is a Notification — the - # server MUST NOT reply. The only notification we act on is - # "notifications/initialized"; every other notification (including - # notification-shaped initialize/tools/list/tools/call) is silently - # ignored, with NO tool side-effects, mirroring the spec. Returning - # nil here suppresses output in server_loop (the `response != nil` gate). - if (id == nil) { - nil + if (invalid != nil) { + invalid } - else { if (method == nil) { - # Request with no method member is an Invalid Request, not an - # unknown method (which would be -32601). - err_response(id, -32600, "invalid request: missing method") + # JSON-RPC 2.0 §4.1: a well-formed message without an `id` member is + # a Notification — the server MUST NOT reply. The only notification + # we act on is "notifications/initialized"; every other notification + # (including notification-shaped initialize/tools/list/tools/call) is + # silently ignored, with NO tool side-effects, mirroring the spec. + # Returning nil suppresses output in server_loop (the `response != + # nil` gate). + else { if (has_member(msg, "id") == 'false') { + nil } else { if (method_s == "initialize") { handle_initialize(id) } + else { if (method_s == "ping") { + # MCP basic/utilities/ping: the receiver MUST respond promptly + # with an empty result. + ok_response(id, %{}) + } else { if (method_s == "notifications/initialized") { - # A request-shaped (id-bearing) "initialized" is unusual but the - # notification path above already handles the normal no-id case. - nil + # A request-shaped (id-bearing) "initialized" is unusual, but + # every request is owed a response — acknowledge it. + ok_response(id, %{}) } else { if (method_s == "tools/list") { handle_tools_list(id) @@ -224,7 +290,7 @@ fun dispatch(msg, opts) { else { # Known shape, unrecognized method name. err_response(id, -32601, "method not found: " ++ method_s) - }}}}}} + }}}}}}} } # ============================================================ diff --git a/src/PathGuard.sw b/src/PathGuard.sw index 0553c2c..069b7fa 100644 --- a/src/PathGuard.sw +++ b/src/PathGuard.sw @@ -24,6 +24,16 @@ module PathGuard # /.config/gh/hosts.yml (gh CLI auth token) # /.config/git/credentials (git credentials) # /.swarm-code/settings.json (our own config) +# /.swarm-code/** (our control files: hooks/ run on every +# tool call, schedule.json queues headless +# runs, .profile_override redirects the +# LLM, sessions/ are replayed as history) +# EXCEPT memory/ and skills/, the data +# the agent legitimately manages +# */.swarm-code.json (project-scope settings) +# +# Matching is case-insensitive: on a case-insensitive filesystem (macOS +# default) ~/.SSH/ and ~/.Swarm-Code/hooks/ ARE the protected paths. # # Blocked on READ (private key material only): # /.gnupg/private-keys* (GPG private key material) @@ -52,7 +62,8 @@ fun is_sensitive(path) { _sensitive_check(rp) } -fun _sensitive_check(rp) { +fun _sensitive_check(rp_raw) { + rp = string_lower(rp_raw) if (string_contains(rp, "/.ssh/") == 'true') { 'true' } else { if (string_contains(rp, "/.gnupg/") == 'true') { 'true' } else { if (string_starts_with(rp, "/etc/") == 'true') { 'true' } @@ -78,7 +89,8 @@ fun validate_write(path) { } } -fun _write_check(rp) { +fun _write_check(rp_raw) { + rp = string_lower(rp_raw) if (string_contains(rp, "/.ssh/") == 'true') { "error: write to sensitive path blocked: refusing to write inside an .ssh directory — set SWARM_CODE_UNSAFE_WRITES=1 to override" } @@ -112,7 +124,40 @@ fun _write_check(rp) { else { if (string_contains(rp, "/.swarm-code/settings.json") == 'true') { "error: write to sensitive path blocked: refusing to write swarm-code's own settings.json — edit it yourself, or set SWARM_CODE_UNSAFE_WRITES=1 to override" } - else { "ok" }}}}}}}}}}} + else { if (_is_control_path(rp) == 'true') { + "error: write to sensitive path blocked: refusing to write swarm-code's own control files (~/.swarm-code/ hooks, schedule, sessions, profile override; ./.swarm-code.json) — only ~/.swarm-code/memory/ and ~/.swarm-code/skills/ are writable; edit it yourself, or set SWARM_CODE_UNSAFE_WRITES=1 to override" + } + else { "ok" }}}}}}}}}}}} +} + +# 'true' for a (lowercased, resolved) path swarm-code reads as control +# input: anything under a .swarm-code/ directory other than its memory/ +# and skills/ data dirs, the directory itself, and a project-scope +# .swarm-code.json. Matches any .swarm-code/ (not just $HOME's) — HOME +# may be unset or a symlink, and over-blocking here is the safe side. +fun _is_control_path(rp) { + if (string_ends_with(rp, "/.swarm-code.json") == 'true') { 'true' } + else { if (string_ends_with(rp, "/.swarm-code") == 'true') { 'true' } + else { + idx = string_index_of(rp, "/.swarm-code/") + if (idx < 0) { 'false' } + else { + start = idx + string_length("/.swarm-code/") + rest = string_sub(rp, start, string_length(rp) - start) + data_dir = if (string_starts_with(rest, "memory/") == 'true') { 'true' } + else { string_starts_with(rest, "skills/") } + # realpath -m is missing on older macOS (resolve_path then + # returns the path as given), so an unresolved ".." must not + # climb out of a data dir: memory/../hooks/pre_tool.sh. + if (data_dir == 'true' && _has_dotdot(rest) == 'false') { 'false' } else { 'true' } + } + }} +} + +fun _has_dotdot(p) { + if (string_contains(p, "/../") == 'true') { 'true' } + else { if (string_ends_with(p, "/..") == 'true') { 'true' } + else { string_starts_with(p, "../") }} } # ------------------------------------------------------------------ @@ -126,7 +171,8 @@ fun validate_read(path) { _read_check(rp) } -fun _read_check(rp) { +fun _read_check(rp_raw) { + rp = string_lower(rp_raw) if (string_contains(rp, "/.gnupg/private-keys") == 'true') { "error: read blocked: refusing to read GPG private key material in .gnupg/private-keys*" } diff --git a/src/SessionSearch.sw b/src/SessionSearch.sw index 4d39df2..e4b0fc0 100644 --- a/src/SessionSearch.sw +++ b/src/SessionSearch.sw @@ -9,7 +9,7 @@ import Util # Layout on disk: # # ~/.swarm-code/sessions/ -# index.db — SQLite with FTS5 virtual table + mtime meta +# index.db — SQLite with FTS5 virtual table + size/mtime meta # journal-.jsonl — existing per-session journals (one JSON per line) # .active — pointer to current session journal # @@ -18,68 +18,79 @@ import Util # System prompts are not journaled so they don't pollute the index. # # How the cache stays fresh: at session start we walk the journals -# directory, compare each file's mtime against the cached value in -# `meta`, and reindex only the ones that changed. A journal is +# directory, compare each file's size + mtime (file_stat — an in-process +# stat(2), no shell) against the values recorded in `meta` when it was +# last indexed, and reindex only the ones that changed. A journal is # rewritten on every turn (`journal_sync`), so the active session's -# new turns land in the index on the next session start. +# new turns land in the index on the next session start — including a +# journal that was still being written by ANOTHER instance when this +# one booted (previously it was marked indexed once and never +# refreshed, because only "present in meta" was checked). # # Usage: # /search QUERY — slash command, prints top hits inline # session_search tool — agent-callable; returns hits as a string -export [init, search, search_render, db_path, sessions_dir] +export [init, init_at, search, search_at, search_render, db_path, sessions_dir] fun sessions_dir() { getenv("HOME") ++ "/.swarm-code/sessions" } -fun db_path() { sessions_dir() ++ "/index.db" } +fun db_path() { db_path_at(sessions_dir()) } +fun db_path_at(dir) { dir ++ "/index.db" } # ------------------------------------------------------------ # init — create schema if missing, then incrementally reindex. -# Idempotent. Called from main.sw at session start. +# Idempotent. Called from main.sw at session start. init_at(dir) is the +# same over an explicit sessions directory (unit tests). # ------------------------------------------------------------ -fun init() { - file_mkdir(sessions_dir()) - db = db_open(db_path()) +fun init() { init_at(sessions_dir()) } + +fun init_at(dir) { + file_mkdir(dir) + db = db_open(db_path_at(dir)) db_exec(db, "CREATE VIRTUAL TABLE IF NOT EXISTS journals USING fts5(" ++ "session UNINDEXED, role UNINDEXED, content)") db_exec(db, "CREATE TABLE IF NOT EXISTS meta(" ++ - "session TEXT PRIMARY KEY, indexed_at INTEGER)") - reindex(db) + "session TEXT PRIMARY KEY, indexed_at INTEGER, size INTEGER, mtime INTEGER)") + # Migrate an index.db created before size/mtime were tracked. On a + # current schema these fail ("duplicate column") — harmless. Old + # rows read back NULL, so every such journal reindexes once. + db_exec(db, "ALTER TABLE meta ADD COLUMN size INTEGER") + db_exec(db, "ALTER TABLE meta ADD COLUMN mtime INTEGER") + reindex(db, dir) db_close(db) 'session_search_ready' } # ------------------------------------------------------------ # reindex — enumerate journals via the file_list builtin (no shell) -# and index any that aren't in the meta cache yet. The currently- -# active session (per the .active marker) is always re-indexed, -# since its turns grow during a run. -# -# Why no mtime check: swarmrt's shell() polls every 1s, so 60 files -# × 1 stat call each = 60s startup. file_list is in-process and free, -# but doesn't surface mtimes — so we trade per-file freshness for -# instant boot. The active session is the only one that mutates -# during a run, and we always refresh it. +# and (re)index every one that is new or whose size / mtime differ +# from what meta recorded at its last indexing. The currently-active +# session (per the .active marker) is always re-indexed as well: its +# turns grow during a run and a same-second rewrite could keep mtime. +# file_stat is an in-process stat(2) — the old "no mtime check, stat +# costs a 1s shell() poll per file" trade-off no longer applies. # ------------------------------------------------------------ -fun reindex(db) { - active = read_active_marker() - names = file_list(sessions_dir()) - reindex_loop(db, names, active) +fun reindex(db, dir) { + active = read_active_marker(dir) + names = file_list(dir) + reindex_loop(db, dir, names, active) } -fun reindex_loop(db, names, active) { +fun reindex_loop(db, dir, names, active) { if (length(names) == 0) { 'ok' } else { n = hd(names) if (is_journal_name(n) == 'true') { - path = sessions_dir() ++ "/" ++ n + path = dir ++ "/" ++ n + sig = file_sig(path) needs = if (path == active) { 'true' } - else { if (is_indexed(db, path) == 'true') { 'false' } + else { if (is_fresh(db, path, sig) == 'true') { 'false' } else { 'true' }} - if (needs == 'true') { index_one(db, path) } + if (needs == 'true') { index_one(db, path, sig) } } - reindex_loop(db, tl(names), active) + reindex_loop(db, dir, tl(names), active) } } @@ -88,15 +99,28 @@ fun is_journal_name(n) { string_ends_with(n, ".jsonl") == 'true' } -fun is_indexed(db, path) { - rows = db_query(db, "SELECT 1 FROM meta WHERE session = ?", [path]) - if (length(rows) == 0) { 'false' } else { 'true' } +# {size, mtime} of a journal right now ({-1, -1} if it vanished). +fun file_sig(path) { + st = file_stat(path) + if (st == nil) { {0 - 1, 0 - 1} } + else { {map_get(st, 'size'), map_get(st, 'mtime')} } +} + +# Indexed AND unchanged since: meta's size + mtime match the file's. +fun is_fresh(db, path, sig) { + rows = db_query(db, "SELECT size, mtime FROM meta WHERE session = ?", [path]) + if (length(rows) == 0) { 'false' } + else { + r = hd(rows) + if (map_get(r, "size") == elem(sig, 0) && map_get(r, "mtime") == elem(sig, 1)) { 'true' } + else { 'false' } + } } # Read the .active pointer to know which journal is the currently # running session (the only one that may have grown since last index). -fun read_active_marker() { - p = sessions_dir() ++ "/.active" +fun read_active_marker(dir) { + p = dir ++ "/.active" if (file_exists(p) == 'false') { nil } else { c = file_read(p) @@ -104,7 +128,10 @@ fun read_active_marker() { } } -fun index_one(db, path) { +# `sig` is the {size, mtime} observed BEFORE reading: if the journal +# grows while we ingest it, the recorded sig is older than the file and +# the next boot reindexes it again (never the reverse). +fun index_one(db, path, sig) { # swarmrt's db_exec() doesn't bind params — only db_query does. # Use db_query for parameterised writes; empty result is harmless. db_query(db, "DELETE FROM journals WHERE session = ?", [path]) @@ -114,8 +141,8 @@ fun index_one(db, path) { ingest_lines(db, path, clines) } db_query(db, - "INSERT OR REPLACE INTO meta(session, indexed_at) VALUES (?, ?)", - [path, timestamp()]) + "INSERT OR REPLACE INTO meta(session, indexed_at, size, mtime) VALUES (?, ?, ?, ?)", + [path, timestamp(), elem(sig, 0), elem(sig, 1)]) 'ok' } @@ -168,8 +195,10 @@ fun ingest_tool_calls(db, path, tcs) { # %{session, role, snippet} # `snippet()` wraps matched terms in >>><<<. # ------------------------------------------------------------ -fun search(query, limit) { - db = db_open(db_path()) +fun search(query, limit) { search_at(sessions_dir(), query, limit) } + +fun search_at(dir, query, limit) { + db = db_open(db_path_at(dir)) rows = db_query(db, "SELECT session, role, " ++ "snippet(journals, 2, '>>>', '<<<', '…', 40) AS snip " ++ diff --git a/src/TestRunner.sw b/src/TestRunner.sw index c25b228..4cf770f 100644 --- a/src/TestRunner.sw +++ b/src/TestRunner.sw @@ -1,10 +1,26 @@ module TestRunner -export [run_tests, parse_output, test_gate, format_result] +import Util + +export [run_tests, run_tests_timed, parse_output, test_gate, format_result, default_timeout_ms] + +# Default / maximum wall clock for a test run (overridable per call with the +# tool's timeout_ms arg, clamped to [1s, 600s]). +fun default_timeout_ms() { 300000 } # Run tests and return structured results -# shell() returns {exit_code, stdout} as a tuple fun run_tests(repo_path, command) { + run_tests_timed(repo_path, command, default_timeout_ms()) +} + +# The suite runs under shell_managed: its own process group, a C-enforced +# timeout that kills the whole tree (the old shell() call left a >120s suite +# running, orphaned, and reported "Exit code: -1" with no output), and ESC +# interrupts it. repo_path is single-quoted (a space in it broke `cd`, and a +# `;` injected a command), and the script is Util.noninteractive_wrap'd: +# stdin from /dev/null, stderr folded in, CI=1 exported, the command on its +# own lines. On timeout the partial output is kept and timed_out is 'true'. +fun run_tests_timed(repo_path, command, timeout_ms) { cmd = if (command == "") { detect_test_command(repo_path) } else { command } if (cmd == "") { # No explicit command and nothing recognizable in the repo. A clear hint @@ -17,13 +33,15 @@ fun run_tests(repo_path, command) { raw: "run_tests: could not auto-detect a test framework in " ++ repo_path ++ " (looked for package.json / Cargo.toml / go.mod / pyproject.toml / setup.py / pytest.ini / test_*.py). " ++ "Pass an explicit 'command', e.g. \"pytest -q\".", - exit_code: 0 + exit_code: 0, + timed_out: 'false' } } else { - full = "cd " ++ repo_path ++ " && " ++ cmd ++ " 2>&1" - result = shell(full) + full = Util.noninteractive_wrap("cd " ++ Util.shell_q(repo_path) ++ " || exit 2\n" ++ cmd) + result = shell_managed(full, timeout_ms) exit_code = elem(result, 0) stdout = elem(result, 1) + interrupted = elem(result, 2) parsed = parse_output(stdout) %{ @@ -35,7 +53,8 @@ fun run_tests(repo_path, command) { duration_ms: parsed.duration_ms, failures: parsed.failures, raw: stdout, - exit_code: exit_code + exit_code: exit_code, + timed_out: interrupted } } } diff --git a/src/ToolExecutor.sw b/src/ToolExecutor.sw index 2b53ae5..76e2a9b 100644 --- a/src/ToolExecutor.sw +++ b/src/ToolExecutor.sw @@ -98,7 +98,9 @@ fun prepare(name, args, opts) { } else { filesystem_hook = Hooks.run_pre_tool(name, args, opts) if (map_get(filesystem_hook, 'veto') == 'true') { - failed("tool '" ++ to_string(name) ++ "' blocked by pre_tool hook") + why = map_get(filesystem_hook, 'reason') + detail = if (why == nil) { "" } else { " — " ++ to_string(why) } + failed("tool '" ++ to_string(name) ++ "' blocked by pre_tool hook" ++ detail) } else { # Hooks may rewrite arguments, so every safety decision below # must inspect the effective arguments, never the originals. @@ -108,9 +110,10 @@ fun prepare(name, args, opts) { if (guard != 'ok') { failed(to_string(guard)) } else { - configured_hook = Config.run_hooks("PreToolUse", name, args_raw, opts) - if (configured_hook == 'block') { - failed("tool '" ++ to_string(name) ++ "' blocked by PreToolUse hook") + configured_hook = Config.run_hooks_verdict("PreToolUse", name, args_raw, opts) + if (configured_hook != 'ok') { + failed("tool '" ++ to_string(name) ++ "' blocked by PreToolUse hook — " ++ + to_string(elem(configured_hook, 1))) } else { %{ok: 'true', args: effective} } @@ -151,10 +154,12 @@ fun permission_gate(name, args, opts) { decision = Config.check_permission(name, args, opts) if (decision == 'allow') { 'ok' } else { if (decision == 'deny') { - "error: permission denied for tool '" ++ to_string(name) ++ "'" + Config.denial_message(name, args, opts) } else { + risk = Config.command_risk(name, args) + why = if (elem(risk, 0) == 'dangerous') { " — flagged dangerous: " ++ elem(risk, 1) } else { "" } "error: tool '" ++ to_string(name) ++ - "' requires interactive permission in this execution context" + "' requires interactive permission in this execution context" ++ why }} } diff --git a/src/ToolGuardrails.sw b/src/ToolGuardrails.sw index eea6989..867f019 100644 --- a/src/ToolGuardrails.sw +++ b/src/ToolGuardrails.sw @@ -11,6 +11,7 @@ module ToolGuardrails # exploration: # 1. identical-call: 5 consecutive calls with the same name+args # 2. same-tool-fail: 8 consecutive failing results from the same tool +# ("error: …", a non-zero "[exit N]", "[timed out …") # # The earlier "no-progress" check (5 consecutive idempotent reads # without a mutation) was removed after it kept firing on legitimate @@ -71,7 +72,7 @@ fun observe_after(opts, name_str, result_str) { table = map_get(opts, 'guardrails_table') if (table == nil) { 'ok' } else { - is_err = string_starts_with(to_string(result_str), "error:") + is_err = is_failure(to_string(result_str)) last_tool = ets_get(table, 'last_tool') fail_count = ets_get(table, 'fail_count') if (is_err == 'true') { @@ -94,6 +95,15 @@ fun observe_after(opts, name_str, result_str) { } } +# A failing tool result: the "error:" sentinel most tools use, a non-zero +# "[exit N]" banner (how bash / the auto-background path report failure — +# the brake never fired for bash before), or a "[timed out" banner. +fun is_failure(s) { + if (string_starts_with(s, "error:") == 'true') { 'true' } + else { if (string_starts_with(s, "[exit ") == 'true' && string_starts_with(s, "[exit 0]") == 'false') { 'true' } + else { string_starts_with(s, "[timed out") }} +} + # Clear all per-turn counters. Called at the start of each user # message in route_input — guardrails are per-turn, not per-session. fun reset(opts) { diff --git a/src/ToolRegistry.sw b/src/ToolRegistry.sw index 9586026..ee56a56 100644 --- a/src/ToolRegistry.sw +++ b/src/ToolRegistry.sw @@ -147,17 +147,23 @@ fun allowed_in(context, name) { if (in_list(subagent_blocked_tools(), to_string(name)) == 'true') { 'false' } else { 'true' } + } else { if (context == "subagent_explore") { + in_list(subagent_explore_tools(), to_string(name)) + } else { if (context == "subagent_bash") { + in_list(subagent_bash_tools(), to_string(name)) } else { if (context == "main") { 'true' } - else { 'false' }}}}} + else { 'false' }}}}}}} } fun names_for(context) { if (context == "mcp_server") { mcp_server_tools() } else { if (context == "council_panel") { council_panel_tools() } else { if (context == "council_judge") { [] } + else { if (context == "subagent_explore") { subagent_explore_tools() } + else { if (context == "subagent_bash") { subagent_bash_tools() } else { if (context == "main" || context == "subagent") { all_names(all_tools(), []) - } else { [] }}}} + } else { [] }}}}}} } fun all_names(entries, acc) { @@ -176,6 +182,15 @@ fun subagent_blocked_tools() { "bg_server", "browser_launch", "git_commit", "todo_write"] } +# Typed subagents (task tool subagent_type). The task schema promises +# "explore (read-only)" and "bash (shell only)" — these lists make that +# a policy instead of a prompt suggestion. explore gets the same +# read-only repository inspection set as a council panelist; no shell, +# writes, web, MCP (unknown side effects) or host-state tools. +fun subagent_explore_tools() { council_panel_tools() } + +fun subagent_bash_tools() { ["bash"] } + # Council panelists inspect the repository independently. Keep the first # prototype read-only and non-recursive: no shell, writes, background work, # memory mutation, browser control, or nested task delegation. diff --git a/src/ToolSchemas.sw b/src/ToolSchemas.sw index 964df0b..bceed9a 100644 --- a/src/ToolSchemas.sw +++ b/src/ToolSchemas.sw @@ -146,7 +146,7 @@ fun multi_edit_s() { fun glob_s() { tool("glob", "Find files matching a glob pattern (e.g. `src/**/*.sw`). Returns " ++ - "paths sorted by modification time, capped at 100 (a cap notice is " ++ + "relative paths (in no particular order), capped at 100 (a cap notice is " ++ "appended when more matched — narrow the pattern to see the rest). " ++ "Use for file discovery.", obj(%{ @@ -157,13 +157,15 @@ fun glob_s() { fun grep_s() { tool("grep", - "Search file contents with a regex. Returns file paths by default; " ++ - "set output_mode=content to see matching lines.", + "Search file contents with a regex (ripgrep syntax). Returns matching lines " ++ + "as `path:line:text` by default; set output_mode=files_with_matches for just " ++ + "the file paths, or count for per-file match counts. An invalid regex is " ++ + "reported as an error.", obj(%{ pattern: s("Regular expression"), path: s("Optional: file or directory to search"), glob: s("Optional: file glob filter (e.g. *.sw)"), - output_mode: s("'files_with_matches' | 'content' | 'count'"), + output_mode: s("'content' (default) | 'files_with_matches' | 'count'"), head_limit: i("Optional: cap results (default 100 lines; output truncated past ~16KB)") }, ["pattern"])) } @@ -309,7 +311,7 @@ fun task_s() { obj(%{ description: s("A short (3-5 word) task label, for the UI"), prompt: s("The full task instructions for the subagent"), - subagent_type: s("'general' (all tools) | 'explore' (read-only: read/glob/grep) | 'bash' (shell only)") + subagent_type: s("'general' (all tools but task/memory/skill/git_commit/background-server writes) | 'explore' (read-only, enforced: read/glob/grep/git_status/git_diff/code_search) | 'bash' (shell only, enforced)") }, ["description", "prompt"])) } @@ -428,7 +430,8 @@ fun run_tests_s() { "changes. If tests fail, fix them before committing.", obj(%{ repo_path: s("Absolute path to the repository root"), - command: s("Optional: test command to run (defaults to npm test, pytest, etc)") + command: s("Optional: test command to run (defaults to npm test, pytest, etc)"), + timeout_ms: i("Optional: wall-clock limit in ms (default 300000, max 600000); on timeout the process group is killed and the partial output returned") }, ["repo_path"])) } @@ -469,16 +472,17 @@ fun log_wait_s() { pattern: s("Regex/substring to wait for"), path: s("Optional: file to watch"), task_id: s("Optional: background task id to watch instead of a file"), - timeout_sec: i("Optional: max seconds to wait (default 30)") + timeout_sec: i("Optional: max seconds to wait (default 60, clamped to 1..600)") }, ["pattern"])) } fun file_watch_s() { tool("file_watch", - "Block until a file changes (or the timeout elapses).", + "Block until a file changes — its mtime or size changes, or it appears " ++ + "or disappears — or the timeout elapses.", obj(%{ path: s("File to watch"), - timeout_sec: i("Optional: max seconds to wait (default 30)") + timeout_sec: i("Optional: max seconds to wait (default 60, clamped to 1..600)") }, ["path"])) } diff --git a/src/agent.sw b/src/agent.sw index b5a986e..b237721 100644 --- a/src/agent.sw +++ b/src/agent.sw @@ -46,7 +46,11 @@ export [run, run_headless, subagent_blocked, SUBAGENT_BLOCKED_TOOLS, drop_to_last_clean_user, trailing_turn_incomplete, trim_incomplete, looks_like_slash_command, is_known_slash_command, bg_normalize_id, get_session_mode, set_session_mode, next_mode, resolve_permission, - show_expand, handle_bg_command, route_input] + show_expand, handle_bg_command, route_input, + skip_remaining_tools, turn_interrupted, + args_malformed, sanitize_tool_calls, turn_cut_reason, refuse_tool_calls, + headless_answer, compact_history, compact_split, mechanical_trim_ex, + profile_to_override] # Maximum tool-call rounds per user turn. fun max_steps() { 200 } @@ -54,13 +58,12 @@ fun max_steps() { 200 } # ------------------------------------------------------------ # Context budget — token-based, sourced from server usage # ------------------------------------------------------------ -fun max_tokens_env() { parse_env_int("SWARM_CODE_MAX_TOKENS", 262144) } -fun output_reserve_env() { parse_env_int("SWARM_CODE_OUTPUT_RESERVE", 16384) } -fun compact_buffer_env() { parse_env_int("SWARM_CODE_COMPACT_BUFFER", 52000) } - -fun context_budget_tokens() { - max_tokens_env() - output_reserve_env() - compact_buffer_env() -} +# ONE definition, in llm.sw (LLM.context_budget_tokens), shared with the +# context meter: the output reserve and compaction buffer scale down with +# the window, so a small SWARM_CODE_MAX_TOKENS no longer yields a NEGATIVE +# budget (32768 - 16384 - 52000 = -35616 compacted on every step). +fun max_tokens_env() { LLM.context_window_tokens() } +fun context_budget_tokens() { LLM.context_budget_tokens() } # When a compaction fires (over context_budget_tokens), trim/summarize down to # THIS lower level — not just barely under the trigger — so the next several tool @@ -74,15 +77,6 @@ fun context_budget_chars_fallback() { context_budget_tokens() * 4 } -fun parse_env_int(name, fallback) { - env = getenv(name) - if (env == nil) { fallback } - else { - parsed = parse_budget_env(env, 0, 0, 'false') - if (parsed < 0) { fallback } else { parsed } - } -} - fun parse_budget_env(s, i, acc, saw_digit) { if (i >= string_length(s)) { if (saw_digit == 'true') { acc } else { 0 - 1 } @@ -303,17 +297,43 @@ fun drop_trailing_tool_turn(msgs) { drop_last(msgs, tools_run + 1) } -# True when any tool_call carries a non-empty arguments blob that fails to -# json_decode (truncated mid-string) — the turn is poisoned and unsendable. +# True when any tool_call carries a non-empty arguments blob that is not one +# complete JSON value (truncated mid-string) — the turn is poisoned and unsendable. fun tcs_args_malformed(tcs) { if (length(tcs) == 0) { 'false' } else { tc = hd(tcs) - raw = to_string(map_get(tc, 'arguments')) - trimmed = string_trim(raw) - bad = if (trimmed == "" || trimmed == "{}" || trimmed == "null") { 'false' } - else { if (json_decode(raw) == nil) { 'true' } else { 'false' } } - if (bad == 'true') { 'true' } else { tcs_args_malformed(tl(tcs)) } + if (args_malformed(map_get(tc, 'arguments')) == 'true') { 'true' } + else { tcs_args_malformed(tl(tcs)) } + } +} + +# A NON-EMPTY arguments blob that is not one complete JSON object/array. +# json_decode alone can't tell: it is lenient and happily decodes +# `{"command":"echo hi` (cut mid-string) into %{command: "echo hi"}, so a +# truncated call used to pass as well-formed and run. The strict structural +# check (Util.json_args_well_formed) catches the unclosed string/container a +# truncation always leaves — even when no finish_reason says so. Empty / "{}" +# / "null" mean "no arguments", not malformed. +fun args_malformed(raw) { + trimmed = string_trim(to_string(raw)) + if (trimmed == "" || trimmed == "{}" || trimmed == "null") { 'false' } + else { if (json_decode(trimmed) == nil) { 'true' } + else { if (Util.json_args_well_formed(trimmed) == 'true') { 'false' } else { 'true' } } } +} + +# History copy of a turn's tool_calls with malformed arguments replaced by +# "{}". Servers that parse assistant tool_call arguments when rendering the +# chat template (vLLM does) reject every later request that still carries a +# cut-off blob, wedging the session. Dispatch uses the RAW calls, so the model +# is still told exactly what was wrong with each one. +fun sanitize_tool_calls(tcs, acc) { + if (length(tcs) == 0) { acc } + else { + tc = hd(tcs) + clean = if (args_malformed(map_get(tc, 'arguments')) == 'true') { map_put(tc, 'arguments', "{}") } + else { tc } + sanitize_tool_calls(tl(tcs), list_append(acc, clean)) } } @@ -325,11 +345,11 @@ fun nth_at(lst, i) { # ------------------------------------------------------------ # F2 — poison / clean-exit markers beside .active # ------------------------------------------------------------ -# A turn that ends on an UNRECOVERED llm error writes .poison (holding the -# journal path) so the next launch knows the recorded session died mid-turn -# and trims the failing turn back to the last clean user message instead of -# replaying the poisoned tool_call. A clean /quit writes .clean_exit. Both -# are cleared when a fresh session starts. +# A turn that ends on an UNRECOVERED fatal llm error writes .poison (holding +# the journal path) so the next launch can say the recorded session ended on +# a request error. (It used to also cut the resumed history back to the last +# user message, losing completed tool results — see run.) A clean /quit +# writes .clean_exit. Both are cleared when a fresh session starts. fun journal_poison_ptr() { session_dir() ++ "/.poison" } fun journal_clean_exit_ptr() { session_dir() ++ "/.clean_exit" } @@ -369,9 +389,10 @@ fun journal_is_poisoned(journal_path) { } # Drop trailing assistant/tool records back to (and including) the last -# role:'user' message — the last clean point. Used on resume of a poisoned -# session so the failing turn never re-fires. Keeps system + everything up -# to and including the last user message. +# role:'user' message — the last clean point. Keeps system + everything up +# to and including the last user message. (No longer applied on a fatal +# error or on resume: it threw away completed tool results whose effects +# were real — both keep completed pairs and trim only the dangling tail.) fun drop_to_last_clean_user(msgs) { idx = last_user_index(msgs, length(msgs) - 1) if (idx < 0) { msgs } @@ -421,12 +442,12 @@ fun run(opts, system_prompt_text) { replay_journal(prev_path) } else { [] } } - # F2: on a poisoned session, trim the failing turn back to the last clean - # user message (drop the poisoned tool_call) BEFORE the usual incomplete - # trim, so resume never re-fires the call that killed the prior session. - resumed_clean = if (prev_poisoned == 'true') { drop_to_last_clean_user(resumed_raw) } - else { resumed_raw } - resumed = trim_incomplete(resumed_clean) + # F2: a poisoned session (it ended on a fatal request error) resumes like + # any other — trim_incomplete drops only the dangling tail. It used to be + # cut back to the last user message, throwing away tool results whose + # effects (written files) are real; nothing in a completed pair can + # re-fire on resume, and cut-off tool arguments are stored as "{}". + resumed = trim_incomplete(resumed_raw) journal_path = if (length(resumed) > 0) { prev_path @@ -445,8 +466,9 @@ fun run(opts, system_prompt_text) { history = if (length(resumed) > 0) { print("") if (prev_poisoned == 'true') { - print(UI.grey_text() ++ " ⏺ previous session ended on an error — trimmed the " ++ - "failing turn; resumed " ++ to_string(length(resumed)) ++ " messages" ++ UI.reset()) + print(UI.grey_text() ++ " ⏺ previous session ended on a request error — resumed " ++ + to_string(length(resumed)) ++ " messages (completed tool results kept; " ++ + "/compact or /reset if the error repeats)" ++ UI.reset()) } else { print(UI.grey_text() ++ " ⏺ resumed crashed session — " ++ to_string(length(resumed)) ++ " messages recovered" ++ UI.reset()) @@ -519,35 +541,68 @@ fun run_headless(opts, system_prompt_text, prompt, json_mode) { } journal_sync(opts_journal, history) - # Auto-accept permissions in headless. - opts_h = map_put(opts_journal, 'headless', 'true') + # Headless: no Reader to ask — an 'ask' is denied unless + # SWARM_CODE_HEADLESS_APPROVE=1 (see resolve_permission). run_id stamps + # every assistant message this run appends (run_turn), so the result can + # only ever be THIS run's answer: resume is the headless default, and the + # resumed history already holds earlier runs' answers — a run whose LLM + # call failed (400, server down) used to report the previous run's reply + # as {"status":"ok"}, exit 0. + run_id = uuid() + opts_h = map_put(map_put(opts_journal, 'headless', 'true'), 'run_id', run_id) final_history = route_input(prompt, history, opts_h) journal_sync(opts_h, final_history) - last_text = last_assistant_text(final_history) - ok = if (string_length(last_text) > 0) { 'true' } else { 'false' } + last_text = headless_answer(final_history, run_id) + # A slash command (/compact, /profile …) is a successful run with no model + # answer — not a failure. + ok = if (string_length(last_text) > 0) { 'true' } + else { if (is_slash_prompt(prompt) == 'true') { 'true' } else { 'false' } } + result_fd = map_get(opts, 'result_fd') if (json_mode == 'true') { status = if (ok == 'true') { "ok" } else { "error" } - print(json_encode(%{status: status, summary: last_text})) - } else { "" } + emit_result(result_fd, json_encode(%{status: status, summary: last_text})) + } else { + # Diverted (piped) stdout gets the final answer; a terminal already + # showed it streamed. + if (ok == 'true' && result_fd != nil && result_fd >= 0) { emit_result(result_fd, last_text) } + else { "" } + } if (ok == 'true') { "" } else { sys_exit(1) } } -fun last_assistant_text(history) { - last_assistant_loop(history, "") +# Write the headless result line to the real stdout (result_fd, set when +# main diverted stdout to stderr), or plain print when nothing was diverted. +fun emit_result(result_fd, s) { + if (result_fd != nil && result_fd >= 0) { fd_write(result_fd, s ++ "\n") } + else { print(s) } } -fun last_assistant_loop(msgs, acc) { - if (length(msgs) == 0) { acc } +# The headless result: the content of the history's FINAL message when it is +# an assistant reply this run produced (stamped with run_id by run_turn) that +# asks for no further tools. Anything else — the turn ended on an LLM failure +# (history ends at the user message or a tool result), hit max steps, or +# appended nothing — is "" (status error), never an earlier run's answer. +# (Was: the last assistant message ANYWHERE in the resumed history.) +fun headless_answer(history, run_id) { + if (length(history) == 0 || run_id == nil) { "" } else { - msg = hd(msgs) - role = map_get(msg, 'role') - na = if (role == 'assistant') { to_string(map_get(msg, 'content')) } else { acc } - last_assistant_loop(tl(msgs), na) + last = hd(take_last(history, 1)) + tcs = map_get(last, 'tool_calls') + if (map_get(last, 'role') == 'assistant' && map_get(last, 'run_id') == run_id && + (tcs == nil || length(tcs) == 0)) { + to_string(map_get(last, 'content')) + } else { "" } } } +fun is_slash_prompt(prompt) { + t = string_trim(to_string(prompt)) + if (string_starts_with(t, "/") == 'true' && is_known_slash_command(first_token(t)) == 'true') { 'true' } + else { 'false' } +} + # ------------------------------------------------------------ # The continuous loop. # ------------------------------------------------------------ @@ -938,14 +993,17 @@ fun slash_dispatch(cmd, history, opts) { else { if (cmd == "/model") { show_model_info(opts) ; history } else { if (string_starts_with(cmd, "/model ") == 'true') { new_model = string_trim(string_sub(cmd, 7, string_length(cmd) - 7)) - apply_model_override(new_model) + apply_model_override(new_model, opts) history } - else { if (cmd == "/profile") { show_active_profile() ; history } + else { if (cmd == "/profile") { show_active_profile(opts) ; history } else { if (cmd == "/profiles") { list_profiles() ; history } else { if (string_starts_with(cmd, "/profile ") == 'true') { name = string_trim(string_sub(cmd, 9, string_length(cmd) - 9)) - apply_profile_override(name) + # "/profile clear" is the documented alias of /profile-clear (this + # prefix branch used to swallow it as a profile named "clear"). + if (name == "clear") { clear_profile_override() } + else { apply_profile_override(name, opts) } history } else { if (cmd == "/profile-clear" || cmd == "/profile clear") { @@ -997,11 +1055,12 @@ fun slash_dispatch(cmd, history, opts) { } else { expr = hd(parsed) prompt = hd(tl(parsed)) - id = Scheduler.add(expr, prompt) - if (id == nil) { - print(UI.warn_text("invalid EXPR — try 30s, 5m, 2h, 1d, hourly, daily 09:00")) + r = Scheduler.add_checked(expr, prompt) + if (elem(r, 0) != 'ok') { + print(UI.warn_text(to_string(elem(r, 1)) ++ + " — EXPR is 30s, 5m, 2h, 1d, hourly or daily HH:MM")) } else { - print(UI.brand_color() ++ "✓ scheduled job " ++ id ++ + print(UI.brand_color() ++ "✓ scheduled job " ++ to_string(elem(r, 1)) ++ " (" ++ expr ++ "): " ++ prompt ++ UI.reset()) } } @@ -1058,8 +1117,11 @@ fun slash_dispatch(cmd, history, opts) { } else { if (cmd == "/compact") { compacted = compact_history(history, opts) - print("\e[2m[history compacted: " ++ to_string(length(history)) ++ - " → " ++ to_string(length(compacted)) ++ " messages]\e[0m") + # Unchanged (nothing old enough / summarizer failed) already said why. + if (length(compacted) != length(history)) { + print("\e[2m[history compacted: " ++ to_string(length(history)) ++ + " → " ++ to_string(length(compacted)) ++ " messages]\e[0m") + } compacted } else { if (cmd == "/save") { @@ -1158,6 +1220,10 @@ fun slash_dispatch(cmd, history, opts) { handle_bg_command(cmd, opts) history } + else { if (cmd == "/trust" || cmd == "/untrust") { + slash_trust(cmd, opts) + history + } else { if (cmd == "/mode") { nxt = cycle_mode(opts) print(UI.brand_color() ++ "✓ mode: " ++ nxt ++ UI.reset()) @@ -1181,7 +1247,26 @@ fun slash_dispatch(cmd, history, opts) { else { print(UI.warn_text("unknown command: " ++ cmd) ++ " (type /help)") history - }}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}} + }}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}} +} + +# /trust, /untrust — the REPL form of `swarm-code trust` for the session's +# directory. Settings are read at launch, so it takes effect next start. +fun slash_trust(cmd, opts) { + add = if (cmd == "/trust") { 'true' } else { 'false' } + dir = to_string(map_get(opts, 'cwd')) + r = Config.set_trusted(dir, add) + if (elem(r, 0) != 'ok') { print(UI.warn_text(to_string(elem(r, 1)))) } + else { + verb = if (add == 'true') { "trusted " } else { "untrusted " } + state = if (elem(r, 2) == 'true') { verb } else { + if (add == 'true') { "already trusted: " } else { "not trusted: " } } + print(UI.brand_color() ++ "✓ " ++ state ++ to_string(elem(r, 1)) ++ UI.reset()) + if (elem(r, 2) == 'true') { + print(UI.grey_text() ++ " restart swarm-code to load this repo's .swarm-code.json " ++ + (if (add == 'true') { "in full" } else { "with only its safe keys" }) ++ UI.reset()) + } else { 'ok' } + } } # /expand — reprint the most recent tool result in full (uncapped), @@ -1230,6 +1315,7 @@ fun show_help() { print(" /bg [tail|kill] [id] list / tail / kill background tasks") print(" /expand reprint the last tool result in full") print(" /mode cycle permission mode (default → auto-accept-edits → plan)") + print(" /trust, /untrust let this repo's .swarm-code.json apply in full (next launch)") print(" /clear clear screen") print(" /reset clear conversation history") print(" /compact summarize history to save context") @@ -1347,14 +1433,16 @@ fun show_debug_tools() { # ------------------------------------------------------------ # Profile / model swap — writes ~/.swarm-code/.profile_override which -# LLM.apply_override consults before every request. No opts threading -# required; the change takes effect on the next LLM call. +# LLM.apply_override consults before every request; the change takes effect +# on the next LLM call. Each write is stamped with this launch's session_id: +# in THIS session the override beats env vars, in a later one an env var +# that is set wins (see LLM.apply_override). # ------------------------------------------------------------ fun profile_override_path() { getenv("HOME") ++ "/.swarm-code/.profile_override" } -fun apply_profile_override(name) { +fun apply_profile_override(name, opts) { if (string_length(name) == 0) { print(UI.warn_text("usage: /profile NAME (try /profiles to list)")) } @@ -1370,7 +1458,7 @@ fun apply_profile_override(name) { print(UI.warn_text("no profile named '" ++ name ++ "' — try /profiles")) } else { - ov = profile_to_override(p) + ov = stamp_override(map_put(profile_to_override(p), 'profile', name), opts) file_write(profile_override_path(), json_encode(ov)) print(UI.brand_color() ++ "✓ profile swapped to " ++ name ++ UI.reset()) print(UI.grey_text() ++ " model : " ++ to_string(map_get(ov, 'model')) ++ UI.reset()) @@ -1381,18 +1469,28 @@ fun apply_profile_override(name) { } } -fun apply_model_override(model_name) { +# Changes ONLY the model: whatever part of an existing override is in effect +# (a /profile's endpoint, api_key, tool_format, chat_template_kwargs) is +# carried over — it used to be overwritten with {model}, silently reverting +# the profile while printing "endpoint/api_key unchanged". +fun apply_model_override(model_name, opts) { if (string_length(model_name) == 0) { print(UI.warn_text("usage: /model NAME")) } else { - ov = %{ model: model_name } + kept = LLM.effective_override(LLM.read_override(), opts) + ov = stamp_override(map_put(kept, 'model', model_name), opts) file_write(profile_override_path(), json_encode(ov)) print(UI.brand_color() ++ "✓ model swapped to " ++ model_name ++ UI.reset()) print(UI.grey_text() ++ " endpoint/api_key unchanged. /profile-clear to revert." ++ UI.reset()) } } +fun stamp_override(ov, opts) { + sid = map_get(opts, 'session_id') + if (sid == nil) { ov } else { map_put(ov, 'session_id', to_string(sid)) } +} + fun clear_profile_override() { p = profile_override_path() if (file_exists(p) == 'true') { @@ -1403,32 +1501,38 @@ fun clear_profile_override() { } } -fun show_active_profile() { +fun show_active_profile(opts) { p = profile_override_path() if (file_exists(p) == 'false') { print("\e[2m(no override active — using launch profile)\e[0m") } else { - raw = file_read(p) - ov = if (raw == nil) { nil } else { json_decode(raw) } + ov = LLM.read_override() if (ov == nil) { print(UI.warn_text("(override file present but unreadable: " ++ p ++ ")")) } else { - print("\e[1mactive override\e[0m") - show_override_field(ov, 'endpoint') - show_override_field(ov, 'model') - show_override_field(ov, 'api_key') - show_override_field(ov, 'tool_format') + prof = map_get(ov, 'profile') + print("\e[1mactive override\e[0m" ++ (if (prof == nil) { "" } else { " (profile " ++ to_string(prof) ++ ")" })) + show_override_field(ov, 'endpoint', opts) + show_override_field(ov, 'model', opts) + show_override_field(ov, 'api_key', opts) + show_override_field(ov, 'tool_format', opts) + show_override_field(ov, 'chat_template_kwargs', opts) print(UI.grey_text() ++ " (use /profile-clear to revert)" ++ UI.reset()) } } } -fun show_override_field(ov, key) { +# A field an env var pins (override left by an earlier session) is shown as +# not applied, so /profile tells the truth about what gets sent. +fun show_override_field(ov, key, opts) { v = map_get(ov, key) if (v == nil) { 'skip' } else { shown = if (key == 'api_key') { "(set)" } else { to_string(v) } - print(" " ++ to_string(key) ++ " : " ++ shown) + applied = LLM.override_applies(opts, ov, key, fn(name) { getenv(name) }) + note = if (applied == 'true') { "" } + else { UI.grey_text() ++ " (not applied — the environment sets it)" ++ UI.reset() } + print(" " ++ to_string(key) ++ " : " ++ shown ++ note) } } @@ -1482,14 +1586,17 @@ fun lookup_profile_loop(keys, values, name) { # Convert a settings.json profile entry into the override-file shape. # Only includes fields actually present in the profile — `apply_override` -# will leave unset fields alone. +# will leave unset fields alone. chat_template_kwargs is a map and is kept +# as one (it used to be omitted, and apply_override then cleared the launch +# value — `/profile qwen2` lost enable_thinking:false). fun profile_to_override(p) { a = profile_field(map_new(), p, 'endpoint') b = profile_field(a, p, 'model') c = profile_field(b, p, 'api_key') d = profile_field(c, p, 'tool_format') e = profile_field(d, p, 'vision') - e + ct = map_get(p, 'chat_template_kwargs') + if (ct == nil) { e } else { map_put(e, 'chat_template_kwargs', ct) } } fun profile_field(acc, p, key) { @@ -1582,44 +1689,119 @@ fun approx_tokens(history) { fun sum_msg_chars(msgs, acc) { history_chars_loop(msgs, acc) } # ------------------------------------------------------------ -# Compaction — summarize oldest messages, keep system + last 16. +# Compaction — summarize the old part, keep the live turn verbatim. # ------------------------------------------------------------ -fun compact_history(history, opts) { - if (length(history) < 10) { history } - else { - sys_msg = hd(history) - rest = tl(history) - keep_tail = take_last(rest, 16) - to_summarize = drop_last(rest, 16) - - summary_prompt = - "Summarize the conversation below to save context. Output EXACTLY " ++ - "these four sections, each 1-3 sentences:\n\n" ++ - "**Current State**: What has been accomplished so far.\n" ++ - "**Working State**: Files modified, commands run, tools used.\n" ++ - "**Key Details**: Important technical decisions, paths, names, configs.\n" ++ - "**Pending**: What still needs to be done, open questions.\n\n" ++ - "Be dense and precise. Preserve exact file paths, function names, " ++ - "and error messages — they are needed for continuity. Under 500 words total.\n\n" ++ - format_for_summary(to_summarize, "") +# Result: [system, summary, ...tail]. The tail is kept verbatim: at most the +# last COMPACT_KEEP_TAIL() messages, pulled back so it always starts at or +# before the MOST RECENT user message (mid-turn compaction used to summarize +# the live request away — requests then went out with no user message at +# all) and never on a role:'tool' result cut off from its assistant. +# Invariants (each was a bug): +# - nothing old enough to summarize → history unchanged, NO LLM call (it +# used to summarize an empty transcript and prepend yet another summary); +# - summarizer failed / said nothing → history unchanged (it used to delete +# the old messages anyway and journal "[compaction failed, messages +# elided]"); +# - an earlier summary is folded into the new one, not stacked. +fun SUMMARY_PREFIX() { "Summary of earlier conversation: " } +fun COMPACT_KEEP_TAIL() { 16 } +fun compact_history(history, opts) { + has_sys = if (length(history) > 0 && map_get(hd(history), 'role') == 'system') { 'true' } else { 'false' } + head = if (has_sys == 'true') { [hd(history)] } else { [] } + rest = if (has_sys == 'true') { tl(history) } else { history } + k = compact_split(rest) + old = take_first(rest, k, []) + keep_tail = drop_first_n(rest, k) + prev_summary = summary_texts(old, "") + fresh = non_summary_msgs(old, []) + if (length(fresh) == 0) { + turn_print(opts, " " ++ UI.dim_text("(nothing old enough to compact — the current turn is kept verbatim)")) + history + } else { ask_msgs = [ LLM.new_message_system("You are a concise summarizer."), - LLM.new_message_user(summary_prompt) + LLM.new_message_user(compact_prompt(prev_summary, fresh)) ] summary = LLM.chat_silent(ask_msgs, opts) - summary_text = if (summary == nil) { - "[compaction failed, messages elided]" - } else { to_string(summary) } + summary_text = if (summary == nil) { "" } else { string_trim(to_string(summary)) } + if (string_length(summary_text) == 0) { + turn_print(opts, " " ++ UI.warn_text("(compaction failed — the summarizer returned nothing; history kept as is)")) + history + } else { + synth = LLM.new_message_assistant(SUMMARY_PREFIX() ++ summary_text, nil, nil) + head ++ [synth] ++ keep_tail + } + } +} + +fun compact_prompt(prev_summary, fresh) { + earlier = if (string_length(prev_summary) == 0) { "" } + else { + "An EARLIER SUMMARY covers what came before the messages below. Fold " ++ + "its facts into your new summary — do not drop them:\n" ++ prev_summary ++ "\n\n" ++ + "Messages since that summary:\n" + } + "Summarize the conversation below to save context. Output EXACTLY " ++ + "these four sections, each 1-3 sentences:\n\n" ++ + "**Current State**: What has been accomplished so far.\n" ++ + "**Working State**: Files modified, commands run, tools used.\n" ++ + "**Key Details**: Important technical decisions, paths, names, configs.\n" ++ + "**Pending**: What still needs to be done, open questions.\n\n" ++ + "Be dense and precise. Preserve exact file paths, function names, " ++ + "and error messages — they are needed for continuity. Under 500 words total.\n\n" ++ + earlier ++ format_for_summary(fresh, "") +} + +# Index in `rest` (history minus the system message) where the verbatim tail +# starts; everything before it gets summarized. 0 = nothing to summarize. +fun compact_split(rest) { + n = length(rest) + k0 = if (n > COMPACT_KEEP_TAIL()) { n - COMPACT_KEEP_TAIL() } else { 0 } + lu = last_user_index(rest, n - 1) + k1 = if (lu >= 0 && lu < k0) { lu } else { k0 } + pair_start(rest, k1) +} + +# Step back over role:'tool' results so the tail opens on the assistant that +# issued them (a tool message without its tool_call is invalid on the wire). +fun pair_start(rest, k) { + if (k <= 0) { 0 } + else { if (map_get(nth_at(rest, k), 'role') == 'tool') { pair_start(rest, k - 1) } else { k } } +} - synth = LLM.new_message_assistant( - "Summary of earlier conversation: " ++ summary_text, - nil, nil) +fun is_summary_msg(m) { + if (map_get(m, 'role') == 'assistant' && + string_starts_with(to_string(map_get(m, 'content')), SUMMARY_PREFIX()) == 'true') { 'true' } + else { 'false' } +} - prepend([sys_msg, synth], keep_tail) +# The text of every earlier summary in `msgs` (older builds stacked several). +fun summary_texts(msgs, acc) { + if (length(msgs) == 0) { acc } + else { + m = hd(msgs) + if (is_summary_msg(m) == 'true') { + c = to_string(map_get(m, 'content')) + t = string_sub(c, string_length(SUMMARY_PREFIX()), string_length(c) - string_length(SUMMARY_PREFIX())) + summary_texts(tl(msgs), if (string_length(acc) == 0) { t } else { acc ++ "\n\n" ++ t }) + } else { summary_texts(tl(msgs), acc) } + } +} + +fun non_summary_msgs(msgs, acc) { + if (length(msgs) == 0) { acc } + else { + m = hd(msgs) + non_summary_msgs(tl(msgs), if (is_summary_msg(m) == 'true') { acc } else { list_append(acc, m) }) } } +fun drop_first_n(lst, n) { + if (n <= 0 || length(lst) == 0) { lst } + else { drop_first_n(tl(lst), n - 1) } +} + fun take_last(lst, n) { if (length(lst) <= n) { lst } else { take_last(tl(lst), n) } @@ -1862,17 +2044,28 @@ fun tcs_chars(tcs, acc) { # → snip before any LLM auto-compaction. Walk OLDEST-first; for each role:'tool' # or role:'user' message whose content exceeds ~8KB, replace the body with a # "[N chars elided]" stub. Stop as soon as approx_tokens(history) < budget. -# Never touches system, assistant prose, tool_calls, or the LAST message (the -# live user turn). Needs no network — works even when the LLM is unreachable. +# Never touches system, assistant prose, tool_calls, the MOST RECENT user +# message (the live request — mid-turn it is no longer the last message, and a +# pasted 10KB request used to get stubbed), or the LAST message (the result the +# model is about to read). Needs no network — works even when the LLM is +# unreachable. fun MECH_TRIM_THRESHOLD_CHARS() { 8000 } -fun mechanical_trim(history, budget_t) { +fun mechanical_trim(history, budget_t) { mechanical_trim_ex(history, budget_t, 'true') } + +# protect_last 'false': the overflow retry (shrink_for_overflow) — the server +# already refused a request containing that last result (a giant `read`), so +# it may be stubbed too; the live user request stays protected. +fun mechanical_trim_ex(history, budget_t, protect_last) { n = length(history) if (n == 0) { history } - else { mech_trim_loop(history, 0, n, budget_t, []) } + else { + live_user = last_user_index(history, n - 1) + mech_trim_loop(history, 0, n, budget_t, [], live_user, protect_last) + } } -fun mech_trim_loop(msgs, i, total, budget_t, acc) { +fun mech_trim_loop(msgs, i, total, budget_t, acc, live_user, protect_last) { if (length(msgs) == 0) { acc } else { m = hd(msgs) @@ -1884,15 +2077,27 @@ fun mech_trim_loop(msgs, i, total, budget_t, acc) { acc ++ msgs } else { role = map_get(m, 'role') - is_last = if (i == total - 1) { 'true' } else { 'false' } - stubbable = if (is_last == 'true') { 'false' } + protected = if (i == live_user) { 'true' } + else { if (protect_last == 'true' && i == total - 1) { 'true' } else { 'false' } } + stubbable = if (protected == 'true') { 'false' } else { if (role == 'tool' || role == 'user') { 'true' } else { 'false' }} new_m = if (stubbable == 'true') { stub_if_large(m) } else { m } - mech_trim_loop(tl(msgs), i + 1, total, budget_t, list_append(acc, new_m)) + mech_trim_loop(tl(msgs), i + 1, total, budget_t, list_append(acc, new_m), live_user, protect_last) } } } +# The server rejected the request as longer than its context. Aim well below +# what was sent — half of it (or the usual compaction target, if lower): +# mechanical trim first (no LLM call; the last result may go too), then +# summarize whatever is old enough if that wasn't enough. +fun shrink_for_overflow(history, opts) { + half = approx_tokens(history) / 2 + target = if (half < compact_target_tokens()) { half } else { compact_target_tokens() } + trimmed = mechanical_trim_ex(history, target, 'false') + if (approx_tokens(trimmed) > target) { compact_history(trimmed, opts) } else { trimmed } +} + # Replace an oversized string body with a stub; small or non-string (multimodal # list) content is left untouched so image blocks never get mangled. fun stub_if_large(m) { @@ -1975,30 +2180,53 @@ fun run_turn(history, opts, step) { # is signalled by exit 1 / the --json status. (Transport detail is # on stderr via the runtime.) Interactive keeps the inline error. if (map_get(opts, 'headless') != 'true') { turn_print(opts, UI.err_text("[error] llm call failed")) } - # F2/F1: distinguish a POISONED context from a TRANSIENT failure. - # - FATAL (a 4xx request rejection or an unparseable body — re-sending - # the identical bytes can't fix it): drop back to the last clean user - # message so the offending tool_call is never journaled/re-fired, and - # flag the journal poisoned so even an immediate resume trims it. - # - TRANSIENT (network blip / 5xx after the retry budget): the context - # is fine, the wire failed. KEEP all completed work; only trim a - # dangling unsendable tail. Do NOT poison — a later resume is valid. + # KEEP every completed assistant+tool pair; only the dangling tail the + # API can't take (a trailing assistant with unanswered tool_calls, a + # partial result set) is dropped — fatal and transient alike. A FATAL + # rejection (4xx / unparseable body) used to drop the turn back to its + # user message: two writes landed on disk, a 400 followed, and after + # resume the model had no record of them. (Cut-off tool arguments — + # the old reason to drop — are stored as "{}" now; see + # sanitize_tool_calls.) A fatal failure still flags the journal + # poisoned, which now only changes the resume notice. + # + # "The context is too long" (context_length_exceeded, "maximum context + # length", …) is recoverable: the server's window can be smaller than + # SWARM_CODE_MAX_TOKENS claims. Shrink — mechanical trim first, then + # summarize — and retry ONCE per turn ('ctx_retry'), if that actually + # made the history smaller. fatal = LLM.last_fail(opts) - recovered = if (fatal == 'fatal') { - mark_poison(opts) - drop_to_last_clean_user(working_hist) + recovered = trim_incomplete(working_hist) + overflow = if (fatal == 'fatal' && map_get(opts, 'ctx_retry') != 'true') { + LLM.last_fail_context_overflow(opts) + } else { 'false' } + shrunk = if (overflow == 'true') { shrink_for_overflow(recovered, opts) } else { recovered } + if (overflow == 'true' && approx_tokens(shrunk) < approx_tokens(recovered)) { + turn_print(opts, " " ++ UI.dim_text("(the server says the context is too long — trimmed to ~" ++ + to_string(approx_tokens(shrunk)) ++ " tokens, retrying once)")) + journal_sync(opts, shrunk) + run_turn(shrunk, map_put(opts, 'ctx_retry', 'true'), step + 1) } else { - trim_incomplete(working_hist) + if (fatal == 'fatal') { mark_poison(opts) } + journal_sync(opts, recovered) + recovered } - journal_sync(opts, recovered) - recovered } else { content = to_string(map_get(result, 'content')) tool_calls_v = map_get(result, 'tool_calls') reasoning = map_get(result, 'reasoning') tool_calls = if (tool_calls_v == nil) { [] } else { tool_calls_v } - - asst_msg = LLM.new_message_assistant(content, tool_calls, reasoning) + # Was the turn cut off — ESC mid-stream, or the output-token limit? + cut = turn_cut_reason(result) + + # History keeps a wire-safe copy of the calls (cut-off arguments + # stored as "{}", see sanitize_tool_calls); dispatch below uses the + # raw ones. + # A headless run stamps its replies with run_id (see run_headless; + # the journal doesn't keep the field). + asst_msg0 = LLM.new_message_assistant(content, sanitize_tool_calls(tool_calls, []), reasoning) + run_id = map_get(opts, 'run_id') + asst_msg = if (run_id == nil) { asst_msg0 } else { map_put(asst_msg0, 'run_id', run_id) } with_assistant = list_append(working_hist, asst_msg) journal_sync(opts, with_assistant) # F2: a turn completed cleanly — clear any poison flag a PRIOR turn in @@ -2008,17 +2236,35 @@ fun run_turn(history, opts, step) { # the good work. clear_poison() is a cheap stat+unlink, idempotent. clear_poison() + # A cut-off turn's tool calls are NEVER dispatched. Their arguments + # may end mid-string, and the lenient json_decode turns that into a + # plausible call — a `write` of half a config file ("ok: wrote 59 + # bytes"), a `bash` command missing its tail. Each call gets a result + # saying it was not run (every tool_call id still needs its tool + # message), then: + # interrupted — the user pressed ESC while it streamed: the turn + # ends, exactly like ESC on a running tool. + # truncated — F4 recovery below (raise max_tokens, then the + # smaller-edits nudge), which tells the model to + # reissue the calls. # F4: length-truncation recovery (finish_reason=length / truncation - # marker) — ONLY when the turn carried no tool_calls. A truncated turn - # that still emitted tool_calls is actionable: fall through and execute - # them normally (appending a user nudge after an assistant-with-open- - # tool_calls would be an invalid native sequence and would skip the - # work). Stage 0: retry ONCE with a raised per-turn max_tokens. Stage 1: - # inject a "continue in smaller append-mode edits" user-message while - # KEEPING the partial assistant output. Stage 2+: give up gracefully. - # Guarded by the 'trunc_retry' counter so we never loop unbounded. - truncated = map_get(result, 'truncated') - if (truncated == 'true' && length(tool_calls) == 0) { + # marker). Stage 0: retry ONCE with a raised per-turn max_tokens. + # Stage 1: inject a "continue in smaller append-mode edits" + # user-message while KEEPING the partial assistant output. Stage 2+: + # give up gracefully. Guarded by the 'trunc_retry' counter so we never + # loop unbounded. The nudge follows the tool results, so the native + # sequence stays valid. + if (length(tool_calls) > 0 && cut != 'complete') { + refused = refuse_tool_calls(tool_calls, with_assistant, cut, opts) + if (cut == 'interrupted') { + journal_sync(opts, refused) + turn_print(opts, " " ++ UI.dim_text("⎿ interrupted · tell the agent what to do instead")) + turn_print(opts, "") + refused + } else { + handle_truncation(refused, tool_calls, opts, step) + } + } else { if (cut == 'truncated') { handle_truncation(with_assistant, tool_calls, opts, step) } else { @@ -2071,9 +2317,15 @@ fun run_turn(history, opts, step) { } else { meter_opts = opts_with_history(opts, with_assistant) post_exec = execute_all(tool_calls, with_assistant, meter_opts) - run_turn(post_exec, meter_opts, step + 1) - } + if (turn_interrupted(post_exec, length(tool_calls)) == 'true') { + turn_print(opts, " " ++ UI.dim_text("⎿ interrupted · tell the agent what to do instead")) + turn_print(opts, "") + post_exec + } else { + run_turn(post_exec, meter_opts, step + 1) + } } + }} } } } @@ -2098,6 +2350,9 @@ fun raised_max_tokens(opts) { if (doubled > ceil) { ceil } else { doubled } } +# `tool_calls` non-empty: the cut came while the model was still writing +# tool calls — refuse_tool_calls already answered each one "not run", so the +# nudge asks for them again instead of "continue where you left off". fun handle_truncation(with_assistant, tool_calls, opts, step) { stage = map_get(opts, 'trunc_retry') stage_n = if (stage == nil) { 0 } else { stage } @@ -2109,8 +2364,14 @@ fun handle_truncation(with_assistant, tool_calls, opts, step) { "max_tokens=" ++ to_string(raised) ++ ")")) retry_opts0 = map_put(opts, 'max_tokens', raised) retry_opts = map_put(retry_opts0, 'trunc_retry', 1) - cont = "Your previous response was cut off at the output token limit. " ++ - "Continue exactly where you left off." + cont = if (length(tool_calls) > 0) { + "Your previous response was cut off at the output token limit while it " ++ + "was still writing tool calls, so none of them ran. Reissue them with " ++ + "complete arguments." + } else { + "Your previous response was cut off at the output token limit. " ++ + "Continue exactly where you left off." + } with_nudge = list_append(with_assistant, LLM.new_message_user(cont)) journal_sync(opts, with_nudge) run_turn(with_nudge, retry_opts, step + 1) @@ -2457,9 +2718,20 @@ fun tool_crash_msg(name, reason) { # ------------------------------------------------------------ # Subagent via the Task tool — synchronous in-process loop. # ------------------------------------------------------------ +# Isolation rules (each a reviewed bug): +# * a FRESH ToolGuardrails table per subagent — sharing the parent's +# let a subagent's 8 failing reads set the parent's halt_reason, so +# the PARENT's turn halted and never saw the subagent's result. The +# subagent's own halt is checked inside run_subagent_loop. +# * per-type tool allow-lists are ENFORCED (execution_context + +# subagent_exec_all), not just promised in the prompt: an explore +# subagent could run bash and write. +# * the result is bounded (subagent_result_cap) and never lossy: every +# abnormal stop (max steps, guardrail halt, LLM failure) returns what +# the subagent had found so far plus the reason. fun handle_task_tool(args, opts) { prompt = map_get(args, 'prompt') - stype = map_get(args, 'subagent_type', "general") + stype = subagent_type_of(map_get(args, 'subagent_type')) if (prompt == nil) { "error: task tool requires a 'prompt' argument" @@ -2467,28 +2739,55 @@ fun handle_task_tool(args, opts) { sub_sys = subagent_system_prompt(stype) sub_history = [ LLM.new_message_system(sub_sys), - LLM.new_message_user(prompt) + LLM.new_message_user(to_string(prompt)) ] + guard = ToolGuardrails.init() sub_opts0 = map_put(opts, 'is_subagent', 'true') - sub_opts = map_put(sub_opts0, 'execution_context', "subagent") + sub_opts1 = map_put(sub_opts0, 'execution_context', subagent_context(stype)) + sub_opts2 = map_put(sub_opts1, 'subagent_type', stype) + sub_opts = map_put(sub_opts2, 'guardrails_table', guard) result = run_subagent_loop(sub_history, sub_opts, 0) - "[subagent:" ++ stype ++ "]\n" ++ result + # ETS tables are a finite resource (1024 live) — one per task call + # would leak without this. + ets_drop(guard) + "[subagent:" ++ stype ++ "]\n" ++ subagent_cap(result) } } +# Normalise the model-supplied type: only "explore" and "bash" narrow +# the tool set; anything else (missing, null, 5, "General") is general. +fun subagent_type_of(v) { + t = if (v == nil) { "general" } else { string_lower(string_trim(to_string(v))) } + if (t == "explore" || t == "bash") { t } else { "general" } +} + +# ToolExecutor enforces the context allow-list on every dispatch, so a +# typed subagent's restriction holds even outside subagent_exec_all. +fun subagent_context(stype) { + if (stype == "explore") { "subagent_explore" } + else { if (stype == "bash") { "subagent_bash" } + else { "subagent" } } +} + fun subagent_max_steps() { 15 } +# Same cap as bash / MCP output: a subagent's answer is a tool result +# the parent re-sends on every later turn (a 320KB answer reached the +# parent uncapped). +fun subagent_result_cap() { 24000 } + fun subagent_system_prompt(stype) { base = "You are a focused subagent spawned from swarm-code for a single " ++ "task. Use tools as needed, then return a concise final answer." if (stype == "explore") { - base ++ "\n\nAllowed tools: read, glob, grep. Do NOT call bash, write, " ++ - "edit, multi_edit, or web_fetch. Your job is to survey the " ++ + base ++ "\n\nYou are READ-ONLY. Allowed tools: " ++ + subagent_join(ToolRegistry.names_for("subagent_explore"), "") ++ + ". Any other tool call is refused. Your job is to survey the " ++ "codebase and report findings." } else { if (stype == "bash") { - base ++ "\n\nAllowed tool: bash. Your job is to run shell commands and " ++ - "report their output." + base ++ "\n\nAllowed tool: bash (any other tool call is refused). Your " ++ + "job is to run shell commands and report their output." } else { base ++ "\n\nYou have the full swarm-code tool set. Be decisive and finish " ++ @@ -2496,50 +2795,226 @@ fun subagent_system_prompt(stype) { }} } +fun subagent_join(names, acc) { + if (length(names) == 0) { acc } + else { + sep = if (string_length(acc) == 0) { "" } else { ", " } + subagent_join(tl(names), acc ++ sep ++ to_string(hd(names))) + } +} + fun subagent_llm_worker(token, parent, history, opts) { result = LLM.chat(history, opts) send(parent, {'llm_result', token, result}) } -fun subagent_await_llm(token, deadline) { +# One LLM round for the subagent, in a MONITORED worker (a crash +# surfaces as a reason instead of a silent 300s wait) bounded by the +# configured LLM timeout (Config.llm_timeout_ms: SWARM_CODE_LLM_TIMEOUT_MS +# → settings llm_timeout_ms → 300s) rather than a hard-coded 300s. +# Returns {'ok', result_map} | {'error', reason_string}. +fun subagent_llm_call(history, opts) { + token = to_string(self()) ++ "/" ++ to_string(timestamp()) ++ "/" ++ + to_string(random_int(1, 1000000000)) + {w, ref} = spawn_monitor(subagent_llm_worker(token, self(), history, opts)) + timeout_ms = Config.llm_timeout_ms(opts) + r = subagent_await_llm(token, ref, timestamp() + timeout_ms) + tag = elem(r, 0) + if (tag == 'ok') { + result = elem(r, 1) + if (result == nil) { {'error', subagent_fail_reason(LLM.last_fail(opts))} } + else { {'ok', result} } + } else { if (tag == 'down') { + {'error', "the LLM worker crashed: " ++ to_string(elem(r, 1))} + } else { + # Give up on the hung call and kill the worker: it is parked in a + # blocking HTTP call, so the kill lands when that returns — which + # still stops LLM.chat's retry loop from re-sending the request + # up to 3 more times. A late reply carries a stale token and is + # dropped by any later await. + demonitor(ref) + exit_proc(w, 'killed') + {'error', "no response from the LLM within " ++ to_string(timeout_ms / 1000) ++ + "s (llm_timeout_ms / SWARM_CODE_LLM_TIMEOUT_MS)"} + }} +} + +fun subagent_fail_reason(why) { + if (why == 'fatal') { "the endpoint rejected the request (4xx / malformed request)" } + else { if (why == 'transient') { "the endpoint kept failing after retries (5xx / dropped connection)" } + else { "the LLM call failed after retries (no usable response)" } } +} + +# Wait for OUR worker's reply ({'ok', r}), its death ({'down', why}), +# or the deadline ({'timeout'}). On success, also consume the worker's +# normal-exit DOWN so it can't linger in the caller's (main's) mailbox. +fun subagent_await_llm(token, ref, deadline) { wait = deadline - timestamp() - if (wait <= 0) { nil } + if (wait <= 0) { {'timeout'} } else { receive { {'llm_result', t, r} -> - if (t == token) { r } - else { subagent_await_llm(token, deadline) } - after wait { nil } + if (t == token) { + receive { + {'DOWN', dref, _kind, _pid, _why} when dref == ref -> 'ok' + after 1000 { demonitor(ref) } + } + {'ok', r} + } + else { subagent_await_llm(token, ref, deadline) } + {'DOWN', dref, _kind, _pid, why} when dref == ref -> + {'down', why} + after wait { {'timeout'} } } } } fun run_subagent_loop(history, opts, step) { if (step >= subagent_max_steps()) { - "[subagent hit max steps without final answer]" + subagent_partial(history, "hit the " ++ to_string(subagent_max_steps()) ++ + "-step limit before a final answer") } else { - token = to_string(timestamp()) - spawn(subagent_llm_worker(token, self(), history, opts)) - result = subagent_await_llm(token, timestamp() + 300000) - if (result == nil) { - "[subagent llm call failed]" + call = subagent_llm_call(history, opts) + if (elem(call, 0) != 'ok') { + subagent_partial(history, "LLM call failed — " ++ to_string(elem(call, 1))) } else { - content = to_string(map_get(result, 'content')) + result = elem(call, 1) + content_v = map_get(result, 'content') + content = if (content_v == nil) { "" } else { to_string(content_v) } tcs_v = map_get(result, 'tool_calls') reasoning = map_get(result, 'reasoning') tool_calls = if (tcs_v == nil) { [] } else { tcs_v } if (length(tool_calls) == 0) { content } else { - asst = LLM.new_message_assistant(content, tool_calls, reasoning) + asst = LLM.new_message_assistant(content, sanitize_tool_calls(tool_calls, []), reasoning) with_assistant = list_append(history, asst) - post_tools = subagent_exec_all(tool_calls, with_assistant, opts) - run_subagent_loop(post_tools, opts, step + 1) + # A cut-off turn's calls are answered, not run (see run_turn). + cut = turn_cut_reason(result) + post_tools = if (cut != 'complete') { refuse_tool_calls(tool_calls, with_assistant, cut, opts) } + else { subagent_exec_all(tool_calls, with_assistant, opts) } + # The subagent's OWN guardrail (fresh table, see + # handle_task_tool): a runaway failure streak stops the + # subagent, never the parent's turn. + guard = map_get(opts, 'guardrails_table') + halt = if (guard == nil) { nil } else { ets_get(guard, 'halt_reason') } + if (halt != nil) { subagent_partial(post_tools, "guardrail halt — " ++ to_string(halt)) } + else { run_subagent_loop(post_tools, opts, step + 1) } } } } } +# An abnormal stop still returns what the subagent had: its latest +# notes (last non-empty assistant text) and a digest of its most recent +# tool calls + results, headed by why it stopped — the old +# "[subagent hit max steps…]" discarded all of it. +fun subagent_partial(history, why) { + notes = subagent_last_text(history, "") + digest = subagent_tool_digest(history, map_new(), []) + recent = subagent_last_n(digest, 4) + head = "[subagent stopped: " ++ why ++ " — partial results below]" + notes_part = if (string_length(string_trim(notes)) == 0) { "" } + else { "\n\nLatest notes from the subagent:\n" ++ notes } + tools_part = if (length(recent) == 0) { "" } + else { "\n\nMost recent tool results (oldest first, " ++ + to_string(length(recent)) ++ " of " ++ to_string(length(digest)) ++ "):" ++ + subagent_render_digest(recent, "") } + if (string_length(notes_part) == 0 && string_length(tools_part) == 0) { + head ++ "\n(no findings yet)" + } else { head ++ notes_part ++ tools_part } +} + +fun subagent_last_text(msgs, acc) { + if (length(msgs) == 0) { acc } + else { + m = hd(msgs) + c = map_get(m, 'content') + next = if (map_get(m, 'role') == 'assistant' && c != nil && + string_length(string_trim(to_string(c))) > 0) { to_string(c) } else { acc } + subagent_last_text(tl(msgs), next) + } +} + +# [{label, result}] in history order. `calls` maps tool_call id → +# "name args" from the assistant turns so each result is labelled. +fun subagent_tool_digest(msgs, calls, acc) { + if (length(msgs) == 0) { acc } + else { + m = hd(msgs) + role = map_get(m, 'role') + if (role == 'assistant') { + tcs = map_get(m, 'tool_calls') + subagent_tool_digest(tl(msgs), subagent_index_calls(if (tcs == nil) { [] } else { tcs }, calls), acc) + } else { if (role == 'tool') { + id = to_string(map_get(m, 'tool_call_id')) + label = map_get(calls, id) + entry = {(if (label == nil) { "tool" } else { label }), to_string(map_get(m, 'content'))} + subagent_tool_digest(tl(msgs), calls, list_append(acc, entry)) + } else { subagent_tool_digest(tl(msgs), calls, acc) } } + } +} + +fun subagent_index_calls(tcs, calls) { + if (length(tcs) == 0) { calls } + else { + tc = hd(tcs) + label = to_string(map_get(tc, 'name')) ++ " " ++ + preview_string(to_string(map_get(tc, 'arguments')), 120) + subagent_index_calls(tl(tcs), map_put(calls, to_string(map_get(tc, 'id')), label)) + } +} + +fun subagent_last_n(lst, n) { + k = length(lst) - n + if (k <= 0) { lst } else { subagent_drop(lst, k) } +} + +fun subagent_drop(lst, k) { + if (k <= 0 || length(lst) == 0) { lst } else { subagent_drop(tl(lst), k - 1) } +} + +fun subagent_render_digest(entries, acc) { + if (length(entries) == 0) { acc } + else { + e = hd(entries) + body = string_trim(elem(e, 1)) + shown = if (string_length(body) > 2000) { + string_sub(body, 0, subagent_utf8_floor(body, 2000)) ++ "\n…[result truncated]" + } else { body } + subagent_render_digest(tl(entries), acc ++ "\n- " ++ elem(e, 0) ++ "\n" ++ shown) + } +} + +# Head+tail cap with an explicit marker; the head (where findings are +# usually summarised) gets two thirds. Cut points are moved back onto +# UTF-8 character boundaries so the parent never re-sends a split +# multi-byte sequence to the LLM. +fun subagent_cap(s) { + n = string_length(s) + cap = subagent_result_cap() + if (n <= cap) { s } + else { + head_len = subagent_utf8_floor(s, cap * 2 / 3) + tail_start = subagent_utf8_floor(s, n - (cap - cap * 2 / 3)) + string_sub(s, 0, head_len) ++ + "\n\n…[" ++ to_string(tail_start - head_len) ++ " bytes of the subagent's answer elided " ++ + "(cap " ++ to_string(cap) ++ ") — give it a narrower task for the details]…\n\n" ++ + string_sub(s, tail_start, n - tail_start) + } +} + +# Largest i' <= i that does not point into the middle of a UTF-8 +# sequence (continuation bytes are 0x80-0xBF). +fun subagent_utf8_floor(s, i) { + if (i <= 0) { 0 } + else { if (i >= string_length(s)) { string_length(s) } + else { + c = codepoint_at(s, i) + if (c >= 128 && c < 192) { subagent_utf8_floor(s, i - 1) } else { i } + }} +} + # Tools the main agent can use but subagents cannot. Anything that # affects long-lived host state (memory, skills, background servers, # git commits, nested task spawns) is locked to the main agent so a @@ -2548,20 +3023,38 @@ fun SUBAGENT_BLOCKED_TOOLS() { ToolRegistry.subagent_blocked_tools() } -fun subagent_blocked(name) { - if (ToolRegistry.allowed_in("subagent", name) == 'true') { 'false' } +# Blocked for a general subagent (the host-state blocklist above). +fun subagent_blocked(name) { subagent_blocked_for("general", name) } + +# Per-type policy: explore = read-only inspection tools, bash = bash, +# general = everything but SUBAGENT_BLOCKED_TOOLS. Backed by the same +# ToolRegistry contexts ToolExecutor enforces on dispatch. +fun subagent_blocked_for(stype, name) { + if (ToolRegistry.allowed_in(subagent_context(stype), to_string(name)) == 'true') { 'false' } else { 'true' } } +fun subagent_block_msg(stype, name_str) { + if (stype == "explore") { + "error: tool '" ++ name_str ++ "' is not available to an explore (read-only) subagent — allowed: " ++ + subagent_join(ToolRegistry.names_for("subagent_explore"), "") + } else { if (stype == "bash") { + "error: tool '" ++ name_str ++ "' is not available to a bash subagent — only bash is allowed" + } else { + "error: tool '" ++ name_str ++ "' is not available to subagents — only the main agent can use it" + }} +} + fun subagent_exec_all(tool_calls, history, opts) { if (length(tool_calls) == 0) { history } else { tc = hd(tool_calls) id = to_string(map_get(tc, 'id')) name_str = to_string(map_get(tc, 'name')) - if (subagent_blocked(name_str) == 'true') { - blocked = "error: tool '" ++ name_str ++ "' is not available to subagents — only the main agent can use it" - tool_msg = LLM.new_message_tool(id, blocked) + stype_v = map_get(opts, 'subagent_type') + stype = if (stype_v == nil) { "general" } else { to_string(stype_v) } + if (subagent_blocked_for(stype, name_str) == 'true') { + tool_msg = LLM.new_message_tool(id, subagent_block_msg(stype, name_str)) new_hist = list_append(history, tool_msg) subagent_exec_all(tl(tool_calls), new_hist, opts) } else { @@ -2569,7 +3062,10 @@ fun subagent_exec_all(tool_calls, history, opts) { args_raw = to_string(map_get(tc, 'arguments')) args_map = json_decode(args_raw) args_map_safe = if (args_map == nil) { map_new() } else { args_map } - sub_result = dispatch_tool(name_atom, args_map_safe, opts) + # Same F5 guard as execute_all: never run a cut-off call. + sub_result = if (args_malformed(args_raw) == 'true') { + "error: the arguments for '" ++ name_str ++ "' were not valid JSON (likely truncated mid-string). Reissue this single tool call with complete, valid JSON arguments." + } else { dispatch_tool(name_atom, args_map_safe, opts) } tool_msg = LLM.new_message_tool(id, sub_result) new_hist = list_append(history, tool_msg) subagent_exec_all(tl(tool_calls), new_hist, opts) @@ -2613,11 +3109,12 @@ fun execute_all(tool_calls, history, opts) { name_atom = string_to_atom(name_str) args_raw = to_string(map_get(tc, 'arguments')) args_map = json_decode(args_raw) - # F5: a NON-EMPTY args blob that fails to parse is a truncated/malformed - # tool call. Don't silently dispatch with empty args (which yields a - # confusing "missing X" for an arg the model DID supply) — tell it to reissue. - args_trim = string_trim(args_raw) - malformed = if (args_map == nil && args_trim != "" && args_trim != "{}" && args_trim != "null") { 'true' } else { 'false' } + # F5: a NON-EMPTY args blob that isn't one complete JSON value is a + # truncated/malformed tool call. Don't dispatch it — with empty args + # (a confusing "missing X" for an arg the model DID supply) or, worse, + # with the lenient decoder's partial value (a command / file content + # cut mid-string) — tell the model to reissue. + malformed = args_malformed(args_raw) result = if (malformed == 'true') { turn_print(opts, "") turn_print(opts, UI.tool_header_str(name_atom, "(malformed / truncated arguments)")) @@ -2635,7 +3132,7 @@ fun execute_all(tool_calls, history, opts) { effective_args = map_get(prepared, 'args') decision = resolve_permission(name_atom, effective_args, opts) if (decision == 'deny') { - denial = "error: permission denied for tool '" ++ name_str ++ "'" + denial = permission_denial(name_atom, name_str, effective_args, opts) turn_print(opts, UI.err_text(denial)) denial } else { @@ -2656,7 +3153,107 @@ fun execute_all(tool_calls, history, opts) { Log.tool_result(name_atom, string_length(result), had_err) hist_done = list_append(history, LLM.new_message_tool(id, result)) journal_sync(opts, hist_done) - execute_all(tl(tool_calls), hist_done, opts) + # ESC/Ctrl-C on a tool stops the whole turn: the rest of this batch + # is recorded as skipped (every tool_call still needs its result + # message) and run_turn hands control back instead of re-calling + # the model. + if (is_user_interrupt(result) == 'true' && map_get(opts, 'headless') != 'true') { + skip_remaining_tools(tl(tool_calls), hist_done, opts) + } else { + execute_all(tl(tool_calls), hist_done, opts) + } + } +} + +fun INTERRUPT_SKIPPED() { "[skipped] the user interrupted this turn before this tool ran." } + +# ------------------------------------------------------------ +# Cut-off turns — a tool call the model didn't finish is never run. +# ------------------------------------------------------------ +# The runtime appends this to the content when ESC/Ctrl-C lands mid-stream +# (and routed mode's interrupted_stream_result mirrors it). +fun INTERRUPT_MARKER() { "[Request interrupted by user]" } + +# 'interrupted' (the user stopped the stream), 'truncated' (output-token +# limit) or 'complete'. llm.sw flags both on the RAW content — inband +# parsing strips everything after the first call marker, marker included — +# the content test is the fallback for a result built elsewhere. +fun turn_cut_reason(result) { + if (map_get(result, 'interrupted') == 'true') { 'interrupted' } + else { if (string_ends_with(string_trim(to_string(map_get(result, 'content'))), + INTERRUPT_MARKER()) == 'true') { 'interrupted' } + else { if (map_get(result, 'truncated') == 'true') { 'truncated' } + else { 'complete' } } } +} + +fun refused_tool_result(reason, raw_args) { + if (reason == 'interrupted') { + "[interrupted] the user pressed ESC while this tool call was still streaming — it was NOT run." + } else { + cut = if (args_malformed(raw_args) == 'true') { + " Its arguments were cut off after " ++ to_string(string_length(to_string(raw_args))) ++ " bytes." + } else { "" } + "error: not executed — your response hit the output token limit before this tool " ++ + "call was complete, so it was NOT run." ++ cut ++ " Reissue it with complete " ++ + "arguments; for large content write a short stub first, then append the rest " ++ + "with `edit` in smaller pieces." + } +} + +# Answer every call of a cut-off turn with a not-run result instead of +# dispatching it (history stays valid: each tool_call id gets its tool +# message). No journal_sync here — a subagent's opts carry the PARENT's +# journal_path; run_turn journals the result itself. +fun refuse_tool_calls(tool_calls, history, reason, opts) { + if (length(tool_calls) == 0) { history } + else { + tc = hd(tool_calls) + id = to_string(map_get(tc, 'id')) + name_atom = string_to_atom(to_string(map_get(tc, 'name'))) + raw = to_string(map_get(tc, 'arguments')) + result = refused_tool_result(reason, raw) + if (map_get(opts, 'is_subagent') != 'true') { + what = if (reason == 'interrupted') { "interrupted while streaming — not run" } + else { "cut off at the output limit — not run" } + turn_print(opts, "") + turn_print(opts, UI.tool_header_str(name_atom, what)) + } + Log.tool_call(name_atom, raw) + Log.tool_result(name_atom, string_length(result), 'true') + refuse_tool_calls(tl(tool_calls), list_append(history, LLM.new_message_tool(id, result)), reason, opts) + } +} + +# Result markers for a tool the user stopped: shell_managed's ESC/Ctrl-C +# kill, and collect_tool_result's interrupt of a non-shell tool. +fun is_user_interrupt(result) { + s = to_string(result) + if (string_starts_with(s, "[stopped by user") == 'true') { 'true' } + else { if (string_starts_with(s, "[interrupted]") == 'true') { 'true' } + else { if (s == INTERRUPT_SKIPPED()) { 'true' } else { 'false' } } } +} + +fun skip_remaining_tools(tool_calls, history, opts) { + if (length(tool_calls) == 0) { history } + else { + id = to_string(map_get(hd(tool_calls), 'id')) + h2 = list_append(history, LLM.new_message_tool(id, INTERRUPT_SKIPPED())) + journal_sync(opts, h2) + skip_remaining_tools(tl(tool_calls), h2, opts) + } +} + +# Did the user interrupt any of the last `n` tool results? +fun turn_interrupted(history, n) { + any_interrupted(take_last(history, n)) +} + +fun any_interrupted(msgs) { + if (length(msgs) == 0) { 'false' } + else { + m = hd(msgs) + if (map_get(m, 'role') == 'tool' && is_user_interrupt(map_get(m, 'content')) == 'true') { 'true' } + else { any_interrupted(tl(msgs)) } } } @@ -2717,7 +3314,13 @@ fun resolve_permission(name, args, opts) { else { if (raw == 'deny') { 'deny' } else { headless = map_get(opts, 'headless') - if (headless == 'true') { 'allow' } + # Headless has nobody to answer an 'ask' — which only arises for a + # dangerous bash command, a user-configured "ask", or an MCP tool + # (default-allowed tools never get here). Deny unless the user + # opted in; permission_denial tells the model how. + if (headless == 'true') { + if (getenv("SWARM_CODE_HEADLESS_APPROVE") == "1") { 'allow' } else { 'deny' } + } else { # Session answer cache FIRST: a user who explicitly said "No, # always deny X this session" must stay denied even after @@ -2757,6 +3360,27 @@ fun resolve_permission(name, args, opts) { }} } +# Model-facing denial text. A headless 'ask' is refused only because no +# one can approve it, so name the opt-in; a hardline / configured 'deny' +# keeps the plain message (no opt-in exists for those). +fun permission_denial(name_atom, name_str, args, opts) { + if (map_get(opts, 'headless') == 'true' && + Config.check_permission(name_atom, args, opts) == 'ask') { + risk = Config.command_risk(name_atom, args) + dangerous = if (elem(risk, 0) == 'dangerous') { 'true' } else { 'false' } + what = if (dangerous == 'true') { " — flagged dangerous: " ++ to_string(elem(risk, 1)) } else { "" } + how = if (dangerous == 'true') { + "run headless with SWARM_CODE_HEADLESS_APPROVE=1" + } else { + "run headless with SWARM_CODE_HEADLESS_APPROVE=1, or set \"" ++ name_str ++ + "\": \"allow\" under permissions in ~/.swarm-code/settings.json" + } + "error: permission denied for tool '" ++ name_str ++ "'" ++ what ++ + " — it needs approval and headless mode has no one to ask. " ++ + "To allow such calls unattended, the user can " ++ how + } else { Config.denial_message(name_str, args, opts) } +} + fun ask_via_reader(name, opts, table, cache_key) { reader_pid = map_get(opts, 'reader_pid') if (reader_pid == nil) { 'deny' } @@ -3228,8 +3852,10 @@ fun is_known_slash_command(cmd) { else { if (cmd == "/flows") { 'true' } else { if (cmd == "/bg") { 'true' } else { if (cmd == "/mode") { 'true' } + else { if (cmd == "/trust") { 'true' } + else { if (cmd == "/untrust") { 'true' } else { if (cmd == "/expand") { 'true' } - else { 'false' }}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}} + else { 'false' }}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}}} } # (preview_string and string_to_atom moved earlier — see below) diff --git a/src/background.sw b/src/background.sw index 9233261..61f92c6 100644 --- a/src/background.sw +++ b/src/background.sw @@ -35,10 +35,11 @@ import Util # '{id}/log_grew_at' → ms timestamp the log size last changed # '{id}/stalled_sent'→ 'true' once a bg_stalled msg has fired (at most once) # 'next_id' → counter +# 'dir' → this session's private task directory (see session_dir) export [ - init, launch, launch_server, status, result, list_all, - log_path_for, tail_log, kill_task, + init, launch, launch_cmd, launch_server, status, result, list_all, + log_path_for, pid_file_for, exit_file_for, session_dir, tail_log, kill_task, poll_and_notify, all_pending_ids, finalize_if_done, wait_for_task, fg_claim, fg_release @@ -56,31 +57,80 @@ fun init() { table } +# Private per-session directory for task files, created on first launch +# with `mktemp -d` (mode 0700) and remembered in the table. Task ids restart +# at bg-0 in every session, so the old fixed /tmp/swarm-code-bg-N.{log,pid, +# exit} paths collided across concurrent sessions: B's launch deleted A's +# files, B's pid overwrote A's (so `bg_kill bg-0` in A killed B's task), and +# a stale exit file could mark a fresh task done. nil if it can't be made. +fun session_dir(table) { + d = ets_get(table, 'dir') + if (d != nil) { d } + else { + tmp = getenv("TMPDIR") + base = if (tmp == nil || string_length(to_string(tmp)) == 0) { "/tmp" } else { to_string(tmp) } + r = shell("mktemp -d " ++ Util.shell_q(base ++ "/swarm-code-bg.XXXXXX") ++ " 2>/dev/null") + made = string_trim(to_string(elem(r, 1))) + if (elem(r, 0) == 0 && string_length(made) > 0 && file_exists(made) == 'true') { + ets_put(table, 'dir', made) + made + } else { nil } + } +} + +fun task_file(table, task_id, ext) { + d = session_dir(table) + if (d == nil) { nil } else { d ++ "/" ++ task_id ++ "." ++ ext } +} + # Path of the log file capturing stdout+stderr. -fun log_path_for(task_id) { - "/tmp/swarm-code-" ++ task_id ++ ".log" +fun log_path_for(table, task_id) { + known = ets_get(table, task_id ++ "/log_file") + if (known != nil) { known } else { task_file(table, task_id, "log") } } -fun pid_file_for(task_id) { - "/tmp/swarm-code-" ++ task_id ++ ".pid" +fun pid_file_for(table, task_id) { + known = ets_get(table, task_id ++ "/pid_file") + if (known != nil) { known } else { task_file(table, task_id, "pid") } +} + +fun exit_file_for(table, task_id) { + known = ets_get(table, task_id ++ "/exit_file") + if (known != nil) { known } else { task_file(table, task_id, "exit") } } # Launch a command detached. Returns the task id string. fun launch(table, command, label) { + launch_cmd(table, command, command, label) +} + +# Launch `run_cmd` detached but record `display_cmd` as the task's command. +# The tool layer passes a wrapped script (Util.noninteractive_wrap) as +# run_cmd and the model's raw command as display_cmd, so /bg listings show +# what the model asked for, not the wrapper. +fun launch_cmd(table, run_cmd, display_cmd, label) { + if (session_dir(table) == nil) { + "error: failed to start background task (could not create a private task directory under $TMPDIR or /tmp)" + } else { launch_in_dir(table, run_cmd, display_cmd, label) } +} + +fun launch_in_dir(table, run_cmd, display_cmd, label) { + command = display_cmd raw_id = ets_get(table, 'next_id') next_id = if (raw_id == nil) { 0 } else { raw_id } task_id = "bg-" ++ to_string(next_id) ets_put(table, 'next_id', next_id + 1) - log_file = log_path_for(task_id) - pid_file = pid_file_for(task_id) - exit_file = "/tmp/swarm-code-" ++ task_id ++ ".exit" + log_file = task_file(table, task_id, "log") + pid_file = task_file(table, task_id, "pid") + exit_file = task_file(table, task_id, "exit") - # The rm -f first is load-bearing: bg-N ids restart at 0 every - # session and /tmp/swarm-code-bg-N.* files are never cleaned, so a - # stale exit file from a previous session would instantly mark this - # fresh task done (poll_and_notify and Flows both file_exists it). - shell("rm -f " ++ exit_file ++ " " ++ pid_file ++ " " ++ log_file) + # Ids are unique within this table and the directory is private to + # it, so nothing stale can be here — the deletes are belt-and-braces + # (a stale exit file would instantly mark the task done). + file_delete(exit_file) + file_delete(pid_file) + file_delete(log_file) # shell_detached double-forks + setsid()s a worker that runs # `( command ); echo $? > exit_file` under /bin/sh with stdin=/dev/null @@ -88,7 +138,7 @@ fun launch(table, command, label) { # since the worker leads its own session/process group). No sleep + # pid-file readback race any more — the runtime hands us the pid # synchronously off its internal pipe. - pid = shell_detached(command, log_file, exit_file) + pid = shell_detached(run_cmd, log_file, exit_file) if (pid == nil) { "error: failed to start background task" } else { @@ -146,7 +196,7 @@ fun result(table, task_id) { # Tail a task's log (N lines). # n_lines is validated upstream (Tools.do_bg_tail) to be a clean integer. -# Log path is internal (/tmp/swarm-code-bg-N.log) but quote anyway as +# Log path is internal (/bg-N.log) but quote anyway as # belt-and-braces. fun tail_log(table, task_id, n_lines) { log_file = ets_get(table, task_id ++ "/log_file") @@ -162,8 +212,8 @@ fun tail_log(table, task_id, n_lines) { } # Kill a task by its OS pid. Validates pid as digits-only before -# using it — pid_file lives in shared /tmp so a stray writer could -# otherwise feed junk here. pid_kill_group SIGTERMs the whole process +# using it — belt-and-braces now that the pid file lives in this +# session's private directory. pid_kill_group SIGTERMs the whole process # group (the worker leads its own pgroup), so child processes the task # forked die too, not just the /bin/sh wrapper. fun kill_task(table, task_id) { diff --git a/src/browser.sw b/src/browser.sw index 8fc087b..f0d943e 100644 --- a/src/browser.sw +++ b/src/browser.sw @@ -1,6 +1,7 @@ module Browser import UI +import PathGuard # ============================================================ # Browser — CDP-over-WebSocket browser control, no Node, no Python @@ -394,6 +395,13 @@ fun interpret_eval_result(result, label) { # ------------------------------------------------------------ fun screenshot(session, path, opts) { p = to_string(path) + # Same write policy as write/edit: a screenshot must not land on a + # credential path or swarm-code's own control files. + guard = PathGuard.validate_write(p) + if (guard != "ok") { guard } else { screenshot_to(session, p, opts) } +} + +fun screenshot_to(session, p, opts) { result = cdp_call_pg(session, "Page.captureScreenshot", %{format: "png"}, 30000, "browser_screenshot", opts) if (result == nil) { "error: screenshot failed (capture timed out or browser unresponsive)" } else { if (string_starts_with(to_string(result), "error:") == 'true') { result } diff --git a/src/config.sw b/src/config.sw index 48da4fc..8535bc8 100644 --- a/src/config.sw +++ b/src/config.sw @@ -1,6 +1,8 @@ module Config import Util +import CommandGuard +import Hooks # ============================================================ # Config — settings.json, SWARM.md, permissions, hooks @@ -10,6 +12,11 @@ import Util # 1. $HOME/.swarm-code/settings.json (user-global) # 2. ./.swarm-code.json (project-local, overrides) # +# The project file arrives with whatever repo you cloned, so it is NOT +# trusted by default: only harmless keys apply and permissions may only +# tighten (see project_scope). Listing the directory under +# "trusted_projects" in the user settings lets it apply in full. +# # Project context is loaded from the first that exists: # 1. ./SWARM.md (swarm-code's preferred name) # 2. ./CLAUDE.md (fallback, most repos have one) @@ -33,19 +40,25 @@ import Util # } export [load, load_project_context, check_permission, run_hooks, is_dangerous_bash, is_hardline_bash, - llm_timeout_ms] + llm_timeout_ms, project_scope, project_ignored_keys, is_trusted_dir, project_notice, + set_trusted, set_trusted_at, trusted_list, project_gated_keys, join_names, + endpoint_host, is_local_endpoint, endpoint_refusal, + command_of, command_risk, denial_message, denial_reason, uses_sudo, + run_hooks_verdict] # ------------------------------------------------------------ # load settings — merged map from user + project config files # ------------------------------------------------------------ fun load() { - user_path = getenv("HOME") ++ "/.swarm-code/settings.json" - project_path = "./.swarm-code.json" - user_settings = load_one(user_path) - project_settings = load_one(project_path) - map_merge(user_settings, project_settings) + user_settings = load_one(user_settings_path()) + project_settings = load_one(project_settings_path()) + trusted = project_trusted(user_settings, project_settings) + map_merge(user_settings, project_scope(user_settings, project_settings, trusted)) } +fun user_settings_path() { getenv("HOME") ++ "/.swarm-code/settings.json" } +fun project_settings_path() { "./.swarm-code.json" } + fun load_one(path) { if (file_exists(path) == 'false') { map_new() @@ -54,12 +67,464 @@ fun load_one(path) { if (file_content == nil) { map_new() } else { + # A non-object top level ("[]", "1") would crash map_merge. decoded = json_decode(file_content) - if (decoded == nil) { map_new() } else { decoded } + if (decoded == nil || is_map(decoded) != 'true') { map_new() } else { decoded } } } } +# ------------------------------------------------------------ +# Project scope — what a repo's ./.swarm-code.json may contribute. +# ------------------------------------------------------------ +# Opening swarm-code inside a cloned repo must not hand that repo's +# author control of the session. Unless the directory is trusted, the +# project file may NOT: +# * run commands — hooks (SessionStart fires at launch) +# * start processes — mcpServers +# * redirect the prompt — endpoint, api_key, providers, profiles, +# fallback_profile +# * loosen permissions — a "permissions" entry only applies when +# it is STRICTER (allow < ask < deny) than +# what the user would otherwise get +# It is an ALLOW-list (project_safe_keys), so a key added later is +# ignored from project scope until someone decides it is harmless. +# +# Opt-in — user settings only, a project can never trust itself: +# "trusted_projects": ["/abs/path/to/repo", ...] +# in ~/.swarm-code/settings.json; that directory's file applies in full. +# ------------------------------------------------------------ +fun project_safe_keys() { + ["model", "max_tokens", "llm_timeout_ms", "chat_template_kwargs", "vision"] +} + +# The part of `project` that may be merged over `user`. trusted='true' +# returns the project unchanged (the old full-override behaviour). +fun project_scope(user, project, trusted) { + if (trusted == 'true') { project } + else { + safe = copy_keys(project, project_safe_keys(), map_new()) + pp = map_get(project, 'permissions') + if (pp == nil || is_map(pp) != 'true') { safe } + else { map_put(safe, 'permissions', tighten_permissions(user_perms(user), pp)) } + } +} + +fun copy_keys(src, keys, acc) { + if (length(keys) == 0) { acc } + else { + k = hd(keys) + v = map_get(src, k) + next = if (v == nil) { acc } else { map_put(acc, k, v) } + copy_keys(src, tl(keys), next) + } +} + +fun user_perms(user) { + up = map_get(user, 'permissions') + if (up == nil || is_map(up) != 'true') { map_new() } else { up } +} + +# User permissions plus every project entry that is stricter than the +# user's effective decision for that tool (configured, else the default). +fun tighten_permissions(uperms, pperms) { + tighten_loop(map_keys(pperms), map_values(pperms), uperms, uperms) +} + +fun tighten_loop(keys, vals, uperms, acc) { + if (length(keys) == 0) { acc } + else { + k = hd(keys) + next = if (perm_stricter(hd(vals), effective_user_perm(uperms, k)) == 'true') { + map_put(acc, k, hd(vals)) + } else { acc } + tighten_loop(tl(keys), tl(vals), uperms, next) + } +} + +fun effective_user_perm(uperms, tool) { + cur = map_get(uperms, tool) + if (cur == nil) { default_permission(tool) } else { string_to_perm(cur) } +} + +# 'true' when the configured value v is stricter than the decision `cur`. +fun perm_stricter(v, cur) { + if (perm_rank(string_to_perm(v)) > perm_rank(cur)) { 'true' } else { 'false' } +} + +fun perm_rank(p) { + if (p == 'deny') { 2 } else { if (p == 'ask') { 1 } else { 0 }} +} + +# Names of the project keys project_scope drops for an untrusted dir — +# "permissions." for a loosening entry. [] when nothing is lost. +fun project_ignored_keys(user, project) { + ignored_loop(map_keys(project), project, user, []) +} + +fun ignored_loop(keys, project, user, acc) { + if (length(keys) == 0) { acc } + else { + k = hd(keys) + ks = to_string(k) + next = if (list_has(project_safe_keys(), ks) == 'true') { acc } + else { if (ks == "permissions") { + pp = map_get(project, k) + if (pp == nil || is_map(pp) != 'true') { list_append(acc, ks) } + else { acc ++ loosening_keys(map_keys(pp), map_values(pp), user_perms(user), []) } + } + else { list_append(acc, ks) }} + ignored_loop(tl(keys), project, user, next) + } +} + +fun loosening_keys(keys, vals, uperms, acc) { + if (length(keys) == 0) { acc } + else { + k = hd(keys) + cur = effective_user_perm(uperms, k) + next = if (perm_rank(string_to_perm(hd(vals))) < perm_rank(cur)) { + list_append(acc, "permissions." ++ to_string(k)) + } else { acc } + loosening_keys(tl(keys), tl(vals), uperms, next) + } +} + +fun join_names(names, acc) { + if (length(names) == 0) { acc } + else { + sep = if (string_length(acc) == 0) { "" } else { ", " } + join_names(tl(names), acc ++ sep ++ to_string(hd(names))) + } +} + +fun list_has(lst, x) { + if (length(lst) == 0) { 'false' } + else { if (hd(lst) == x) { 'true' } else { list_has(tl(lst), x) }} +} + +# Is the current directory listed in the user's "trusted_projects"? +# Costs one shell() (for the real cwd), paid only when there IS a +# project file and the user HAS a trust list. +fun project_trusted(user, project) { + tp = map_get(user, 'trusted_projects') + if (map_size(project) == 0 || tp == nil || is_list(tp) != 'true') { 'false' } + else { + # Both spellings of the cwd: logical (symlinks kept, as the user + # typed it) and physical (pwd -P), so either form of entry matches. + cwds = string_split(string_trim(to_string(elem(shell("pwd; pwd -P"), 1))), "\n") + is_trusted_dir(tp, cwds) + } +} + +# Pure core of project_trusted: does any entry of `trusted` equal any +# of `cwds` (trailing slashes ignored)? Entries are absolute paths. +fun is_trusted_dir(trusted, cwds) { + if (length(trusted) == 0) { 'false' } + else { + t = strip_trailing_slash(string_trim(to_string(hd(trusted)))) + if (string_length(t) > 0 && dir_in(cwds, t) == 'true') { 'true' } + else { is_trusted_dir(tl(trusted), cwds) } + } +} + +fun dir_in(cwds, t) { + if (length(cwds) == 0) { 'false' } + else { if (strip_trailing_slash(string_trim(to_string(hd(cwds)))) == t) { 'true' } + else { dir_in(tl(cwds), t) }} +} + +fun strip_trailing_slash(s) { + if (string_length(s) > 1 && string_ends_with(s, "/") == 'true') { + strip_trailing_slash(string_sub(s, 0, string_length(s) - 1)) + } else { s } +} + +# ------------------------------------------------------------ +# swarm-code trust / untrust — edit "trusted_projects" for the user. +# ------------------------------------------------------------ +# set_trusted_at(settings_path, dir, add) → {'ok', abs_dir, changed} | +# {'error', message}. `dir` is resolved to an absolute path (cd + pwd, so +# symlinks and relative paths work); add='true' appends it once, +# add='false' removes every matching entry. Every other key in the file +# is kept, and a settings.json that doesn't parse as an object is never +# overwritten (the user would lose it) — the error says to fix it first. +fun set_trusted_at(settings_path, dir, add) { + r = shell("cd " ++ Util.shell_q(to_string(dir)) ++ " 2>/dev/null && pwd") + abs = string_trim(to_string(elem(r, 1))) + if (elem(r, 0) != 0 || string_length(abs) == 0) { + {'error', "no such directory: " ++ to_string(dir)} + } else { + raw = if (file_exists(settings_path) == 'true') { file_read(settings_path) } else { nil } + decoded = if (raw == nil || string_length(string_trim(raw)) == 0) { map_new() } else { json_decode(raw) } + if (decoded == nil || is_map(decoded) != 'true') { + {'error', settings_path ++ " is not a valid JSON object — fix it by hand first " ++ + "(swarm-code won't overwrite it)"} + } else { + cur = map_get(decoded, 'trusted_projects') + tp = if (cur != nil && is_list(cur) == 'true') { cur } else { [] } + had = is_trusted_dir(tp, [abs]) + next = if (add == 'true') { + if (had == 'true') { tp } else { list_append(tp, abs) } + } else { drop_dir(tp, abs, []) } + changed = if (add == 'true') { bool_not_str(had) } else { had } + if (changed == 'false') { {'ok', abs, 'false'} } + else { + file_mkdir(dirname_of(settings_path)) + rc = file_atomic_write(settings_path, Util.json_pretty(map_put(decoded, 'trusted_projects', next))) + if (rc == 'ok') { {'ok', abs, 'true'} } + else { {'error', "could not write " ++ settings_path} } + } + } + } +} + +fun set_trusted(dir, add) { set_trusted_at(user_settings_path(), dir, add) } + +fun trusted_list() { + tp = map_get(load_one(user_settings_path()), 'trusted_projects') + if (tp != nil && is_list(tp) == 'true') { tp } else { [] } +} + +fun drop_dir(tp, abs, acc) { + if (length(tp) == 0) { acc } + else { + t = strip_trailing_slash(string_trim(to_string(hd(tp)))) + next = if (t == abs) { acc } else { list_append(acc, hd(tp)) } + drop_dir(tl(tp), abs, next) + } +} + +fun bool_not_str(b) { if (b == 'true') { 'false' } else { 'true' } } + +fun dirname_of(p) { + i = last_slash(p, string_length(p) - 1) + if (i <= 0) { "/" } else { string_sub(p, 0, i) } +} + +fun last_slash(p, i) { + if (i < 0) { 0 - 1 } + else { if (string_sub(p, i, 1) == "/") { i } else { last_slash(p, i - 1) } } +} + +# The keys a directory's .swarm-code.json sets that only apply once it is +# trusted — shown by `swarm-code trust` so the user sees what they grant. +fun project_gated_keys(dir) { + project = load_one(to_string(dir) ++ "/.swarm-code.json") + if (map_size(project) == 0) { nil } + else { project_ignored_keys(map_new(), project) } +} + +# One-line notice naming what an untrusted ./.swarm-code.json tried to +# set, or nil when nothing was dropped. main prints it once at startup. +fun project_notice() { + project = load_one(project_settings_path()) + if (map_size(project) == 0) { nil } + else { + user = load_one(user_settings_path()) + if (project_trusted(user, project) == 'true') { nil } + else { + ignored = project_ignored_keys(user, project) + if (length(ignored) == 0) { nil } + else { + "./.swarm-code.json: ignored untrusted project settings (" ++ + join_names(ignored, "") ++ ") — if you trust this repo, run " ++ + "`swarm-code trust` here (or /trust) and restart" + } + } + } +} + +# ------------------------------------------------------------ +# Network isolation — may this LLM endpoint URL be dialed? +# ------------------------------------------------------------ +# The ONE local-network check, used by main.sw's startup gate (primary +# endpoint) and by llm.sw at the point of dial, so fallback, providers[] +# and ~/.swarm-code/.profile_override endpoints get the same treatment. +# The URL is read the way curl will read it and anything ambiguous is +# refused: +# * scheme http:// or https:// only (any case) +# * no userinfo: curl dials "http://127.0.0.1@evil" at evil, so an '@' +# anywhere in the authority is refused outright +# * host lowercased, port (digits only) stripped, host chars limited +# to [a-z0-9._-] — no %-escapes, braces or other curl-isms +# * IPv6 only in brackets, and only ::1, fc00::/7 and fe80::/10 +# * a host whose LAST label is numeric (or 0x…) is an IPv4 literal and +# must be a strict dotted quad: curl dials 3221225985, 0x7f.1 and +# 010.0.0.1 (octal) as other addresses than they appear to be +# * names: localhost, *.local (mDNS), *.ts.net, or a bare dot-less +# name (MagicDNS / /etc/hosts) — "127.0.0.1.evil.com" is a name +# IPv4: loopback 127/8, RFC1918 10/8 172.16/12 192.168/16, CGNAT 100.64/10. +# ------------------------------------------------------------ + +# nil when `url` may be dialed, else a one-line refusal reason. The only +# opt-out is SWARM_CODE_ALLOW_REMOTE=1 — an API key alone is not one. +fun endpoint_refusal(url) { + if (getenv("SWARM_CODE_ALLOW_REMOTE") == "1") { nil } + else { if (is_local_endpoint(url) == 'true') { nil } + else { + host = endpoint_host(url) + why = if (host == nil) { "not a plain http(s)://host[:port] URL" } + else { "host " ++ host ++ " is not on the local network" } + "network isolation: refusing to contact " ++ to_string(url) ++ " (" ++ why ++ + ") — set SWARM_CODE_ALLOW_REMOTE=1 to allow remote endpoints" + }} +} + +fun is_local_endpoint(url) { + host = endpoint_host(url) + if (host == nil) { 'false' } else { is_local_host(host) } +} + +# Lowercased host of an http(s) URL (an IPv6 literal without brackets), +# or nil when the URL is not a plain http(s)://host[:port][/...] form. +fun endpoint_host(url) { + s = string_lower(string_trim(to_string(url))) + rest = if (string_starts_with(s, "http://") == 'true') { string_sub(s, 7, string_length(s) - 7) } + else { if (string_starts_with(s, "https://") == 'true') { string_sub(s, 8, string_length(s) - 8) } + else { nil }} + if (rest == nil) { nil } + else { + auth = authority_of(rest, 0, string_length(rest)) + if (string_length(auth) == 0 || string_contains(auth, "@") == 'true') { nil } + else { host_of_authority(auth) } + } +} + +# Everything before the first '/', '?' or '#'. +fun authority_of(s, i, n) { + if (i >= n) { s } + else { + ch = string_sub(s, i, 1) + if (ch == "/" || ch == "?" || ch == "#") { string_sub(s, 0, i) } + else { authority_of(s, i + 1, n) } + } +} + +fun host_of_authority(a) { + if (string_starts_with(a, "[") == 'true') { + close = string_index_of(a, "]") + if (close < 0) { nil } + else { + inner = string_sub(a, 1, close - 1) + port_part = string_sub(a, close + 1, string_length(a) - close - 1) + if (valid_port_suffix(port_part) == 'true' && string_contains(inner, ":") == 'true' && + all_chars_in(inner, "0123456789abcdef:.") == 'true') { inner } + else { nil } + } + } else { + colon = string_index_of(a, ":") + host = if (colon < 0) { a } else { string_sub(a, 0, colon) } + port_part = if (colon < 0) { "" } else { string_sub(a, colon, string_length(a) - colon) } + if (valid_port_suffix(port_part) == 'true' && valid_hostname(host) == 'true') { host } + else { nil } + } +} + +# "" or ':' followed by 1-5 digits. +fun valid_port_suffix(s) { + n = string_length(s) + if (n == 0) { 'true' } + else { if (string_starts_with(s, ":") == 'true' && n >= 2 && n <= 6) { + all_chars_in(string_sub(s, 1, n - 1), "0123456789") + } else { 'false' }} +} + +# Non-empty, dot-separated labels of [a-z0-9_-]. +fun valid_hostname(h) { + if (string_length(h) == 0) { 'false' } + else { if (all_chars_in(h, "abcdefghijklmnopqrstuvwxyz0123456789.-_") == 'false') { 'false' } + else { if (string_starts_with(h, ".") == 'true' || string_ends_with(h, ".") == 'true' || + string_contains(h, "..") == 'true') { 'false' } + else { 'true' }}} +} + +fun all_chars_in(s, allowed) { aci_loop(s, allowed, 0, string_length(s)) } + +fun aci_loop(s, allowed, i, n) { + if (i >= n) { 'true' } + else { if (string_contains(allowed, string_sub(s, i, 1)) == 'true') { aci_loop(s, allowed, i + 1, n) } + else { 'false' }} +} + +# `h` is a validated, lowercased host from endpoint_host. +fun is_local_host(h) { + if (string_contains(h, ":") == 'true') { is_local_ipv6(h) } + else { if (ipv4_like(h) == 'true') { is_private_ipv4(h) } + else { is_local_name(h) }} +} + +# The last label decides: numeric (or 0x-hex) means curl parses the whole +# host as an IPv4 address, however few dots it has. +fun ipv4_like(h) { + last = last_label(h, string_length(h) - 1) + if (string_starts_with(last, "0x") == 'true') { 'true' } + else { all_chars_in(last, "0123456789") } +} + +fun last_label(h, i) { + if (i < 0) { h } + else { if (string_sub(h, i, 1) == ".") { string_sub(h, i + 1, string_length(h) - i - 1) } + else { last_label(h, i - 1) }} +} + +fun is_private_ipv4(h) { + parts = string_split(h, ".") + if (length(parts) != 4) { 'false' } + else { if (strict_octets(parts) == 'false') { 'false' } + else { + a = to_int(hd(parts)) + b = to_int(hd(tl(parts))) + if (a == 127 || a == 10) { 'true' } + else { if (a == 192 && b == 168) { 'true' } + else { if (a == 172 && b >= 16 && b <= 31) { 'true' } + else { if (a == 100 && b >= 64 && b <= 127) { 'true' } + else { 'false' }}}} + }} +} + +# Decimal 0-255, 1-3 digits, no leading zero (curl reads 010 as octal 8). +fun strict_octets(parts) { + if (length(parts) == 0) { 'true' } + else { + p = hd(parts) + n = string_length(p) + ok = if (n < 1 || n > 3) { 'false' } + else { if (all_chars_in(p, "0123456789") == 'false') { 'false' } + else { if (n > 1 && string_starts_with(p, "0") == 'true') { 'false' } + else { if (to_int(p) > 255) { 'false' } else { 'true' }}}} + if (ok == 'true') { strict_octets(tl(parts)) } else { 'false' } + } +} + +# Loopback, ULA fc00::/7, link-local fe80::/10. The first group must be +# written out in full: "fd::1" is 00fd::1 — a public address. +fun is_local_ipv6(h) { + if (h == "::1" || h == "0:0:0:0:0:0:0:1") { 'true' } + else { + colon = string_index_of(h, ":") + first = if (colon < 0) { h } else { string_sub(h, 0, colon) } + if (string_length(first) != 4) { 'false' } + else { + p2 = string_sub(first, 0, 2) + p3 = string_sub(first, 0, 3) + if (p2 == "fc" || p2 == "fd") { 'true' } + else { if (p3 == "fe8" || p3 == "fe9" || p3 == "fea" || p3 == "feb") { 'true' } + else { 'false' }} + } + } +} + +fun is_local_name(h) { + if (h == "localhost") { 'true' } + else { if (string_ends_with(h, ".local") == 'true') { 'true' } + else { if (string_ends_with(h, ".ts.net") == 'true') { 'true' } + # Bare (dot-less) name: mDNS / Tailscale MagicDNS / /etc/hosts. It can't + # be told apart from a public bare host without DNS; on dev machines the + # local case is overwhelmingly the common one. + else { string_contains(h, ".") == 'false' }}} +} + # ------------------------------------------------------------ # llm_timeout_ms — inactivity window (ms) for the worker-routed LLM # stream (llm.sw stream_call). If no stream message (chunk / reason / @@ -127,7 +592,7 @@ fun load_project_context() { # Permissions — decide whether a tool call should run. # Returns an atom: 'allow', 'deny', or 'ask'. # -# Policy (updated): +# Policy: # 1. Default-allow for every BUILT-IN tool. The user explicitly asked # for "all allowed by default" — prompting on every bash/write/edit # was breaking flow during tool test runs. The one exception is MCP @@ -136,49 +601,55 @@ fun load_project_context() { # 2. settings.permissions[tool_name] in settings.json can downgrade # a specific tool to 'ask' or 'deny' if the user wants tighter # control on one tool (e.g. "bash": "ask"). -# 3. The dangerous-bash hard gate still fires regardless. Commands -# that look like `rm -rf`, `sudo`, `curl | sh`, `mkfs`, force -# push, or hard resets still prompt even in default-allow mode. -# We are not giving a model root access to the box. +# 3. Every tool that runs a model-supplied shell command (command_of: +# bash, background, bg_server, run_tests.command) goes through the +# CommandGuard classifier, which parses the command like sh does: +# * hardline (rm -r on /, mkfs, dd to a raw disk, halt/reboot, +# chmod -R on /, fork bomb) → 'deny', unconditionally — before +# any settings/env lookup, so SWARM_CODE_ALLOW_DANGEROUS=1 +# cannot turn it off. +# * dangerous (sudo, rm -rf on ~ or a system path, dd to a +# device) → escalates 'allow' to 'ask' (a configured 'deny' +# stays 'deny'); 'deny' outright when SWARM_CODE_DENY_DANGEROUS=1 +# (unattended /flows children); skipped entirely when +# SWARM_CODE_ALLOW_DANGEROUS=1. +# denial_message() names the matched pattern so the model can adapt. # ------------------------------------------------------------ fun check_permission(tool_name, args, opts) { - # HARDLINE: unbypassable deny for catastrophic patterns (mkfs, dd - # to disk, shutdown/reboot, fork bomb, rm -rf /*). Fires BEFORE - # any settings/env lookup so SWARM_CODE_ALLOW_DANGEROUS=1 cannot - # turn it off. See is_hardline_bash for the pattern list. - if (tool_name == 'bash' && is_hardline_bash(args) == 'true') { + risk = command_risk(tool_name, args) + level = elem(risk, 0) + if (level == 'hardline') { 'deny' } else { - settings = map_get(opts, 'settings') - perms = if (settings == nil) { nil } else { map_get(settings, 'permissions') } - - # settings.json is decoded with atom keys, so pass tool_name directly. - configured = if (perms == nil) { - nil - } else { - map_get(perms, tool_name) - } - - decision = if (configured != nil) { - string_to_perm(configured) - } else { - default_permission(tool_name) - } + decision = configured_decision(tool_name, opts) - # Hard-gate dangerous bash commands regardless of config. - # Headless converts 'ask' to 'allow' (agent.resolve_permission), - # so unattended children (/flows fan-out sets - # SWARM_CODE_DENY_DANGEROUS=1) turn this gate into a hard deny - # instead of silently auto-approving. - if (tool_name == 'bash' && is_dangerous_bash(args) == 'true') { - if (getenv("SWARM_CODE_DENY_DANGEROUS") == "1") { 'deny' } else { 'ask' } + # Headless denies an 'ask' unless SWARM_CODE_HEADLESS_APPROVE=1 + # (agent.resolve_permission); SWARM_CODE_DENY_DANGEROUS=1 (set for + # /flows fan-out and scheduled children) makes this a hard deny + # even then. + if (level == 'dangerous' && dangerous_bypassed() == 'false') { + if (decision == 'deny') { 'deny' } + else { if (getenv("SWARM_CODE_DENY_DANGEROUS") == "1") { 'deny' } else { 'ask' } } } else { decision } } } +# settings.permissions[tool] if set, else the built-in default. +fun configured_decision(tool_name, opts) { + settings = if (opts == nil) { nil } else { map_get(opts, 'settings') } + perms = if (settings == nil) { nil } else { map_get(settings, 'permissions') } + # settings.json is decoded with atom keys, so pass tool_name directly. + configured = if (perms == nil) { nil } else { map_get(perms, tool_name) } + if (configured != nil) { string_to_perm(configured) } else { default_permission(tool_name) } +} + +fun dangerous_bypassed() { + if (getenv("SWARM_CODE_ALLOW_DANGEROUS") == "1") { 'true' } else { 'false' } +} + fun string_to_perm(s) { if (s == "allow") { 'allow' } else { if (s == "deny") { 'deny' } @@ -197,142 +668,67 @@ fun default_permission(tool_name) { else { 'allow' } } -# Return 'true' ONLY for truly catastrophic, unambiguous patterns. -# Scoped down from a broad "destructive commands" net because the old -# version was flagging perfectly normal dev workflows like -# `rm -rf ./build-dir`, `git push --force` on feature branches, and -# `git reset --hard HEAD~1`. The model is a coding assistant; those -# are its daily bread. -# -# What still trips the gate (after much narrowing): -# * rm -rf targeting `/` or `~` or `$HOME` literally -# * mkfs (formatting a block device) -# * dd if=... writing to /dev/disk, /dev/sd, /dev/nvme, /dev/rdisk -# * sudo (privilege escalation is always worth a beat) -# -# You can fully disable even this minimal gate by exporting -# SWARM_CODE_ALLOW_DANGEROUS=1 before launching swarm-code. Everything -# runs, nothing prompts. YOLO mode. -fun is_dangerous_bash(args) { - bypass = getenv("SWARM_CODE_ALLOW_DANGEROUS") - if (bypass == "1") { 'false' } - else { - cmd = map_get(args, 'command') - if (cmd == nil) { 'false' } - else { - # rm targeting the filesystem root or user home literally. - # We look for "rm " ++ anything ++ " /" at word boundary - # rather than the broad "rm -rf" string match. A simple - # conservative approach: flag only the specific dangerous - # literal suffixes. - if (string_contains(cmd, "rm -rf /") == 'true' && - string_contains(cmd, "rm -rf /tmp") == 'false' && - string_contains(cmd, "rm -rf /var/") == 'false' && - string_contains(cmd, "rm -rf /Users/") == 'false' && - string_contains(cmd, "rm -rf /home/") == 'false' && - string_contains(cmd, "rm -rf /opt/") == 'false') { 'true' } - else { if (string_contains(cmd, "rm -rf ~") == 'true') { 'true' } - else { if (string_contains(cmd, "rm -rf $HOME") == 'true') { 'true' } - else { if (string_contains(cmd, "sudo ") == 'true') { 'true' } - else { if (string_contains(cmd, "mkfs") == 'true') { 'true' } - else { if (string_contains(cmd, "dd if=") == 'true' && - string_contains(cmd, "of=/dev/") == 'true') { 'true' } - else { 'false' }}}}}} - } - } +# The shell command a tool call will run, or nil for tools that don't run +# one. Every tool listed here gets the same hardline/dangerous gate as bash — +# before, `background` / `bg_server` / `run_tests.command` ran ungated. +fun command_of(tool_name, args) { + t = to_string(tool_name) + if (args == nil || is_map(args) == 'false') { nil } + else { if (t == "bash" || t == "background" || t == "bg_server" || t == "run_tests") { + c = map_get(args, 'command') + if (c == nil) { nil } else { to_string(c) } + } else { nil }} } -# ------------------------------------------------------------ -# HARDLINE blocklist — UNBYPASSABLE bash patterns. -# ------------------------------------------------------------ -# Unlike is_dangerous_bash, this CANNOT be turned off with -# SWARM_CODE_ALLOW_DANGEROUS=1. If your agent is asking to mkfs a -# disk or reboot the box, no env var should let it through. -# -# Categories: -# * Filesystem destruction: mkfs, mkswap -# * Disk wipe: dd if=... of=/dev/{sd,nvme,disk,rdisk} -# * System halt: shutdown, reboot, halt, poweroff, init 0, init 6 -# * Filesystem lockout: chmod 000 /, chown -R 0:0 / -# * Fork bomb literal: :(){:|:&};: -# * Whole-disk rm: rm -rf /* -# ------------------------------------------------------------ -fun is_hardline_bash(args) { - cmd = map_get(args, 'command') - if (cmd == nil) { 'false' } - else { - s = to_string(cmd) - # Filesystem destruction - if (string_contains(s, "mkfs") == 'true') { 'true' } - else { if (string_contains(s, "mkswap") == 'true') { 'true' } - # dd writing to a raw disk node - else { if (string_contains(s, "dd if=") == 'true' && - string_contains(s, "of=/dev/sd") == 'true') { 'true' } - else { if (string_contains(s, "dd if=") == 'true' && - string_contains(s, "of=/dev/nvme") == 'true') { 'true' } - else { if (string_contains(s, "dd if=") == 'true' && - string_contains(s, "of=/dev/disk") == 'true') { 'true' } - else { if (string_contains(s, "dd if=") == 'true' && - string_contains(s, "of=/dev/rdisk") == 'true') { 'true' } - # System halt — matched as whole command words (not bare substrings), - # so `cat asphalt_survey.csv` / `vim shutdown_handler.py` are NOT - # blocked while `shutdown -h now`, `/sbin/reboot`, `poweroff` still are. - else { if (contains_command_word(s, "shutdown") == 'true') { 'true' } - else { if (contains_command_word(s, "reboot") == 'true') { 'true' } - else { if (contains_command_word(s, "halt") == 'true') { 'true' } - else { if (contains_command_word(s, "poweroff") == 'true') { 'true' } - else { if (contains_command_word(s, "init 0") == 'true') { 'true' } - else { if (contains_command_word(s, "init 6") == 'true') { 'true' } - # telinit N is the SysV alias (telinit 0 halts, telinit 6 reboots) — - # word-boundary "init 0" misses it ("init" preceded by 'l'), so match - # the verb directly. Keep this as long as "init 0"/"init 6" are blocked. - else { if (contains_command_word(s, "telinit") == 'true') { 'true' } - # Filesystem lockout - else { if (string_contains(s, "chmod 000 /") == 'true') { 'true' } - else { if (string_contains(s, "chown -R 0:0 /") == 'true') { 'true' } - # Fork bomb - else { if (string_contains(s, ":(){:|:&};:") == 'true') { 'true' } - # Whole-disk wipe - else { if (string_contains(s, "rm -rf /*") == 'true') { 'true' } - else { 'false' }}}}}}}}}}}}}}}}} - } +# {'hardline'|'dangerous'|'ok', reason} for a tool call. +fun command_risk(tool_name, args) { + cmd = command_of(tool_name, args) + if (cmd == nil) { {'ok', ""} } else { CommandGuard.risk_of(cmd) } } -# Whole-word match for a catastrophic verb: the word must be bounded by a -# non-identifier char (or string edge) on both sides, so it isn't matched as -# a substring of a larger filename/identifier (asphalt, rebooter, -# shutdown_handler). Over-blocks rare cases like `cat shutdown.sh` — the safe -# direction for an unbypassable floor (never under-blocks a real `shutdown`). -fun contains_command_word(s, word) { - cw_scan(s, word, string_length(word), string_length(s), 0) -} +fun uses_sudo(cmd) { CommandGuard.uses_sudo(cmd) } -fun cw_scan(s, word, wlen, slen, i) { - if (i + wlen > slen) { 'false' } - else { - if (string_sub(s, i, wlen) == word) { - prev_ch = cw_char_at(s, i - 1, slen) - next_ch = cw_char_at(s, i + wlen, slen) - if (cw_boundary(prev_ch) == 'true' && cw_boundary(next_ch) == 'true') { 'true' } - else { cw_scan(s, word, wlen, slen, i + 1) } - } else { cw_scan(s, word, wlen, slen, i + 1) } - } +# "error: permission denied for tool 'X' — ". The why names the +# matched pattern (and the offending simple command), so the model can +# pick a narrower command instead of retrying the same call. +fun denial_message(tool_name, args, opts) { + "error: permission denied for tool '" ++ to_string(tool_name) ++ "'" ++ + denial_reason(tool_name, args, opts) } -fun cw_char_at(s, idx, slen) { - if (idx < 0 || idx >= slen) { "" } - else { string_sub(s, idx, 1) } +fun denial_reason(tool_name, args, opts) { + risk = command_risk(tool_name, args) + level = elem(risk, 0) + if (level == 'hardline') { + " — blocked by the hardline safety floor: " ++ elem(risk, 1) ++ + ". This is never allowed (no setting or env var lifts it); use a narrower command." + } else { if (configured_decision(tool_name, opts) == 'deny') { + " — denied by settings.json (permissions." ++ to_string(tool_name) ++ " = \"deny\")." + } else { if (level == 'dangerous' && dangerous_bypassed() == 'false' && + getenv("SWARM_CODE_DENY_DANGEROUS") == "1") { + " — flagged dangerous: " ++ elem(risk, 1) ++ + ". SWARM_CODE_DENY_DANGEROUS=1 is set for this unattended run, so it is denied; use a narrower command." + } else { if (level == 'dangerous' && dangerous_bypassed() == 'false') { + " — flagged dangerous: " ++ elem(risk, 1) ++ "; it was not approved." + } else { + " — not approved at the permission prompt (don't retry the same call; ask the user or try another approach)." + }}}} } -fun cw_boundary(ch) { - if (ch == "") { 'true' } - else { if (cw_is_ident(ch) == 'true') { 'false' } else { 'true' }} +# Kept for callers/tests that ask about a bash args map directly. +# 'true' for dangerous OR hardline commands (a hardline command is certainly +# dangerous); 'false' when SWARM_CODE_ALLOW_DANGEROUS=1 (YOLO mode). +fun is_dangerous_bash(args) { + if (dangerous_bypassed() == 'true') { 'false' } + else { + level = elem(command_risk('bash', args), 0) + if (level == 'ok') { 'false' } else { 'true' } + } } -fun cw_is_ident(ch) { - if ((ch >= "a" && ch <= "z") || (ch >= "A" && ch <= "Z") - || (ch >= "0" && ch <= "9") || ch == "_") { 'true' } - else { 'false' } +# 'true' for the unbypassable tier (see CommandGuard for the categories). +fun is_hardline_bash(args) { + if (elem(command_risk('bash', args), 0) == 'hardline') { 'true' } else { 'false' } } # ------------------------------------------------------------ @@ -340,21 +736,38 @@ fun cw_is_ident(ch) { # # Settings shape: # "hooks": { -# "PreToolUse": [ {"matcher": "bash", "command": "..."} ], -# "PostToolUse": [ {"matcher": "edit|write", "command": "..."} ], +# "PreToolUse": [ {"matcher": "bash", "command": "..."} ], (+ background, bg_server, run_tests, …) +# "PostToolUse": [ {"matcher": "edit|write", "command": "..."} ], (+ multi_edit) # "UserPromptSubmit": [ {"command": "..."} ], # "Stop": [ {"command": "..."} ] # } # -# Matcher is a literal substring of the tool name, or "*" for all. -# Hooks receive context through environment variables set before shell(): -# SWARM_CODE_EVENT, SWARM_CODE_TOOL, SWARM_CODE_ARGS +# Matcher: "*" (or none) for all tools, else "|"-separated alternatives, +# each a case-insensitive substring of the tool name or a tool FAMILY — +# "bash" also fires for background/bg_server/run_tests/file_watch/log_wait, +# "edit" also for multi_edit. See matches() for the exact rules. +# Hooks receive context through the environment and stdin: +# SWARM_CODE_EVENT, SWARM_CODE_TOOL — event and tool name +# stdin, and the private 0600 file $SWARM_CODE_ARGS_FILE — the args JSON +# (always; read one of these to see every call) +# SWARM_CODE_ARGS — the same JSON inline, only when it is under 100KB; +# above that it is unset and SWARM_CODE_ARGS_OMITTED=1 +# (The args used to be spliced into the command line + env: a >128KB +# write hit E2BIG, the hook never ran, and shell() polled 120s before +# reporting a block.) Hooks time out after hook_cmd_timeout_ms(). # # Returns 'ok' normally. Returns 'block' if any PreToolUse hook exited -# non-zero (blocking the tool call). +# non-zero, timed out, or could not be run at all (fail closed) — +# blocking the tool call. run_hooks_verdict also says why. # ------------------------------------------------------------ fun run_hooks(event, tool_name, args_json, opts) { - settings = map_get(opts, 'settings') + v = run_hooks_verdict(event, tool_name, args_json, opts) + if (v == 'ok') { 'ok' } else { 'block' } +} + +# 'ok' | {'block', reason} +fun run_hooks_verdict(event, tool_name, args_json, opts) { + settings = if (opts == nil) { nil } else { map_get(opts, 'settings') } if (settings == nil) { 'ok' } else { hooks = map_get(settings, 'hooks') @@ -411,8 +824,8 @@ fun run_matching_hooks(hooks_list, tool_name, args_json, event) { } else { if (matches(matcher, tool_name) == 'true') { result = run_hook_cmd(cmd, event, tool_name, args_json) - if (result == 'block') { - 'block' + if (result != 'ok') { + result } else { run_matching_hooks(tl(hooks_list), tool_name, args_json, event) } @@ -423,41 +836,88 @@ fun run_matching_hooks(hooks_list, tool_name, args_json, event) { } } +# Matcher semantics (PreToolUse / PostToolUse): +# nil, "", "*" → every tool +# "a|b" / "a,b" → any alternative matches +# an alternative matches a tool — case-insensitively, so Claude-Code-style +# "Bash" / "Edit|Write" / "MultiEdit" work — when it is a substring of the +# tool name (so "browser" covers every browser_* tool and "mcp__github" +# every tool of that server; underscores are ignored, "WebFetch" ≈ +# "web_fetch"), or when it names a FAMILY the tool belongs to: +# bash (or shell) → bash, background, bg_server, run_tests, file_watch, +# log_wait — every tool that runs a shell command or +# shell poll loop, so a "bash" guard hook can't be +# sidestepped by calling `background` instead +# edit → edit, multi_edit +# multiedit → multi_edit +# (The old check was "matcher ⊂ tool or tool ⊂ matcher" on the raw string: +# "edit|write" never fired for multi_edit, "bash" never for background.) fun matches(matcher, tool_name) { if (matcher == nil) { 'true' } else { - if (matcher == "*") { 'true' } + m = string_lower(string_trim(to_string(matcher))) + if (m == "" || m == "*") { 'true' } else { - # Two semantics, both supported via "either direction" - # check: - # 1. Pipe-alternation: matcher "bash|edit" matches tool - # "bash" because tool_name is a substring of matcher. - # 2. Substring: matcher "ed" matches tool "edit" because - # matcher is a substring of tool_name. - # The original code only did (1); the audit miscalled this - # as a bug. Doing both makes the obvious matchers work - # whichever way the user expected. - m = to_string(matcher) - t = to_string(tool_name) - if (string_contains(m, t) == 'true') { 'true' } - else { string_contains(t, m) } + alts = string_split(string_replace(m, ",", "|"), "|") + any_alt_matches(alts, string_lower(to_string(tool_name))) } } } -# Run a single hook command. Wrap with env exports for context. -# If the command exits non-zero, treat as a block signal. -# Args JSON is exposed as SWARM_CODE_ARGS so hooks can inspect the -# tool payload (e.g. a `bash` hook that greps the command). Quoted -# with shell_q_local because args_json contains arbitrary JSON -# (including single quotes inside strings). +fun any_alt_matches(alts, t) { + if (length(alts) == 0) { 'false' } + else { if (alt_matches(string_trim(hd(alts)), t) == 'true') { 'true' } + else { any_alt_matches(tl(alts), t) }} +} + +fun alt_matches(a, t) { + if (a == "") { 'false' } + else { if (a == "*") { 'true' } + else { if (hook_list_has(hook_family(a), t) == 'true') { 'true' } + else { if (string_contains(t, a) == 'true') { 'true' } + else { string_contains(string_replace(t, "_", ""), string_replace(a, "_", "")) }}}} +} + +fun hook_family(a) { + if (a == "bash" || a == "shell") { + ["bash", "background", "bg_server", "run_tests", "file_watch", "log_wait"] + } else { if (a == "edit") { ["edit", "multi_edit"] } + else { if (a == "multiedit") { ["multi_edit"] } + else { [] }}} +} + +fun hook_list_has(lst, item) { + if (length(lst) == 0) { 'false' } + else { if (hd(lst) == item) { 'true' } else { hook_list_has(tl(lst), item) } } +} + +# Run a single hook command with the context described above run_hooks. +# 'ok' on exit 0; {'block', reason} on a non-zero exit, a timeout, or when +# the hook could not be run at all — the reason carries the exit code and +# the start of the hook's output (stdout+stderr) so the model can see why. +fun hook_cmd_timeout_ms() { 60000 } + fun run_hook_cmd(cmd, event, tool_name, args_json) { - args_safe = Util.shell_q(to_string(args_json)) - full = "export SWARM_CODE_EVENT=" ++ Util.shell_q(to_string(event)) ++ "; " ++ - "export SWARM_CODE_TOOL=" ++ Util.shell_q(to_string(tool_name)) ++ "; " ++ - "export SWARM_CODE_ARGS=" ++ args_safe ++ "; " ++ - cmd - result = shell(full) - code = elem(result, 0) - if (code == 0) { 'ok' } else { 'block' } + exports = "export SWARM_CODE_EVENT=" ++ Util.shell_q(to_string(event)) ++ "; " ++ + "export SWARM_CODE_TOOL=" ++ Util.shell_q(to_string(tool_name)) ++ ";" + res = Hooks.run_with_payload(cmd, to_string(args_json), "SWARM_CODE_ARGS", exports, + hook_cmd_timeout_ms(), hook_tmp_prefix()) + label = to_string(event) ++ " hook `" ++ string_truncate(to_string(cmd), 80) ++ "`" + if (map_get(res, 'ran') != 'true') { + {'block', label ++ " could not run — " ++ to_string(map_get(res, 'error'))} + } else { if (map_get(res, 'interrupted') == 'true') { + {'block', label ++ " timed out after " ++ to_string(hook_cmd_timeout_ms() / 1000) ++ "s"} + } else { if (map_get(res, 'code') == 0) { + 'ok' + } else { + out = string_trim(to_string(map_get(res, 'out'))) + tail = if (string_length(out) == 0) { "" } else { ": " ++ string_truncate(out, 500) } + {'block', label ++ " exited " ++ to_string(map_get(res, 'code')) ++ tail} + }}} +} + +fun hook_tmp_prefix() { + t = getenv("TMPDIR") + base = if (t == nil || string_length(to_string(t)) == 0) { "/tmp" } else { to_string(t) } + base ++ "/swarm-code-hook-" } diff --git a/src/llm.sw b/src/llm.sw index 5ead1d0..74759ed 100644 --- a/src/llm.sw +++ b/src/llm.sw @@ -6,6 +6,7 @@ import Markdown import Hooks import Config import UI +import Prompts # ============================================================ # LLM — OpenAI-compatible chat completions client @@ -43,13 +44,17 @@ export [ last_prompt_tokens, last_reasoning, last_fail, record_usage, record_reasoning, extract_usage, extract_reasoning, extract_content, extract_finish_reason, + last_fail_detail, last_fail_context_overflow, is_context_overflow_msg, new_message_system, new_message_user, new_message_assistant, new_message_tool, parse_inband_tool_calls, inband_assistant_text, api_tool_calls_to_internal, repair_history, apply_override, + read_override, apply_override_map, override_applies, effective_override, + retarget_system_prompt, inject_context_status, build_status_string, - routed_collect, maybe_large_context_hint + routed_collect, maybe_large_context_hint, + context_window_tokens, context_budget_tokens, budget_for_window ] # ============================================================ @@ -85,55 +90,90 @@ fun new_message_tool(tool_call_id, content) { # URL ending in `/chat/completions` and we'll use it verbatim. # ------------------------------------------------------------ # ------------------------------------------------------------ -# In-session profile override. Slash command /profile NAME writes -# ~/.swarm-code/.profile_override with {endpoint, model, api_key, -# tool_format}; every LLM call consults this file and applies the -# overrides before serialising the request. File-based (rather than -# threading opts through main_loop) so it's a 5-line touch instead of -# a deep refactor — `rm ~/.swarm-code/.profile_override` reverts. +# Profile override. /profile NAME and /model NAME write +# ~/.swarm-code/.profile_override ({endpoint, model, api_key, tool_format, +# vision, chat_template_kwargs, profile, session_id}); every LLM call +# consults it before serialising the request. File-based (rather than +# threading opts through main_loop); /profile-clear or deleting the file +# reverts. +# +# Precedence: the file is global and outlives the session that wrote it. +# An override written by THIS session (its session_id is opts.session_id, +# minted per launch in main.sw) is the user's explicit in-session choice and +# beats everything. One left over from an EARLIER session is only a default: +# an env var that is set (SWARM_CODE_MODEL, …) wins over it, per field — a +# stale `/profile qwen3-b` used to silently replace SWARM_CODE_MODEL=test. # ------------------------------------------------------------ fun apply_override(opts) { + ov = read_override() + if (ov == nil) { opts } else { apply_override_map(opts, ov, fn(name) { getenv(name) }) } +} + +fun read_override() { home = getenv("HOME") - if (home == nil) { opts } + if (home == nil) { nil } else { path = home ++ "/.swarm-code/.profile_override" - if (file_exists(path) == 'false') { opts } + if (file_exists(path) == 'false') { nil } else { raw = file_read(path) - if (raw == nil) { opts } - else { - ov = json_decode(raw) - if (ov == nil) { opts } - else { - a = override_field(opts, ov, 'endpoint') - b = override_field(a, ov, 'model') - c = override_field(b, ov, 'api_key') - d = override_field(c, ov, 'tool_format') - d = override_field(d, ov, 'vision') - # chat_template_kwargs is a map (not a string), so it - # bypasses override_field's to_string coercion. Always - # replace, even with nil, so switching to a profile that - # doesn't set thinking-off actually clears it. - ct_v = map_get(ov, 'chat_template_kwargs') - d = map_put(d, 'chat_template_kwargs', ct_v) - # Re-derive temperature when model changes — Kimi K2.x - # rejects any temperature other than 1.0. - if (map_get(ov, 'model') != nil) { - new_model = to_string(map_get(d, 'model')) - new_temp = if (string_starts_with(new_model, "kimi") == 'true') { 1.0 } - else { 0.2 } - map_put(d, 'temperature', new_temp) - } else { d } - } - } + if (raw == nil) { nil } else { json_decode(raw) } } } } -fun override_field(opts, ov, key) { - v = map_get(ov, key) - if (v == nil) { opts } +# The env var that pins each overridable field (nil: none). +fun override_env_var(key) { + if (key == 'endpoint') { "SWARM_CODE_ENDPOINT" } + else { if (key == 'model') { "SWARM_CODE_MODEL" } + else { if (key == 'api_key') { "SWARM_CODE_API_KEY" } + else { if (key == 'tool_format') { "SWARM_CODE_TOOL_FORMAT" } + else { nil }}}} +} + +# Does the override's `key` take effect? It is in the file AND (this session +# wrote it, or no env var pins the field). env_fn: name -> value | nil (the +# real getenv, or a stub in tests). +fun override_applies(opts, ov, key, env_fn) { + if (map_get(ov, key) == nil) { 'false' } else { + sid = map_get(ov, 'session_id') + mine = if (sid != nil && to_string(sid) == to_string(map_get(opts, 'session_id'))) { 'true' } else { 'false' } + var = override_env_var(key) + if (mine == 'true' || var == nil) { 'true' } + else { if (env_fn(var) == nil) { 'true' } else { 'false' } } + } +} + +fun apply_override_map(opts, ov, env_fn) { + a = override_field(opts, ov, 'endpoint', env_fn) + b = override_field(a, ov, 'model', env_fn) + c = override_field(b, ov, 'api_key', env_fn) + d = override_field(c, ov, 'tool_format', env_fn) + e = override_field(d, ov, 'vision', env_fn) + # chat_template_kwargs is a map (not a string), so it bypasses + # override_field's to_string coercion. A /profile swap (the file names its + # profile) REPLACES it — even with nil, so switching to a profile without + # thinking-off really clears the launch profile's; the profile's own + # kwargs are now written to the file (they used to be dropped, so + # `/profile qwen2` lost enable_thinking:false). A /model-only override + # leaves it alone. + ct = map_get(ov, 'chat_template_kwargs') + f = if (map_get(ov, 'profile') != nil || ct != nil) { map_put(e, 'chat_template_kwargs', ct) } else { e } + # Re-derive temperature when the model changes — Kimi K2.x rejects any + # temperature other than 1.0. + if (override_applies(opts, ov, 'model', env_fn) == 'true') { + new_model = to_string(map_get(f, 'model')) + new_temp = if (string_starts_with(new_model, "kimi") == 'true') { 1.0 } + else { 0.2 } + map_put(f, 'temperature', new_temp) + } else { f } +} + +fun override_field(opts, ov, key, env_fn) { + if (override_applies(opts, ov, key, env_fn) == 'false') { opts } + else { + v = map_get(ov, key) val = if (key == 'tool_format') { s = to_string(v) if (s == "native") { 'native' } else { 'inband' } @@ -142,6 +182,29 @@ fun override_field(opts, ov, key) { } } +# The part of an override that takes effect right now, for /model to carry +# forward when it restamps the file as this session's: it changes ONLY the +# model. Keeping env-shadowed fields would promote them over env; dropping the +# active ones silently reverted a /profile's endpoint and api_key (while +# printing "endpoint/api_key unchanged"). +fun effective_override(ov, opts) { + if (ov == nil) { map_new() } + else { + env_fn = fn(name) { getenv(name) } + keep_applying(ov, opts, ['endpoint', 'model', 'api_key', 'tool_format', 'vision', + 'chat_template_kwargs', 'profile'], env_fn, map_new()) + } +} + +fun keep_applying(ov, opts, keys, env_fn, acc) { + if (length(keys) == 0) { acc } + else { + k = hd(keys) + next = if (override_applies(opts, ov, k, env_fn) == 'true') { map_put(acc, k, map_get(ov, k)) } else { acc } + keep_applying(ov, opts, tl(keys), env_fn, next) + } +} + fun chat_completions_url(endpoint) { # Strip any trailing slash so suffix checks are deterministic. base = if (string_ends_with(endpoint, "/") == 'true') { @@ -367,7 +430,7 @@ fun build_request_body(messages, opts) { # trailing unmatched tool_calls, and collapses consecutive user # messages. Keeps the API from rejecting requests after a crash # mid-turn or a journal load that landed mid-conversation. - repaired = repair_history(messages) + repaired = repair_history(retarget_system_prompt(messages, tool_format)) # Drain any pending image attachments (from read_image) and inject # them into the LAST user message as OpenAI multimodal content @@ -416,6 +479,36 @@ fun build_request_body(messages, opts) { json_encode(final_req) } +# ------------------------------------------------------------ +# System prompt ↔ wire format +# ------------------------------------------------------------ +# main.sw builds the system prompt ONCE, for the launch tool_format. A +# /profile override, a fallback endpoint or a provider chain can send a +# request in the OTHER format — and a native prompt ("your tools are in this +# request's `tools` array") on an inband request, which carries no tools +# array, left the model with no usable tools at all. Swap the prompt's +# tool-use sections (Prompts.tool_sections — pure text, so an exact replace) +# to match the format this request is actually sent in. The history's system +# message is untouched; only the outbound copy changes. +fun retarget_system_prompt(messages, tool_format) { + if (length(messages) == 0) { messages } + else { + first = hd(messages) + if (map_get(first, 'role') != 'system') { messages } + else { + want = if (tool_format == 'native') { "native" } else { "inband" } + other = if (want == "native") { "inband" } else { "native" } + content = to_string(map_get(first, 'content')) + from = Prompts.tool_sections(other) + if (string_contains(content, from) == 'false') { messages } + else { + fixed = map_put(first, 'content', string_replace(content, from, Prompts.tool_sections(want))) + [fixed | tl(messages)] + } + } + } +} + # ------------------------------------------------------------ # Attachment injection — turn the last user message in `messages` # into a multimodal content list, with image_url blocks before the @@ -504,7 +597,7 @@ fun inject_context_status(messages, opts) { else { # Bracketed marker (not XML) — avoids any future # JSON-escape weirdness on `<`/`>` and reads cleaner - # to humans glancing at /tmp/swarm-code-last-body.json. + # to humans glancing at the SWARM_CODE_DEBUG body dump. new_content = to_string(content) ++ "\n\n[" ++ status ++ "]" new_msg = map_put(last_msg, 'content', new_content) replace_at(messages, last_idx, new_msg, 0, []) @@ -515,14 +608,9 @@ fun inject_context_status(messages, opts) { } fun build_status_string(messages, opts) { - max_tok = parse_env_int_local("SWARM_CODE_MAX_TOKENS", 262144) - out_res = parse_env_int_local("SWARM_CODE_OUTPUT_RESERVE", 16384) - buf = parse_env_int_local("SWARM_CODE_COMPACT_BUFFER", 52000) - # Clamp to a floor of 1 — a degenerate env (reserve + buffer >= max_tokens) - # would otherwise make this 0 or negative and the division below PANICS - # (swarmrt traps integer divide-by-zero) on the very first turn. - raw_budget = max_tok - out_res - buf - tok_budget = if (raw_budget < 1) { 1 } else { raw_budget } + # The same budget the compactor triggers on (always >= 1, so the + # division below can't trap on a degenerate env). + tok_budget = context_budget_tokens() # Use the server's real prompt-token count once we have it; fall back to a # char/4 estimate before the first response. @@ -543,6 +631,34 @@ fun build_status_string(messages, opts) { fmt_k(tok_used) ++ "/" ++ fmt_k(tok_budget) ++ " tok" } +# ------------------------------------------------------------ +# Context budget — the ONE definition (Agent's compaction trigger and the +# context meter above both use it; agent.sw imports llm.sw, not the reverse). +# ------------------------------------------------------------ +# budget = window − output reserve − compaction buffer. The reserve (16384, +# Kimi's max output) and buffer (52000) defaults are sized for a 262K window; +# subtracted from a small one they drove the budget NEGATIVE — at +# SWARM_CODE_MAX_TOKENS=32768 it was −35,616, so every step read "context at +# ~1197 / -35616 tokens, compacting" and paid a summarizer call. Each is now +# capped at a quarter of the window (explicit env values too), so the budget +# is always at least half the window: 8K→4096, 32K→16384, 128K→81920, +# 262K→193760 (unchanged). Floor of 1 for a degenerate window. +fun context_window_tokens() { parse_env_int_local("SWARM_CODE_MAX_TOKENS", 262144) } + +fun context_budget_tokens() { + budget_for_window(context_window_tokens(), + parse_env_int_local("SWARM_CODE_OUTPUT_RESERVE", 16384), + parse_env_int_local("SWARM_CODE_COMPACT_BUFFER", 52000)) +} + +fun budget_for_window(window, reserve_cfg, buffer_cfg) { + quarter = window / 4 + reserve = if (reserve_cfg < quarter) { reserve_cfg } else { quarter } + buffer = if (buffer_cfg < quarter) { buffer_cfg } else { quarter } + b = window - reserve - buffer + if (b < 1) { 1 } else { b } +} + # Estimate token count from char count (≈4 chars/token) for use on # the first turn before the server has reported usage. Multimodal # content lists are skipped (images aren't text-counted). @@ -594,28 +710,6 @@ fun parse_positive_int_local(s, i, acc, saw_digit) { } } -# inband mode fallback: pull the text block out of a multimodal -# content list. Images are dropped silently — inband protocol -# (Gemma 4 in-band tool calls) is text-only. -# Workaround for a swarmrt-side bug: http_post_stream's content -# accumulator doesn't decode \uXXXX JSON escapes for ASCII-range -# characters, so prose like `df -h / && free` comes through as -# `df -h / u0026u0026 free` (the backslash is gone, but the `uXXXX` -# stays). Until swarmrt's stream emitter is fixed, replace the most -# common offenders on the prose we extracted. Limited to chars that -# almost never appear as a literal `uXXXX` in real prose, so the -# false-positive risk is negligible. -fun fix_json_unicode_escapes(s) { - s1 = string_replace(s, "u0026", "&") - s2 = string_replace(s1, "u003c", "<") - s3 = string_replace(s2, "u003e", ">") - s4 = string_replace(s3, "u0027", "'") - s5 = string_replace(s4, "u0022", "\"") - s6 = string_replace(s5, "u002f", "/") - s7 = string_replace(s6, "u003d", "=") - string_replace(s7, "u005c", "\\") -} - # True when the only content the model produced is the C-side truncation # marker — a reasoning-only turn that hit max_tokens. Lets the recovery # surface the reasoning instead of showing a near-blank turn. @@ -629,6 +723,9 @@ fun is_truncation_marker_only(prose) { } } +# inband mode fallback: pull the text block out of a multimodal +# content list. Images are dropped silently — inband protocol +# (Gemma 4 in-band tool calls) is text-only. fun extract_text_block(content_list) { if (length(content_list) == 0) { "" } else { @@ -909,6 +1006,42 @@ fun last_fail(opts) { if (table == nil) { nil } else { ets_get(table, 'fail_reason') } } +# The server's message for the last FATAL rejection ("" when none) — lets the +# agent tell "the context is too long" (recoverable: compact and retry) from +# any other 4xx. +fun record_fail_detail(opts, msg) { + table = map_get(opts, 'llm_stats_table') + if (table == nil) { 'ok' } + else { ets_put(table, 'fail_detail', to_string(msg)) } +} + +fun last_fail_detail(opts) { + table = map_get(opts, 'llm_stats_table') + v = if (table == nil) { nil } else { ets_get(table, 'fail_detail') } + if (v == nil) { "" } else { v } +} + +# Did the last fatal rejection say the prompt exceeds the model's context? +fun last_fail_context_overflow(opts) { + if (last_fail(opts) != 'fatal') { 'false' } + else { is_context_overflow_msg(last_fail_detail(opts)) } +} + +# Wording used by OpenAI-compatible servers (OpenAI / vLLM "maximum context +# length", "context_length_exceeded"), Anthropic ("prompt is too long"), +# llama.cpp ("exceeds the available context size") and others. +fun is_context_overflow_msg(msg) { + low = string_lower(to_string(msg)) + if (string_contains(low, "maximum context length") == 'true') { 'true' } + else { if (string_contains(low, "context_length_exceeded") == 'true') { 'true' } + else { if (string_contains(low, "too many tokens") == 'true') { 'true' } + else { if (string_contains(low, "prompt is too long") == 'true') { 'true' } + else { if (string_contains(low, "exceeds the available context size") == 'true') { 'true' } + else { if (string_contains(low, "context window") == 'true' && + string_contains(low, "exceed") == 'true') { 'true' } + else { 'false' }}}}}} +} + fun retry_delay_ms(attempt) { base = if (attempt == 0) { 1000 } else { if (attempt == 1) { 2000 } @@ -1353,16 +1486,32 @@ fun use_routed_stream(opts) { else { 'true' }}}}}}} } +# Network-isolation gate at the point of dial. main.sw's startup check +# only sees the primary endpoint; the fallback, providers[] and +# .profile_override endpoints are chosen later, so every URL the LLM +# layer contacts is re-checked here (Config.endpoint_refusal). Returns +# nil when the URL may be dialed; otherwise prints the reason (diag() is +# silent in headless) and returns it. +fun dial_refusal(url) { + reason = Config.endpoint_refusal(url) + if (reason != nil) { eprint(" " ++ UI.warn_text("⚠ " ++ reason)) } + reason +} + fun stream_call(url, hdrs, body, opts) { routed_ok = use_routed_stream(opts) - if (map_get(opts, 'wake_turn') == 'true' && routed_ok == 'true') { + # A refused URL surfaces as a FATAL (4xx) failure: never retried; the + # fallback / next provider still gets its own (gated) attempt. + if (dial_refusal(url) != nil) { + {'error', 403, "blocked by network isolation (see above)"} + } else { if (map_get(opts, 'wake_turn') == 'true' && routed_ok == 'true') { wake_sync_stream(url, hdrs, body, opts) } else { if (routed_ok == 'true') { routed_stream(url, hdrs, body, opts) } else { record_stream_mode(opts, 'sync', "") http_post_stream(url, hdrs, body) - }} + }}} } # Which mode the JUST-COMPLETED stream used + what it painted. Read by @@ -1747,7 +1896,7 @@ fun chat_native(messages, opts) { body_chars = string_length(body) file_mkdir(getenv("HOME") ++ "/.swarm-code") - file_write(getenv("HOME") ++ "/.swarm-code/last-body.json", body) + dump_last_body(body) Log.llm_request(to_string(model), length(messages), body_chars) # Substantive live-wait line shown for the whole TTFT window. # Suppressed on wake turns — the Reader is pinned in read_line and @@ -1770,6 +1919,7 @@ fun chat_native(messages, opts) { else { list_append(base_hdrs, {"Authorization", "Bearer " ++ api_key}) } record_fail(opts, "fail") + record_fail_detail(opts, "") record_retry_after(opts, 0) cls = classify_stream(stream_call(url, hdrs, body, opts)) latency = timestamp() - start_ms @@ -1786,6 +1936,7 @@ fun chat_native(messages, opts) { f_status = elem(cls, 1) f_msg = elem(cls, 2) record_fail(opts, 'fatal') + record_fail_detail(opts, f_msg) diag(opts, " " ++ UI.err_text("✗ request rejected (HTTP " ++ to_string(f_status) ++ ") — not retrying: " ++ to_string(f_msg))) Log.llm_error("fatal request error " ++ to_string(f_status), to_string(f_msg)) @@ -1816,6 +1967,7 @@ fun chat_native(messages, opts) { # identical body just re-fails. Mark fatal so the retry # loop stops instead of hammering the same poisoned body. record_fail(opts, 'fatal') + record_fail_detail(opts, to_string(em) ++ " " ++ to_string(map_get(err, 'code'))) diag(opts, " " ++ UI.err_text("[llm error] " ++ to_string(em))) Log.llm_error("server error", to_string(em)) nil @@ -1835,8 +1987,11 @@ fun chat_native(messages, opts) { msg_obj = map_get(choice0, 'message') raw_content = map_get(msg_obj, 'content') raw_tool_calls = map_get(msg_obj, 'tool_calls') - prose_raw = if (raw_content == nil) { "" } - else { fix_json_unicode_escapes(to_string(raw_content)) } + # Content is used verbatim: the runtime already decodes \uXXXX + # escapes. (A u003c -> "<" "repair" pass left over from an old + # runtime bug turned "\u003c" in code into "\<" and mangled + # prose that mentions such escapes.) + prose_raw = if (raw_content == nil) { "" } else { to_string(raw_content) } # Length-truncation recovery: when the server stops on # max_tokens with an empty content but populated # reasoning_content (Kimi K2 thinking mid-stream), we'd @@ -1866,6 +2021,10 @@ fun chat_native(messages, opts) { else { if (string_contains(to_string(prose_raw), "[Response truncated at max_tokens") == 'true') { 'true' } else { 'false' }} + # ESC landed mid-stream (the runtime appends the marker; the + # sync path still hands back the tool calls streamed so far, + # arguments possibly cut mid-string). run_turn must not run them. + interrupted = stream_interrupted(prose_raw) had_tools = if (length(tool_calls) > 0) { 'true' } else { 'false' } Log.llm_response(latency, string_length(prose), had_tools) @@ -1887,7 +2046,8 @@ fun chat_native(messages, opts) { content: prose, tool_calls: tool_calls, reasoning: reason_text, - truncated: truncated + truncated: truncated, + interrupted: interrupted } } } @@ -1895,6 +2055,15 @@ fun chat_native(messages, opts) { } } +# The user stopped this stream (ESC/Ctrl-C): the runtime — and routed mode's +# interrupted_stream_result — append "[Request interrupted by user]" to the +# content. Checked on the RAW content: inband parsing cuts the prose at the +# first call marker, which would drop the marker along with it. +fun stream_interrupted(raw) { + if (string_ends_with(string_trim(to_string(raw)), "[Request interrupted by user]") == 'true') { 'true' } + else { 'false' } +} + # Convert OpenAI tool_calls → internal flat shape. # Wire shape: [%{id, type:"function", function:%{name, arguments: }}, ...] # Internal flat: [%{id, name, arguments: }, ...] @@ -1923,6 +2092,26 @@ fun api_tool_calls_to_internal(raw, acc) { }} } +# Debug-only copy of the outbound request body — it carries the whole +# conversation (prompts, tool output, any secrets in them), so it is +# written only with SWARM_CODE_DEBUG=1 and never world-readable: mkstemp +# (file_temp) creates the file 0600 and the rename keeps that mode. +# Nothing reads it back; it is for a human debugging a request. +fun dump_last_body(body) { + home = getenv("HOME") + if (getenv("SWARM_CODE_DEBUG") != "1" || home == nil) { 'skip' } + else { + dir = home ++ "/.swarm-code" + file_mkdir(dir) + tmp = file_temp(dir ++ "/.last-body.") + if (tmp == nil) { 'skip' } + else { + file_write(tmp, body) + file_rename(tmp, dir ++ "/last-body.json") + } + } +} + # ============================================================ # Inband streaming path — parse markers ONCE into structured form # ============================================================ @@ -1934,7 +2123,7 @@ fun chat_inband(messages, opts) { body = build_request_body(messages, opts) body_chars = string_length(body) - file_write("/tmp/swarm-code-last-body.json", body) + dump_last_body(body) Log.llm_request(to_string(model), length(messages), body_chars) # Same live-wait line for the inband (Gemma-style) path. Suppressed # on wake turns (Reader pinned in read_line); newline-less in @@ -1953,6 +2142,7 @@ fun chat_inband(messages, opts) { else { list_append(base_hdrs, {"Authorization", "Bearer " ++ api_key}) } record_fail(opts, "fail") + record_fail_detail(opts, "") record_retry_after(opts, 0) cls = classify_stream(stream_call(url, hdrs, body, opts)) latency = timestamp() - start_ms @@ -1965,6 +2155,7 @@ fun chat_inband(messages, opts) { # retry loop (chat_inband_retry checks last_fail), transient ones retry. if (cls_tag == 'fatal') { record_fail(opts, 'fatal') + record_fail_detail(opts, elem(cls, 2)) diag(opts, " " ++ UI.err_text("✗ request rejected (HTTP " ++ to_string(elem(cls, 1)) ++ ") — not retrying: " ++ to_string(elem(cls, 2)))) Log.llm_error("fatal request error (inband, HTTP " ++ to_string(elem(cls, 1)) ++ ")", to_string(elem(cls, 2))) @@ -1985,7 +2176,10 @@ fun chat_inband(messages, opts) { reason_text = extract_reasoning(resp) record_reasoning(opts, reason_text) - parsed = parse_inband_tool_calls(fix_json_unicode_escapes(to_string(raw_content))) + # Verbatim, like chat_native — the runtime already decoded \uXXXX, so + # a JSON escape the model put INSIDE call:NAME{...} arguments + # (\u003c for "<") is decoded once, by the arguments' own parse. + parsed = parse_inband_tool_calls(to_string(raw_content)) prose_raw = map_get(parsed, 'content') tool_calls = map_get(parsed, 'tool_calls') @@ -2029,10 +2223,13 @@ fun chat_inband(messages, opts) { # post_stream_render. post_stream_render(opts, to_string(prose), had_tools) - # F4: same length-truncation signal as chat_native. + # F4: same length-truncation signal as chat_native. Tested on the RAW + # content: when calls were parsed, prose_raw ends at the first call + # marker and the runtime's truncation/interrupt marker (appended at + # the very end) is no longer in it. fin = extract_finish_reason(resp) truncated = if (fin == "length") { 'true' } - else { if (string_contains(to_string(prose_raw), + else { if (string_contains(to_string(raw_content), "[Response truncated at max_tokens") == 'true') { 'true' } else { 'false' }} @@ -2040,7 +2237,8 @@ fun chat_inband(messages, opts) { content: prose, tool_calls: tool_calls, reasoning: reason_text, - truncated: truncated + truncated: truncated, + interrupted: stream_interrupted(raw_content) } } } @@ -2151,7 +2349,7 @@ fun chat_silent(messages, opts) { hdrs = if (api_key == nil) { base_hdrs } else { list_append(base_hdrs, {"Authorization", "Bearer " ++ api_key}) } - resp = http_post(url, hdrs, body) + resp = if (dial_refusal(url) != nil) { nil } else { http_post(url, hdrs, body) } if (resp == nil) { nil } else { usage = extract_usage(resp) @@ -2200,7 +2398,8 @@ fun chat_for_subagent(messages, opts, target_pid, name) { hdrs = if (api_key == nil) { base_hdrs } else { list_append(base_hdrs, {"Authorization", "Bearer " ++ api_key}) } - resp = unwrap_stream(http_post_stream(url, hdrs, body, target_pid, name)) + resp = if (dial_refusal(url) != nil) { nil } + else { unwrap_stream(http_post_stream(url, hdrs, body, target_pid, name)) } latency = timestamp() - start_ms if (resp == nil) { @@ -2223,8 +2422,7 @@ fun chat_for_subagent(messages, opts, target_pid, name) { choice0 = hd(choices) msg_obj = map_get(choice0, 'message') raw_content = map_get(msg_obj, 'content') - prose_raw = if (raw_content == nil) { "" } - else { fix_json_unicode_escapes(to_string(raw_content)) } + prose_raw = if (raw_content == nil) { "" } else { to_string(raw_content) } tool_format = map_get(opts, 'tool_format') result_map = if (tool_format == 'native') { %{ diff --git a/src/log.sw b/src/log.sw index 5c52bc1..4cbcad1 100644 --- a/src/log.sw +++ b/src/log.sw @@ -36,7 +36,7 @@ export [ tool_call, tool_result, bg_done, bg_stalled, compaction, permission, tail_recent, summarize, - redact + redact, redact_value ] # Ensure the telemetry directory exists and return the log path. @@ -55,12 +55,14 @@ fun path() { } # Core writer: serialize a map to JSON + write a single line. -# The map should already contain a 'type' key. The encoded line passes -# through redact() so secrets in previews/args never reach disk — this -# single funnel covers every event constructor below. +# The map should already contain a 'type' key. Every string VALUE is +# redacted BEFORE encoding (redact_value) so secrets in previews/args +# never reach disk — this single funnel covers every event constructor +# below. Redacting the ENCODED line instead corrupted it: a masked run +# could swallow the letter of a \n escape, leaving an invalid `\[`. fun event(data) { with_ts = map_put(data, 'ts', timestamp()) - line = redact(json_encode(with_ts)) ++ "\n" + line = json_encode(redact_value(with_ts)) ++ "\n" file_append(path(), line) } @@ -229,16 +231,26 @@ fun truncate(s, max_len) { # Secret redaction # ------------------------------------------------------------ # -# redact(s) masks common secret shapes in a string before it is -# written to disk. Applied to every events.jsonl line (see event()) -# and to trajectory exports (Trajectory module). Layered, cheapest -# and most-precise first: +# redact(s) masks common secret shapes in a PLAIN-TEXT string before it +# is written to disk; redact_value(v) applies it to every string inside +# a decoded value (maps / lists walked recursively, keys kept) and is +# what callers use BEFORE json_encode — events.jsonl (event()) and +# trajectory exports (Trajectory module). Never run redact over encoded +# JSON: escapes (\n, \") would be masked into invalid JSON. +# Layered, cheapest and most-precise first: # 1. exact match on live SWARM_CODE_API_KEY / SWARM_CODE_EMBED_KEY -# 2. known token prefixes (sk-, mk_live_, AKIA, ghp_, xoxb-, ...) -# 3. "Bearer " authorization headers -# 4. key/value shapes (_key":"..., password":"..., api_key=..., -# + JSON-escaped forms) -# 5. long blobs: >= 40 contiguous [A-Za-z0-9+/=_-] mixing letters+digits +# 2. PEM blocks (-----BEGIN …----- … -----END …-----): body masked +# wholesale — base64 lines with '/' dodged the blob heuristic +# 3. URL userinfo: scheme://user:pass@host → scheme://user:[REDACTED]@host +# 4. known token prefixes (sk-, mk_live_, AKIA, ghp_, xoxb-, ...) +# 5. "Bearer " authorization headers +# 6. key/value shapes, case-insensitive, for identifiers ENDING in a +# secret word (password, passwd, secret, token, api_key, apikey, +# access_key, private_key, _key, …: PGPASSWORD=, AWS_SECRET_ACCESS_KEY=, +# "password": "…", password: … (YAML), \"token\":\"…\" (JSON in text)) +# with optional quotes / whitespace around = : or => +# 7. long blobs: >= 40 contiguous [A-Za-z0-9+/=_-] mixing letters+digits; +# a backslash and the character it escapes end a run # sw has no regex, so these are recursive string_index_of/string_sub # scans (tail calls, flat stack). Stateless; thresholds are conservative # so file paths and short hashes stay readable. @@ -246,24 +258,36 @@ fun truncate(s, max_len) { fun redact(s) { if (s == nil) { "" } else { - s1 = redact_exact(s, getenv("SWARM_CODE_API_KEY")) + s1 = redact_exact(to_string(s), getenv("SWARM_CODE_API_KEY")) s2 = redact_exact(s1, getenv("SWARM_CODE_EMBED_KEY")) - s3 = redact_prefixes(s2, [ + s3 = redact_pem(s2, 0) + s4 = redact_userinfo(s3, 0) + s5 = redact_prefixes(s4, [ "sk-", "mk_live_", "mk_test_", "AKIA", "ghp_", "gho_", "github_pat_", "xoxb-", "xoxp-" ]) - s4 = redact_bearers(s3, ["Bearer ", "bearer "]) - s5 = redact_kvs(s4, [ - "_key\":\"", "token\":\"", "secret\":\"", "password\":\"", - "authorization\":\"", "Authorization\":\"", - "_key\\\":\\\"", "token\\\":\\\"", "secret\\\":\\\"", "password\\\":\\\"", - "authorization\\\":\\\"", "Authorization\\\":\\\"", - "api_key=" - ]) - redact_blobs(s5, string_length(s5), 0, 0, 'false', 'false') + s6 = redact_bearers(s5, ["Bearer ", "bearer "]) + s7 = redact_kvs(s6, rd_secret_words()) + redact_blobs(s7, string_length(s7), 0, 0, 'false', 'false') } } +# Redact every string inside a decoded JSON-shaped value. Map keys are +# structure, not data, and are kept verbatim. +fun redact_value(v) { + if (v == nil) { v } + else { if (typeof(v) == "string") { redact(v) } + else { if (is_list(v) == 'true') { map(fn(x) { redact_value(x) }, v) } + else { if (is_map(v) == 'true') { + rd_map_loop(map_keys(v), map_values(v), map_new()) + } else { v } } } } +} + +fun rd_map_loop(keys, vals, acc) { + if (length(keys) == 0) { acc } + else { rd_map_loop(tl(keys), tl(vals), map_put(acc, hd(keys), redact_value(hd(vals)))) } +} + # Layer 1: exact-match a live key value (zero false positives). fun redact_exact(s, k) { if (k == nil) { s } @@ -273,7 +297,81 @@ fun redact_exact(s, k) { } } -# Layer 2: known token prefixes. Keeps the prefix visible (so logs +# Layer 2: PEM blocks. The BEGIN/END lines stay (they say WHAT was +# there); everything between is masked. A block with no END (a +# truncated preview) is masked to the end of the string. +fun redact_pem(s, from) { + slen = string_length(s) + idx = rd_index_from(s, "-----BEGIN ", from) + if (idx < 0) { s } + else { + close = rd_index_from(s, "-----", idx + 11) + body_start = if (close < 0) { slen } else { close + 5 } + end_idx = rd_index_from(s, "-----END ", body_start) + if (end_idx < 0) { + string_sub(s, 0, body_start) ++ "\n[REDACTED]" + } else { + ns = string_sub(s, 0, body_start) ++ "\n[REDACTED]\n" ++ + string_sub(s, end_idx, slen - end_idx) + redact_pem(ns, body_start + 12 + 9) + } + } +} + +# Layer 3: URL userinfo. In scheme://user:pass@host the password is +# masked (the user stays: it is rarely secret and useful context). +fun redact_userinfo(s, from) { + slen = string_length(s) + idx = rd_index_from(s, "://", from) + if (idx < 0) { s } + else { + auth_start = idx + 3 + auth_end = rd_authority_end(s, auth_start, slen) + at = rd_last_at(s, auth_start, auth_end, 0 - 1) + colon = if (at < 0) { 0 - 1 } else { rd_index_in(s, 58, auth_start, at) } + if (colon < 0 || at - colon - 1 < 1 || + string_sub(s, colon + 1, at - colon - 1) == "[REDACTED]") { + redact_userinfo(s, auth_start) + } else { + ns = string_sub(s, 0, colon + 1) ++ "[REDACTED]" ++ string_sub(s, at, slen - at) + redact_userinfo(ns, colon + 11) + } + } +} + +# End of a URL authority: first / ? # whitespace quote < > \ ) ] or end. +fun rd_authority_end(s, i, slen) { + if (i >= slen) { i } + else { + c = codepoint_at(s, i) + if (c == 47 || c == 63 || c == 35 || c == 32 || c == 9 || c == 10 || c == 13 || + c == 34 || c == 39 || c == 60 || c == 62 || c == 92 || c == 41 || c == 93) { i } + else { rd_authority_end(s, i + 1, slen) } + } +} + +fun rd_last_at(s, i, stop, found) { + if (i >= stop) { found } + else { rd_last_at(s, i + 1, stop, (if (codepoint_at(s, i) == 64) { i } else { found })) } +} + +# First index of byte `c` in [i, stop), or -1. +fun rd_index_in(s, c, i, stop) { + if (i >= stop) { 0 - 1 } + else { if (codepoint_at(s, i) == c) { i } else { rd_index_in(s, c, i + 1, stop) } } +} + +# string_index_of with a start offset (absolute result, -1 if absent). +fun rd_index_from(s, needle, from) { + slen = string_length(s) + if (from >= slen) { 0 - 1 } + else { + idx = string_index_of(string_sub(s, from, slen - from), needle) + if (idx < 0) { 0 - 1 } else { from + idx } + } +} + +# Layer 4: known token prefixes. Keeps the prefix visible (so logs # still show what KIND of key was masked), masks the token body when # prefix + body is >= 16 chars. fun redact_prefixes(s, prefixes) { @@ -312,7 +410,7 @@ fun redact_prefix_from(s, prefix, from) { } } -# Layer 3: "Bearer " — mask the token after the marker when it +# Layer 5: "Bearer " — mask the token after the marker when it # is >= 12 chars. fun redact_bearers(s, markers) { if (length(markers) == 0) { s } @@ -340,39 +438,148 @@ fun redact_bearer_from(s, marker, from) { } } -# Layer 4: key/value shapes. Covers bare field names (password":", -# token":", secret":" — token/secret also catch the _token/_secret -# suffixed forms), the _key suffix (api_key etc.; bare key":" would -# false-positive on words like monkey), and the backslash-escaped -# forms (_key\":\") that appear once tool args are embedded -# inside a JSON-encoded line. Values >= 8 chars are masked up to the -# next quote/backslash/whitespace/delimiter. -fun redact_kvs(s, markers) { - if (length(markers) == 0) { s } - else { redact_kvs(redact_kv_from(s, hd(markers), 0), tl(markers)) } +# Layer 6: key/value secrets. Each entry is {word, min_value_len}: an +# identifier (case-insensitive) ENDING in `word` — PGPASSWORD, +# db_password, AWS_SECRET_ACCESS_KEY, X-Api-Key — followed by an +# optional closing quote (" ' or \"), blanks, a separator (= : =>, but +# not == or ::), blanks and an optional opening quote. Words that are +# unambiguous secrets mask any non-trivial value; the generic suffixes +# (token, _key) keep the old >= 8 threshold so `max_token: 4096` or +# `sort_key: id` stay readable. +fun rd_secret_words() { + [{"password", 1}, {"passwd", 1}, {"passphrase", 1}, {"secret", 1}, + {"api_key", 1}, {"apikey", 1}, {"api-key", 1}, {"access_key", 1}, + {"private_key", 1}, {"authorization", 1}, + {"token", 8}, {"_key", 8}, {"-key", 8}, {"credential", 8}, {"credentials", 8}] } -fun redact_kv_from(s, marker, from) { +fun redact_kvs(s, words) { + if (length(words) == 0) { s } + else { + w = hd(words) + redact_kvs(redact_kv_from(s, string_lower(s), elem(w, 0), elem(w, 1), 0), tl(words)) + } +} + +# `low` is string_lower(s): same byte length (ASCII-only lowering), so +# offsets found in it apply to `s`. +fun redact_kv_from(s, low, word, min_len, from) { slen = string_length(s) - if (from >= slen) { s } + idx = rd_index_from(low, word, from) + if (idx < 0) { s } else { - idx = string_index_of(string_sub(s, from, slen - from), marker) - if (idx < 0) { s } + wend = idx + string_length(word) + ident_ends = if (wend >= slen) { 'true' } + else { if (rd_is_token(codepoint_at(s, wend)) == 'true') { 'false' } else { 'true' } } + span = if (ident_ends == 'true') { rd_kv_span(s, wend, slen, min_len <= 1) } else { {0 - 1, 0 - 1} } + vs = elem(span, 0) + ve = elem(span, 1) + if (vs < 0) { redact_kv_from(s, low, word, min_len, wend) } + else { if (rd_kv_maskable(s, vs, ve, min_len) == 'false') { + redact_kv_from(s, low, word, min_len, (if (ve > wend) { ve } else { wend })) + } else { + ns = string_sub(s, 0, vs) ++ "[REDACTED]" ++ string_sub(s, ve, slen - ve) + redact_kv_from(ns, string_lower(ns), word, min_len, vs + 10) + }} + } +} + +# Value span {start, end} after a key ending at i, or {-1, -1} when no +# key/value separator follows. Quoted values run to the matching quote; +# unquoted `=` values (env / query / ini) stop at whitespace and +# delimiters; unquoted `:` values (YAML / headers) run to end of line +# for the unambiguous secret words (`to_eol`: "password: two words") +# but stop at whitespace for the generic suffixes, so prose such as +# "max_token: 4096 sort_key: id" keeps its short values. +fun rd_kv_span(s, i, slen, to_eol) { + j0 = rd_skip_key_quote(s, i, slen) + j1 = rd_skip_blanks(s, j0, slen) + if (j1 >= slen) { {0 - 1, 0 - 1} } + else { + c = codepoint_at(s, j1) + n1 = if (j1 + 1 < slen) { codepoint_at(s, j1 + 1) } else { 0 } + sep_end = if (c == 58 && n1 != 58) { j1 + 1 } # : (not ::) + else { if (c == 61 && n1 == 62) { j1 + 2 } # => + else { if (c == 61 && n1 != 61) { j1 + 1 } # = (not ==) + else { 0 - 1 } } } + if (sep_end < 0) { {0 - 1, 0 - 1} } else { - val_start = from + idx + string_length(marker) - val_end = rd_value_end(s, val_start, slen) - if (val_end - val_start >= 8) { - ns = string_sub(s, 0, val_start) ++ "[REDACTED]" ++ - string_sub(s, val_end, slen - val_end) - redact_kv_from(ns, marker, val_start + 10) + j = rd_skip_blanks(s, sep_end, slen) + q = if (j < slen) { codepoint_at(s, j) } else { 0 } + q2 = if (j + 1 < slen) { codepoint_at(s, j + 1) } else { 0 } + if (q == 34 || q == 39) { + {j + 1, rd_until_quote(s, j + 1, slen, q)} + } else { if (q == 92 && q2 == 34) { + {j + 2, rd_until_quote(s, j + 2, slen, 92)} + } else { if (c == 58 && to_eol == 'true') { + {j, rd_trim_end(s, j, rd_until_eol(s, j, slen))} } else { - redact_kv_from(s, marker, val_start) - } + {j, rd_value_end(s, j, slen)} + }}} } } } -# Layer 5: long-blob heuristic. Any contiguous run of base64-ish chars +fun rd_skip_key_quote(s, i, slen) { + if (i >= slen) { i } + else { + c = codepoint_at(s, i) + if (c == 34 || c == 39) { i + 1 } + else { if (c == 92 && i + 1 < slen && codepoint_at(s, i + 1) == 34) { i + 2 } + else { i } } + } +} + +fun rd_skip_blanks(s, i, slen) { + if (i >= slen) { i } + else { + c = codepoint_at(s, i) + if (c == 32 || c == 9) { rd_skip_blanks(s, i + 1, slen) } else { i } + } +} + +# Up to (not including) quote byte q, a backslash, or a newline. +fun rd_until_quote(s, i, slen, q) { + if (i >= slen) { i } + else { + c = codepoint_at(s, i) + if (c == q || c == 92 || c == 10 || c == 13) { i } else { rd_until_quote(s, i + 1, slen, q) } + } +} + +# Up to end of line, or a flow-style delimiter , } ] or quote. +fun rd_until_eol(s, i, slen) { + if (i >= slen) { i } + else { + c = codepoint_at(s, i) + if (c == 10 || c == 13 || c == 44 || c == 125 || c == 93 || c == 34 || c == 39) { i } + else { rd_until_eol(s, i + 1, slen) } + } +} + +fun rd_trim_end(s, start, e) { + if (e <= start) { e } + else { + c = codepoint_at(s, e - 1) + if (c == 32 || c == 9) { rd_trim_end(s, start, e - 1) } else { e } + } +} + +# Worth masking: long enough for its class, not already masked, and not +# a bare boolean/null ("secret: true" is configuration, not a secret). +fun rd_kv_maskable(s, vs, ve, min_len) { + n = ve - vs + if (n < min_len || n < 1) { 'false' } + else { + v = string_lower(string_sub(s, vs, n)) + # "[redacted" — the span may stop at the marker's own "]". + if (string_starts_with(v, "[redacted") == 'true') { 'false' } + else { if (v == "true" || v == "false" || v == "null" || v == "none" || v == "nil") { 'false' } + else { 'true' } } + } +} + +# Layer 7: long-blob heuristic. Any contiguous run of base64-ish chars # [A-Za-z0-9+/=_-] that is >= 40 long AND mixes letters with digits is # masked. Threshold 40 keeps file paths and 7-char short hashes # readable; "[REDACTED]" (letters only) can never re-match itself. @@ -380,6 +587,9 @@ fun redact_kv_from(s, marker, from) { # exactly 40/64 chars (full git SHA-1/SHA-256) and '/'-bearing runs # under 80 chars (digit-containing file paths — '/' stays in the # charset because base64 secrets contain it, but those run long). +# A backslash ends a run AND the byte it escapes is skipped: in text +# that embeds JSON / C strings, "\nAKIA…" must not become a run +# starting with the escape letter (masking it yields "\[REDACTED]"). fun redact_blobs(s, slen, i, start, seen_alpha, seen_digit) { if (i >= slen) { if (rd_blob_hit(s, start, slen, seen_alpha, seen_digit) == 'true') { @@ -392,12 +602,14 @@ fun redact_blobs(s, slen, i, start, seen_alpha, seen_digit) { nd = if (rd_is_digit(c) == 'true') { 'true' } else { seen_digit } redact_blobs(s, slen, i + 1, start, na, nd) } else { + skip = if (c == 92) { 2 } else { 1 } if (rd_blob_hit(s, start, i, seen_alpha, seen_digit) == 'true') { ns = string_sub(s, 0, start) ++ "[REDACTED]" ++ string_sub(s, i, slen - i) - redact_blobs(ns, string_length(ns), start + 10, start + 10, 'false', 'false') + nstart = start + 10 + skip + redact_blobs(ns, string_length(ns), nstart, nstart, 'false', 'false') } else { - redact_blobs(s, slen, i + 1, i + 1, 'false', 'false') + redact_blobs(s, slen, i + skip, i + skip, 'false', 'false') } } } @@ -448,7 +660,7 @@ fun rd_token_end(s, i, slen) { } # End (exclusive) of a secret value: stops at whitespace, quotes, -# backslash (start of a JSON escape), and ,/}/]/& delimiters. +# backslash (start of a JSON escape), and , } ] & ; ) delimiters. fun rd_value_end(s, i, slen) { if (i >= slen) { i } else { @@ -471,8 +683,8 @@ fun rd_is_alpha(c) { (c >= 65 && c <= 90) || (c >= 97 && c <= 122) } fun rd_is_digit(c) { c >= 48 && c <= 57 } -# space tab nl cr " ' \ , } ] & +# space tab nl cr " ' \ , } ] & ; ) fun rd_is_stop(c) { c == 32 || c == 9 || c == 10 || c == 13 || c == 34 || c == 39 || - c == 92 || c == 44 || c == 125 || c == 93 || c == 38 + c == 92 || c == 44 || c == 125 || c == 93 || c == 38 || c == 59 || c == 41 } diff --git a/src/main.sw b/src/main.sw index 85f1a82..50b3483 100644 --- a/src/main.sw +++ b/src/main.sw @@ -18,6 +18,8 @@ module Main # SWARM_CODE_API_KEY default: (none) # SWARM_CODE_MAX_TOKENS default: 262144 (Kimi K2.7 context window) # SWARM_CODE_OUTPUT_RESERVE default: 16384 (Kimi K2.7 max output) +# SWARM_CODE_COMPACT_BUFFER default: 52000 — reserve and buffer are each +# capped at 1/4 of SWARM_CODE_MAX_TOKENS # SWARM_CODE_TEMP default: "0.2" (string — parsed to float) # SWARM_CODE_CWD default: "." (used in system prompt) # @@ -54,6 +56,9 @@ fun main() { # server exposing bash/read/write/edit/glob/grep/web_fetch tools. # No LLM, no agent loop — pure tool execution for orchestrators. if (has_flag(os_args(), "--mcp-server") == 'true') { + # stdout is JSON-RPC framing — the notice goes to stderr. + mcp_pn = Config.project_notice() + if (mcp_pn != nil) { eprint("swarm-code: " ++ mcp_pn) } mcp_server_opts = %{ cwd: resolve_cwd(), settings: Config.load(), @@ -69,8 +74,9 @@ fun main() { # Headless mode: `swarm -p ""` runs one task to completion # and exits — no TUI, no Reader. `--json` adds a final result line. - # The prompt comes from the argument after -p, OR from stdin when - # no argument follows (or `-p -`) — so an orchestrator can pipe a + # The prompt comes from the argument after -p (or, when a flag follows + # -p, the first positional argument: `-p --json "..."`), OR from stdin + # when there is none (or `-p -`) — so an orchestrator can pipe a # multi-line prompt without shell-quoting it as an argument. cli_args = os_args() p_present = if (has_flag(cli_args, "-p") == 'true') { 'true' } @@ -84,6 +90,15 @@ fun main() { no_resume = if (has_flag(cli_args, "--no-resume") == 'true') { 'true' } else { 'false' } arg_prompt = get_print_arg(cli_args) headless = p_present + # Headless stdout carries only the result: the --json line, or the final + # answer when stdout is piped. The transcript (tool calls, streamed + # tokens, notices) moves to stderr, so `swarm -p ... | jq` and + # `x=$(swarm -p ...)` get clean output. A terminal stdout keeps the live + # transcript as before (-1 = not diverted). + result_fd = if (headless == 'true' && + (json_mode == 'true' || stdout_is_tty() == 'false')) { + stdout_to_stderr() + } else { -1 } headless_prompt = if (p_present == 'false') { nil } else { if (arg_prompt != nil) { arg_prompt } else { read_all_stdin("") }} @@ -99,10 +114,9 @@ fun main() { # Network isolation: verify the configured endpoint is on a local / # private / Tailscale network. Refuse to run against a public-internet - # endpoint unless SWARM_CODE_ALLOW_REMOTE=1 is set OR the user has - # supplied an explicit API key (which is itself an intentional opt-in - # to a remote provider). This guarantees no conversation data leaves - # the user's network by accident. + # endpoint unless SWARM_CODE_ALLOW_REMOTE=1 is set (an API key alone is + # NOT an opt-in). This guarantees no conversation data leaves the + # user's network by accident. llm.sw re-checks every URL it dials. endpoint_url = to_string(map_get(base_opts, 'endpoint')) # Optional full-terminal alt-screen mode (interactive only) @@ -123,7 +137,12 @@ fun main() { # in_alt tells it to LEAVE the alt screen first so the explanation # isn't discarded with the alt buffer (silent-exit bug). in_alt = if (tui_env == "1" && headless == 'false') { 'true' } else { 'false' } - verify_network_isolation(endpoint_url, map_get(base_opts, 'api_key'), in_alt) + verify_network_isolation(endpoint_url, in_alt) + + # An untrusted ./.swarm-code.json had keys dropped (Config.project_scope) + # — say so once, so a repo's config never half-applies silently. + project_note = Config.project_notice() + if (project_note != nil) { print(" " ++ UI.warn_text("⚠ " ++ project_note)) } # Load settings (user-global + project-local merged) and project context. settings = Config.load() @@ -189,7 +208,10 @@ fun main() { context_env = getenv("SWARM_CODE_EXECUTION_CONTEXT") execution_context = if (context_env == nil) { "main" } else { to_string(context_env) } - opts0 = map_put(base_opts, 'execution_context', execution_context) + # session_id: stamps /profile and /model overrides so THIS session's beat + # env vars while one left over from an earlier session doesn't (LLM.apply_override). + opts0 = map_put(map_put(base_opts, 'execution_context', execution_context), + 'session_id', uuid()) opts = map_put(opts0, 'cwd', cwd) opts2 = map_put(opts, 'todos_table', todos_table) opts3 = map_put(opts2, 'perms_table', perms_table) @@ -273,7 +295,7 @@ fun main() { } if (headless == 'true') { - Agent.run_headless(map_put(opts5, 'no_resume', no_resume), + Agent.run_headless(map_put(map_put(opts5, 'no_resume', no_resume), 'result_fd', result_fd), system_prompt_text, headless_prompt, json_mode) } else { Agent.run(opts5, system_prompt_text) @@ -332,29 +354,74 @@ fun has_flag(args, flag) { else { has_flag(tl(args), flag) }} } -# Value of the -p / --print flag — the inline headless prompt — or -# nil. nil means either the flag is absent OR it's present with no -# inline prompt (a flag-looking next arg, or nothing), in which case -# the caller reads the prompt from stdin instead. +# Value of the -p / --print flag — the inline headless prompt — or nil +# (absent, or read the prompt from stdin): +# -p "prompt" that argument — even when it starts with "-" +# ("-- summarize the diff", "-x ..."): Scheduler and +# Flows pass user prompts right after -p +# -p - stdin +# -p --json "prompt" a KNOWN flag right after -p: the first positional +# argument after it is the prompt (README's form) +# -p / -p --json no positional argument: stdin +# Only an exact "-" or a known flag is "not the prompt"; any other leading +# "-" used to switch to stdin, so `-p --json "..."` sent nothing and exited 1 +# and a prompt starting with "-" was silently dropped. fun get_print_arg(args) { - if (length(args) == 0) { nil } + idx = print_arg_index(args) + if (idx < 0) { nil } else { list_nth(args, idx) } +} + +# argv index of the inline -p prompt, or -1. +fun print_arg_index(args) { pai_find(args, 0) } + +fun pai_find(args, i) { + if (length(args) == 0) { 0 - 1 } else { h = hd(args) - if (h == "-p" || h == "--print") { - rest = tl(args) - if (length(rest) == 0) { nil } - else { - cand = hd(rest) - # A flag-looking next arg (`--json`, `-`) is not the - # prompt — that's stdin mode. - if (string_starts_with(cand, "-") == 'true') { nil } else { cand } - } - } else { - get_print_arg(tl(args)) - } + if (h == "-p" || h == "--print") { pai_after(tl(args), i + 1) } + else { pai_find(tl(args), i + 1) } + } +} + +# args: what follows -p; i: argv index of hd(args). +fun pai_after(args, i) { + if (length(args) == 0) { 0 - 1 } + else { + cand = hd(args) + if (cand == "-") { 0 - 1 } + else { if (is_cli_flag(cand) == 'true') { first_positional(args, i) } + else { i } } + } +} + +fun first_positional(args, i) { + if (length(args) == 0) { 0 - 1 } + else { + a = hd(args) + if (a == "-") { 0 - 1 } + else { if (is_cli_flag(a) == 'true') { + if (flag_takes_value(a) == 'true' && length(tl(args)) > 0) { first_positional(tl(tl(args)), i + 2) } + else { first_positional(tl(args), i + 1) } + } else { i } } } } +# Every flag main() understands — never mistaken for a prompt or a profile. +fun is_cli_flag(a) { + if (a == "--json" || a == "--no-resume" || a == "-p" || a == "--print") { 'true' } + else { if (a == "--profile" || a == "-P" || a == "--mcp-server" || a == "--print-config") { 'true' } + else { if (a == "--help" || a == "-h" || a == "--version" || a == "-V" || a == "--doctor") { 'true' } + else { 'false' } } } +} + +fun flag_takes_value(a) { + if (a == "--profile" || a == "-P") { 'true' } else { 'false' } +} + +fun list_nth(lst, i) { + if (i <= 0) { hd(lst) } else { list_nth(tl(lst), i - 1) } +} + # Slurp all of stdin into one string (newline-joined). Used for the # headless prompt when `-p` has no inline argument. read_line returns # nil at EOF (pipe closed / Ctrl-D), which ends the loop. @@ -385,7 +452,55 @@ fun handle_cli_flags(args) { has_flag(args, "--doctor") == 'true') { sys_exit(run_doctor()) } - else { "ok" }}}} + else { if (subcommand(args) == "trust" || subcommand(args) == "untrust") { + sys_exit(run_trust(subcommand(args), tl(tl(args)))) + } + else { "ok" }}}}} +} + +# First argument after the program name, or nil. +fun subcommand(args) { + if (length(args) < 2) { nil } else { list_nth(args, 1) } +} + +# swarm-code trust [DIR] | untrust [DIR] | trust --list +# A repo's ./.swarm-code.json only applies its safe keys (model, +# max_tokens, …) until its directory is listed in "trusted_projects" in +# ~/.swarm-code/settings.json. This edits that list for the user. +fun run_trust(cmd, rest) { + if (cmd == "trust" && length(rest) > 0 && hd(rest) == "--list") { + tp = Config.trusted_list() + if (length(tp) == 0) { print("no trusted projects") } else { print_lines(tp) } + 0 + } else { + dir = if (length(rest) > 0) { hd(rest) } else { "." } + add = if (cmd == "trust") { 'true' } else { 'false' } + r = Config.set_trusted(dir, add) + if (elem(r, 0) != 'ok') { + print("swarm-code " ++ cmd ++ ": " ++ to_string(elem(r, 1))) + 1 + } else { + abs = elem(r, 1) + changed = elem(r, 2) + if (add == 'true') { + if (changed == 'true') { print("trusted " ++ abs) } else { print("already trusted: " ++ abs) } + gated = Config.project_gated_keys(abs) + if (gated == nil) { + print(" (no .swarm-code.json there yet — one added later applies in full)") + } else { if (length(gated) > 0) { + print(" its .swarm-code.json now also applies: " ++ Config.join_names(gated, "") ++ + " — hooks and MCP servers run commands, so only trust repos you control") + } else { "" } } + } else { + if (changed == 'true') { print("untrusted " ++ abs) } else { print("not trusted: " ++ abs) } + } + 0 + } + } +} + +fun print_lines(xs) { + if (length(xs) == 0) { 'ok' } else { print(to_string(hd(xs))) ; print_lines(tl(xs)) } } fun print_usage() { @@ -398,7 +513,10 @@ fun print_usage() { print(" swarm -p \"\" run one task headless, then exit") print(" swarm -p \"...\" --json headless + a final JSON result line") print(" swarm -p \"...\" --no-resume start fresh, ignore .active session") + print(" swarm -p - < prompt.txt read the headless prompt from stdin") print(" swarm doctor validate config, endpoint, dirs, version") + print(" swarm trust [DIR] let DIR's .swarm-code.json apply in full (hooks, endpoint, MCP);") + print(" untrust [DIR] reverses it, trust --list shows the list") print(" swarm --mcp-server start as a stdio MCP tool server (JSON-RPC 2.0)") print(" swarm --help, -h show this help and exit") print(" swarm --version, -V print the version and exit") @@ -414,6 +532,7 @@ fun print_usage() { print(" SWARM_CODE_API_KEY API key for a remote provider") print(" SWARM_CODE_TOOL_FORMAT native | inband (else auto-detected)") print(" SWARM_CODE_ALLOW_REMOTE set to 1 to permit non-local endpoints") + print(" SWARM_CODE_HEADLESS_APPROVE set to 1 to auto-approve 'ask' tools in -p runs") print(" SWARM_CODE_CWD working directory shown to the model") print(" SWARM_CODE_PLAN=auto|on|off plan mode (default: auto)") print(" SWARM_CODE_EMBED_ENDPOINT embedding API URL (enables semantic recall)") @@ -645,30 +764,29 @@ fun resolve_profile_name(args, profiles) { else { if (profiles == nil) { nil } else { rest = if (length(args) == 0) { args } else { tl(args) } - find_positional_profile(rest, profiles) + find_positional_profile(rest, 1, print_arg_index(args), profiles) }} } -fun find_positional_profile(args, profiles) { +# i: argv index of hd(args). The -p prompt (prompt_idx, the SAME argument +# get_print_arg returns — `-p --json gemma` means prompt "gemma") and the +# values of --profile / -P are skipped; anything else starting with "-" is a +# flag. +fun find_positional_profile(args, i, prompt_idx, profiles) { if (length(args) == 0) { nil } else { a = hd(args) - if (string_starts_with(a, "-") == 'true') { - rest = tl(args) - value_taking = if (a == "-p") { 'true' } - else { if (a == "--print") { 'true' } - else { if (a == "--profile") { 'true' } - else { if (a == "-P") { 'true' } - else { if (a == "--mcp-server") { 'false' } - else { 'false' }}}}} - if (value_taking == 'true' && length(rest) > 0) { - find_positional_profile(tl(rest), profiles) + rest = tl(args) + if (i == prompt_idx) { find_positional_profile(rest, i + 1, prompt_idx, profiles) } + else { if (string_starts_with(a, "-") == 'true') { + if (flag_takes_value(a) == 'true' && length(rest) > 0) { + find_positional_profile(tl(rest), i + 2, prompt_idx, profiles) } else { - find_positional_profile(rest, profiles) + find_positional_profile(rest, i + 1, prompt_idx, profiles) } } else { if (lookup_string_key(profiles, a) != nil) { a } else { nil } - } + }} } } @@ -756,12 +874,12 @@ fun resolve_cwd() { " && pwd || echo " ++ Util.shell_q(to_string(cwd_env))), 1) string_trim(to_string(raw)) } else { - pwd_env = getenv("PWD") - if (pwd_env != nil) { - to_string(pwd_env) - } else { - string_trim(elem(shell("pwd"), 1)) - } + # Ask /bin/sh rather than trusting $PWD: a launcher that chdir()s + # without updating PWD (IDE integrations, subprocess(cwd=...), cron) + # leaves it naming another directory, and the system prompt would + # send the model's absolute paths there. sh keeps $PWD only when it + # really is the current directory, so symlinked paths survive. + string_trim(elem(shell("pwd"), 1)) } } @@ -780,25 +898,23 @@ fun resolve_cwd() { # reports, NO hardcoded "phone home" URLs. This check enforces that the # LLM endpoint is local/private by default. Set SWARM_CODE_ALLOW_REMOTE=1 # to bypass (e.g., when running against a remote LAN box intentionally). -fun verify_network_isolation(url, api_key, in_alt) { +# +# The URL rules live in Config.is_local_endpoint. This startup check only +# sees the primary endpoint — so the refusal is explained up front — and +# llm.sw applies the same rule at the point of dial to every other URL +# (fallback, providers[], .profile_override). +fun verify_network_isolation(url, in_alt) { bypass = getenv("SWARM_CODE_ALLOW_REMOTE") - has_auth = if (api_key == nil) { 'false' } - else { if (string_length(to_string(api_key)) == 0) { 'false' } - else { 'true' }} - host = extract_host(url) - is_local = is_local_host(host) - if (bypass == "1") { print(" " ++ UI.warn_text("⚠ SWARM_CODE_ALLOW_REMOTE=1 — network isolation disabled")) "ok" } - else { if (is_local == 'true') { - "ok" - } - else { if (has_auth == 'true') { + else { if (Config.is_local_endpoint(url) == 'true') { "ok" } else { + host = Config.endpoint_host(url) + host_shown = if (host == nil) { "(not a plain http(s)://host[:port] URL)" } else { host } # Hard refusal — if we're inside the optional alt screen, leave it # BEFORE printing: sys_exit would otherwise discard the explanation # with the alt buffer and leave the terminal un-restored. @@ -807,117 +923,23 @@ fun verify_network_isolation(url, api_key, in_alt) { print(UI.brand_color() ++ "\e[1m⏺ swarm-code" ++ UI.reset() ++ " refuses to contact non-local endpoints.") print("") print(" endpoint : " ++ url) - print(" host : " ++ host) + print(" host : " ++ host_shown) print("") print(" This is to guarantee no conversation data leaves your") - print(" network. swarm-code only allows:") + print(" network. swarm-code only allows http(s) URLs (no user@) to:") print(" - loopback: 127.*, ::1, localhost") print(" - private RFC1918: 10.*, 172.16-31.*, 192.168.*") print(" - Tailscale CGNAT: 100.64.0.0/10") + print(" - IPv6 ULA / link-local: [fc00::/7], [fe80::/10]") print(" - .local hostnames (mDNS)") print(" - Tailscale MagicDNS hosts (*.ts.net, bare hostnames)") print("") - print(" Set SWARM_CODE_API_KEY=... or SWARM_CODE_ALLOW_REMOTE=1 to bypass.") + print(" Set SWARM_CODE_ALLOW_REMOTE=1 to use a remote endpoint") + print(" (an API key alone does not lift this check).") print("") sys_exit(1) "denied" - }}} -} - -# Extract the host portion from a URL. Handles http:// and https://, -# and recognises bracketed IPv6 literals (`http://[::1]:8000/...`). -fun extract_host(url) { - after_scheme = if (string_starts_with(url, "https://") == 'true') { - string_sub(url, 8, string_length(url) - 8) - } else { - if (string_starts_with(url, "http://") == 'true') { - string_sub(url, 7, string_length(url) - 7) - } else { - url - } - } - # IPv6 literal: [::1] or [fd00::1]. Pull the body between brackets - # before any path/port splitting. - if (string_starts_with(after_scheme, "[") == 'true') { - rest = string_sub(after_scheme, 1, string_length(after_scheme) - 1) - close_parts = string_split(rest, "]") - hd(close_parts) - } else { - no_path_parts = string_split(after_scheme, "/") - host_with_port = hd(no_path_parts) - port_parts = string_split(host_with_port, ":") - hd(port_parts) - } -} - -# Return 'true' if host is on a local/private/Tailscale network. -fun is_local_host(host) { - if (host == "localhost") { 'true' } - else { if (host == "127.0.0.1") { 'true' } - else { if (host == "::1") { 'true' } - else { if (string_starts_with(host, "127.") == 'true') { 'true' } - else { if (string_starts_with(host, "10.") == 'true') { 'true' } - else { if (string_starts_with(host, "192.168.") == 'true') { 'true' } - else { if (is_172_private(host) == 'true') { 'true' } - else { if (is_100_cgnat(host) == 'true') { 'true' } - else { if (string_ends_with(host, ".local") == 'true') { 'true' } - else { if (string_ends_with(host, ".ts.net") == 'true') { 'true' } - else { if (is_ipv6_private(host) == 'true') { 'true' } - else { - # Bare hostname (no dots, no colons) = probably mDNS or - # Tailscale MagicDNS. We can't tell apart a public bare host - # from a private one without DNS resolution; mDNS / MagicDNS - # is the overwhelmingly common case on dev laptops, so allow. - if (string_contains(host, ".") == 'false' - && string_contains(host, ":") == 'false') { 'true' } - else { 'false' } - }}}}}}}}}}} -} - -# IPv6 private ranges (best-effort): loopback (already above), ULA -# fc00::/7 (starts with `fc` or `fd`), link-local fe80::/10. -fun is_ipv6_private(host) { - if (string_starts_with(host, "fc") == 'true') { 'true' } - else { if (string_starts_with(host, "fd") == 'true') { 'true' } - else { if (string_starts_with(host, "fe8") == 'true') { 'true' } - else { if (string_starts_with(host, "fe9") == 'true') { 'true' } - else { if (string_starts_with(host, "fea") == 'true') { 'true' } - else { if (string_starts_with(host, "feb") == 'true') { 'true' } - else { 'false' }}}}}} -} - -fun is_172_private(host) { - # 172.16.0.0/12 = 172.16.* through 172.31.* - if (string_starts_with(host, "172.") == 'false') { 'false' } - else { - parts172 = string_split(host, ".") - if (length(parts172) < 2) { 'false' } - else { - second172 = hd(tl(parts172)) - n172 = parse_int_simple(second172) - if (n172 < 16) { 'false' } - else { - if (n172 > 31) { 'false' } else { 'true' } - } - } - } -} - -fun is_100_cgnat(host) { - # 100.64.0.0/10 = 100.64.* through 100.127.* - if (string_starts_with(host, "100.") == 'false') { 'false' } - else { - parts100 = string_split(host, ".") - if (length(parts100) < 2) { 'false' } - else { - second100 = hd(tl(parts100)) - n100 = parse_int_simple(second100) - if (n100 < 64) { 'false' } - else { - if (n100 > 127) { 'false' } else { 'true' } - } - } - } + }} } fun parse_int_simple(s) { @@ -1014,6 +1036,11 @@ fun run_doctor() { api_key = map_get(opts, 'api_key') print(" ✓ model: " ++ model) print(" ✓ endpoint: " ++ endpoint) + if (Config.endpoint_refusal(endpoint) != nil) { + print(" ⚠ endpoint is not on the local network — swarm-code will refuse to start") + print(" unless SWARM_CODE_ALLOW_REMOTE=1 (an API key alone is not enough)") + warnings = warnings + 1 + } if (api_key == nil || string_length(to_string(api_key)) == 0) { print(" ⚠ api_key: (not set — fine for local endpoints, fatal for remote)") warnings = warnings + 1 @@ -1036,8 +1063,13 @@ fun run_doctor() { hdrs = if (api_key == nil) { [{"Accept", "application/json"}] } else { [{"Accept", "application/json"}, {"Authorization", "Bearer " ++ to_string(api_key)}] } - resp = http_get(models_url, hdrs) - if (resp == nil) { + # Same gate as the LLM layer: no request (and no API key) goes to a + # remote host unless SWARM_CODE_ALLOW_REMOTE=1. + refused = Config.endpoint_refusal(models_url) + resp = if (refused != nil) { nil } else { http_get(models_url, hdrs) } + if (refused != nil) { + print(" - skipped: not contacting a remote endpoint without SWARM_CODE_ALLOW_REMOTE=1") + } else { if (resp == nil) { print(" ✗ " ++ models_url ++ " — no response (network down? endpoint typo?)") errors = errors + 1 } else { @@ -1063,7 +1095,7 @@ fun run_doctor() { to_string(length(data)) ++ " model(s) advertised") } }} - } + }} print("") # 6. Skills / memory inventory diff --git a/src/mcp.sw b/src/mcp.sw index a354b91..e2655a8 100644 --- a/src/mcp.sw +++ b/src/mcp.sw @@ -468,7 +468,17 @@ fun mcp_owner_loop(name, handle, next_id) { } } -# Issue one tools/call and return the result as a tool-result string. +# Issue one tools/call. Returns a STRUCTURED outcome, never a bare +# string, so health bookkeeping (mcp_note_result) keys on what actually +# happened on the wire — not on the tool's output text, which is +# arbitrary data (a log line reading "db connection lost" is a +# SUCCESSFUL result and must not mark the server failed): +# {'ok', text} server answered with a result +# {'error', text} server answered, but with a JSON-RPC error / +# isError result / malformed response — alive +# {'timeout', text} no answer within mcp_call_timeout_ms (strike) +# {'lost', text} write failed or the pipe hit EOF — process gone +# `text` is the tool-result string handed to the model in every case. fun mcp_do_call(handle, id, tool_name, args) { args_obj = if (args == nil) { %{} } else { args } req = %{ @@ -479,61 +489,141 @@ fun mcp_do_call(handle, id, tool_name, args) { } sent = subprocess_send_line(handle, json_encode(req)) if (sent != 'ok') { - "error: MCP server connection lost (write failed)" + {'lost', "error: MCP server connection lost (write failed)"} } else { - resp = mcp_read_response(handle, id, timestamp() + mcp_call_timeout_ms()) - if (resp == nil) { - "error: MCP server did not respond (timed out or connection closed)" + r = mcp_read_reply(handle, id, timestamp() + mcp_call_timeout_ms()) + tag = elem(r, 0) + if (tag == 'ok') { mcp_format_result(elem(r, 1)) } + else { if (tag == 'lost') { + {'lost', "error: MCP server connection lost (the server exited or closed its output)"} } else { - mcp_format_result(resp) - } + {'timeout', "error: MCP server did not respond within " ++ + to_string(mcp_call_timeout_ms() / 1000) ++ "s"} + }} } } -# Drain JSON-RPC lines off the pipe until the response with `want_id` -# arrives, skipping notifications and server→client requests (whose -# id won't match). Bounded by `deadline` (a timestamp() value). -fun mcp_read_response(handle, want_id, deadline) { +# How early (ms before the deadline) a nil from subprocess_recv_line +# must arrive to count as EOF rather than a timeout. recv_line returns +# nil EARLY only on EOF / read error; a real timeout waits out the full +# window (then up to one 100ms select tick more), so anything clearly +# before the deadline is the server's stdout closing. +fun mcp_eof_slack_ms() { 50 } + +# Drain JSON-RPC lines off the pipe until the RESPONSE to `want_id` +# arrives. Tagged outcome, bounded by `deadline` (a timestamp() value): +# {'ok', msg} our response (result, error, or a malformed reply) +# {'timeout'} deadline passed with the pipe still open +# {'lost'} EOF / read error — the server process is gone +# JSON-RPC ids are per-direction: a server numbering its OWN requests +# can reuse the id of ours (e.g. {"id":100,"method":"roots/list"} while +# our call 100 is in flight), so id alone never identifies our reply — +# a message carrying `method` is a server→client request (id present) +# or notification (no id) and is never a response. Requests get an +# answer so the server is not left blocked waiting on us; then we keep +# reading for the real response. +fun mcp_read_reply(handle, want_id, deadline) { now = timestamp() - if (now >= deadline) { nil } + if (now >= deadline) { {'timeout'} } else { line = subprocess_recv_line(handle, deadline - now) if (line == nil) { - nil + if (timestamp() < deadline - mcp_eof_slack_ms()) { {'lost'} } + else { {'timeout'} } } else { decoded = json_decode(string_trim(line)) - if (decoded == nil) { - # Non-JSON line on stdout (a stray banner) — skip it. - mcp_read_response(handle, want_id, deadline) - } else { - rid = map_get(decoded, 'id') - # Compare ids as strings — some servers echo numeric ids - # back as strings, which would otherwise never match. - if (rid != nil && to_string(rid) == to_string(want_id)) { decoded } - else { mcp_read_response(handle, want_id, deadline) } + kind = mcp_msg_kind(decoded, want_id) + if (kind == 'response') { {'ok', decoded} } + else { + if (kind == 'request') { mcp_answer_server_request(handle, decoded) } + # 'notification' / 'other' (stray banner, stale reply + # to an earlier timed-out call) — skip and keep reading. + mcp_read_reply(handle, want_id, deadline) } } } } -# Turn a JSON-RPC tools/call response into a string for the model. +# Handshake-path wrapper: the decoded response, or nil on timeout/EOF +# (mcp_handshake_one / mcp_list_tools report either as a boot failure). +fun mcp_read_response(handle, want_id, deadline) { + r = mcp_read_reply(handle, want_id, deadline) + if (elem(r, 0) == 'ok') { elem(r, 1) } else { nil } +} + +# Classify one decoded line relative to the request we are waiting on: +# 'request' server→client request (method + non-null id) +# 'notification' server→client notification (method, no id) +# 'response' no method, id matches ours (compared as strings — +# some servers echo numeric ids back as strings) +# 'other' junk / non-object / a reply to some other id +# A matching-id message with neither result nor error is still +# returned as 'response' so the call fails fast ("carried no result") +# instead of waiting out the whole call timeout. +fun mcp_msg_kind(msg, want_id) { + if (msg == nil) { 'other' } + else { if (is_map(msg) == 'false') { 'other' } + else { + rid = map_get(msg, 'id') + if (mcp_has_key(msg, "method") == 'true') { + if (rid == nil) { 'notification' } else { 'request' } + } else { + if (rid != nil && to_string(rid) == to_string(want_id)) { 'response' } + else { 'other' } + } + }} +} + +# Key presence by name. map_has_key reports 'false' for a key whose +# value is JSON null, and json_decode keys are atoms — scan map_keys. +fun mcp_has_key(m, k) { mcp_key_scan(map_keys(m), k) } + +fun mcp_key_scan(keys, k) { + if (length(keys) == 0) { 'false' } + else { if (to_string(hd(keys)) == k) { 'true' } + else { mcp_key_scan(tl(keys), k) } } +} + +# Answer a server→client request. `ping` MUST get an empty result (MCP +# spec, basic/utilities/ping); we advertise no client capabilities +# (roots / sampling / elicitation), so everything else gets -32601 — +# an explicit error the server can act on, rather than silence that +# leaves it blocked on a reply that never comes. +fun mcp_server_request_reply(msg) { + id = map_get(msg, 'id') + method = to_string(map_get(msg, 'method')) + if (method == "ping") { + %{jsonrpc: "2.0", id: id, result: %{}} + } else { + %{jsonrpc: "2.0", id: id, + error: %{code: -32601, message: "method not found: " ++ method}} + } +} + +fun mcp_answer_server_request(handle, msg) { + subprocess_send_line(handle, json_encode(mcp_server_request_reply(msg))) +} + +# Turn a JSON-RPC tools/call response into {'ok'|'error', text}. fun mcp_format_result(resp) { err = map_get(resp, 'error') if (err != nil) { - em = map_get(err, 'message') - "error: MCP — " ++ (if (em == nil) { "call failed" } else { to_string(em) }) + em = if (is_map(err) == 'true') { map_get(err, 'message') } else { err } + {'error', "error: MCP — " ++ (if (em == nil) { "call failed" } else { to_string(em) })} } else { result = map_get(resp, 'result') if (result == nil) { - "error: MCP response carried no result" + {'error', "error: MCP response carried no result"} + } else { if (is_map(result) == 'false') { + {'error', "error: MCP response result is not an object"} } else { content = map_get(result, 'content') text = if (content == nil) { "" } else { mcp_extract_text(content, "") } body = if (string_length(string_trim(text)) == 0) { "(tool returned no text content)" } else { text } is_err = map_get(result, 'isError') - if (is_err == 'true') { "error: " ++ body } else { body } - } + if (is_err == 'true') { {'error', "error: " ++ body} } else { {'ok', body} } + }} } } @@ -579,34 +669,23 @@ fun mcp_reply_deadline_ms() { 260000 } # Wait for the owner's reply carrying our exact correlation token. A # reply with any other token is a stale result from an earlier call # that already timed out — drop it and keep waiting until the deadline. +# Returns the owner's {status, text} outcome (see mcp_do_call); an owner +# that never replies counts as a timeout strike against the server. fun mcp_await_result(server, token, deadline) { wait = deadline - timestamp() - if (wait <= 0) { - "error: MCP call to '" ++ server ++ "' got no reply (owner stalled or server hung)" - } else { + stalled = {'timeout', "error: MCP call to '" ++ server ++ + "' got no reply (owner stalled or server hung)"} + if (wait <= 0) { stalled } + else { receive { {'mcp_result', t, r} -> if (t == token) { r } else { mcp_await_result(server, token, deadline) } - after wait { - "error: MCP call to '" ++ server ++ "' got no reply (owner stalled or server hung)" - } + after wait { stalled } } } } -# A write failure means the pipe (and so the server process) is gone. -fun mcp_is_conn_lost(r) { - if (string_contains(to_string(r), "connection lost") == 'true') { 'true' } - else { 'false' } -} - -# A read timeout — the server may just have been slow on this one call. -fun mcp_is_timeout_result(r) { - if (string_contains(to_string(r), "did not respond") == 'true') { 'true' } - else { 'false' } -} - # A failed server gets one automatic reconnect attempt when a call next # needs it, at most once per this cooldown — beyond that it stays # fast-fail until the user runs /mcp reconnect. @@ -627,15 +706,18 @@ fun mcp_maybe_lazy_reconnect(table, server, opts) { ets_get(table, server ++ "/status") } -# Record a call outcome against the server's health. A lost connection -# marks it failed at once; a timeout only after two in a row (one slow -# call shouldn't kill the server); anything else resets the strikes. -fun mcp_note_result(table, server, r) { - if (mcp_is_conn_lost(r) == 'true') { +# Record a call outcome against the server's health, keyed on the +# STRUCTURED status from mcp_do_call — never on result text. A lost +# connection (EOF / write failure) marks it failed at once so the next +# call takes the lazy-reconnect path; a timeout only after two in a row +# (one slow call shouldn't kill the server); 'ok' and 'error' (the +# server answered, even if with an error) reset the strikes. +fun mcp_note_result(table, server, status) { + if (status == 'lost') { ets_put(table, server ++ "/status", "failed") - ets_put(table, server ++ "/error", "stopped responding mid-session") + ets_put(table, server ++ "/error", "connection lost mid-session (server exited)") } else { - if (mcp_is_timeout_result(r) == 'true') { + if (status == 'timeout') { prev = ets_get(table, server ++ "/fail_count") n = if (prev == nil) { 1 } else { prev + 1 } ets_put(table, server ++ "/fail_count", n) @@ -680,8 +762,8 @@ fun call_tool(prefixed, args, opts) { token = to_string(self()) ++ "/" ++ to_string(timestamp()) send(owner, {'mcp_call', tool, args, self(), token}) r = mcp_await_result(server, token, timestamp() + mcp_reply_deadline_ms()) - mcp_note_result(table, server, r) - r + mcp_note_result(table, server, elem(r, 0)) + elem(r, 1) } } } diff --git a/src/plan.sw b/src/plan.sw index 5b0f193..e3d0007 100644 --- a/src/plan.sw +++ b/src/plan.sw @@ -35,6 +35,7 @@ module Plan # then continue to list_append + run_turn as normal. import UI +import Config export [ init, generate, display, confirm, @@ -377,7 +378,8 @@ fun generate(user_msg, history, opts) { show_wait = if (map_get(opts, 'headless') == 'true') { 'false' } else { if (map_get(opts, 'is_subagent') == 'true') { 'false' } else { 'true' } } if (show_wait == 'true') { print_inline("\r\e[K \e[38;5;240m⋯ generating plan…\e[0m") } - resp = http_post(url, hdrs, body) + # Same network-isolation gate as every LLM dial (llm.sw dial_refusal). + resp = if (Config.endpoint_refusal(url) != nil) { nil } else { http_post(url, hdrs, body) } if (show_wait == 'true') { UI.tool_progress_clear() } if (resp == nil) { nil } else { plan_extract_content(resp) } diff --git a/src/prompts.sw b/src/prompts.sw index 8921f4d..943f8ff 100644 --- a/src/prompts.sw +++ b/src/prompts.sw @@ -15,7 +15,7 @@ module Prompts # own. When we eventually serve Claude-backed models too, we can # branch on model name and emit a matching shape. -export [system_prompt] +export [system_prompt, tool_sections] # system_prompt now takes a second arg: tool_format ('native' | 'inband'). # In native mode we skip the in-band protocol section (model gets the @@ -23,6 +23,25 @@ export [system_prompt] # native-mode instruction. Tool descriptions stay — they add context for # the long-tail tools that don't have JSON schemas yet. fun system_prompt(cwd, tool_format) { + # The sw idiom guide is ~3k tokens and only useful when the project is a + # sw/swarmrt codebase. Gate it on a cheap cwd probe so non-sw projects + # don't pay the prefill every turn (every extra token brings the + # prefill/compaction deadlock closer). + sw_section = if (is_sw_project(cwd) == 'true') { + "\n\n=== WRITING SW CODE ===\n" ++ sw_guide() + } else { "" } + preamble() ++ + "\n\n=== ENVIRONMENT ===\n" ++ environment_section(cwd) ++ + tool_sections(tool_format) ++ + sw_section ++ + "\n\n=== RULES ===\n" ++ rules() +} + +# The tool-use part of the system prompt for one wire format ("native" / +# "inband") — a pure function of the format, so llm.sw can swap it by exact +# text when a request goes out in the OTHER format than the prompt was built +# for (/profile switching tool_format, a fallback endpoint, a provider chain). +fun tool_sections(tool_format) { protocol_section = if (tool_format == "native") { "\n\n=== TOOL USE ===\n" ++ "Function tools are provided via this request's `tools` array. When " ++ @@ -43,19 +62,7 @@ fun system_prompt(cwd, tool_format) { } else { "\n\n=== AVAILABLE TOOLS ===\n" ++ tool_descriptions() } - # The sw idiom guide is ~3k tokens and only useful when the project is a - # sw/swarmrt codebase. Gate it on a cheap cwd probe so non-sw projects - # don't pay the prefill every turn (every extra token brings the - # prefill/compaction deadlock closer). - sw_section = if (is_sw_project(cwd) == 'true') { - "\n\n=== WRITING SW CODE ===\n" ++ sw_guide() - } else { "" } - preamble() ++ - "\n\n=== ENVIRONMENT ===\n" ++ environment_section(cwd) ++ - protocol_section ++ - tools_section ++ - sw_section ++ - "\n\n=== RULES ===\n" ++ rules() + protocol_section ++ tools_section } # Does cwd look like a sw / swarmrt project? True when cwd OR cwd/src holds diff --git a/src/scheduler.sw b/src/scheduler.sw index 9b5e5c9..9073bca 100644 --- a/src/scheduler.sw +++ b/src/scheduler.sw @@ -1,6 +1,7 @@ module Scheduler import Util +import JsonCheck # ============================================================ # Scheduler — interval-based recurring agent runs @@ -25,11 +26,28 @@ import Util # Supported expressions (all interpreted internally as milliseconds # because timestamp() — and therefore every comparison here — is ms): # -# 30s, 5m, 2h, 1d interval +# 30s, 5m, 2h, 1d interval — a whole positive count + one unit +# (strict: "1.5h", "10x5m", "5 m", "-1h" are rejected) # daily HH:MM fire once per day at the given UTC time -# hourly every hour on the hour (alias for 1h) +# (H or HH, exactly MM: "daily 9:05", "daily 23:59") +# hourly at the top of every hour (UTC wallclock, :00) — +# NOT "60 minutes after creation"; use 1h for that # daily every 24h (alias for 1d) # +# Robustness: schedule.json is hand-editable, so every read validates. +# The file must hold a JSON array; entries that are not objects, or +# whose id / expr / prompt / numeric fields don't validate, are SKIPPED +# (left untouched on disk) with a one-time warning — a bad entry never +# panics main's heartbeat handler. A file that exists but does not +# parse as an array is never overwritten: /schedule refuses to write +# instead of silently replacing every job. Writes go through +# file_atomic_write (temp file + rename), so a crash mid-write can't +# truncate the schedule. +# +# Dispatched children run with SWARM_CODE_DENY_DANGEROUS=1 (like +# /flows): a headless child otherwise auto-approves every 'ask', and a +# job firing unattended must not auto-run dangerous bash. +# # Dispatcher: hooked off the Heartbeat. Every tick the heartbeat # calls Scheduler.tick(opts), which walks jobs, computes "is this due # now?", and shells out `swarm-code -p ""` in the background @@ -41,9 +59,10 @@ import Util # scheduled--.out so the user can `tail` later. export [ - load, add, remove, list_all, tick, + load, add, add_checked, remove, list_all, tick, schedule_path, jobs_dir, parse_expr, parse_interval, daily_time_ms, compute_next_fire, + expr_error, read_state_at, normalize_job, add_at, tick_at, dispatch_cmd, swarm_binary_path, prune_old_out_files, pause_job, resume_job ] @@ -62,80 +81,248 @@ fun load() { } # ------------------------------------------------------------ -# Read jobs from disk. Returns [] if file missing/unparseable. +# Reading schedule.json # ------------------------------------------------------------ -fun list_all() { - p = schedule_path() - if (file_exists(p) == 'false') { [] } +# read_state_at(path) → {'ok', raw_entries} | {'corrupt', why} +# missing or blank file → {'ok', []} +# unreadable / not JSON / not an array → {'corrupt', why} +# "Not JSON" is decided by the STRICT JsonCheck.valid, not by +# json_decode returning nil: the runtime decoder guesses its way +# through damage ("[{…} ,,, oops" decodes to a list padded with nils), +# and rewriting such a guess would destroy the jobs it mangled. +# The raw entries are returned UNVALIDATED so writers can round-trip +# entries they don't understand instead of dropping them; readers go +# through list_all() / normalize_job. +fun read_state_at(p) { + if (file_exists(p) == 'false') { {'ok', []} } else { c = file_read(p) - if (c == nil) { [] } + if (c == nil) { {'corrupt', "could not be read"} } else { - decoded = json_decode(string_trim(c)) - if (decoded == nil) { [] } else { decoded } + trimmed = string_trim(c) + if (string_length(trimmed) == 0) { {'ok', []} } + else { if (JsonCheck.valid(trimmed) == 'false') { {'corrupt', "is not valid JSON"} } + else { + decoded = json_decode(trimmed) + if (decoded == nil) { {'corrupt', "is not valid JSON"} } + else { if (is_list(decoded) == 'false') { + {'corrupt', "must be a JSON array of jobs (found " ++ typeof(decoded) ++ ")"} + } else { {'ok', decoded} }} + }} } } } -fun save_all(jobs) { - file_write(schedule_path(), json_encode(jobs)) - 'ok' +# Every valid, normalized job (invalid entries skipped). [] when the +# file is missing or corrupt — readers never see a non-list. +fun list_all() { valid_jobs_of(read_state_at(schedule_path())) } + +fun valid_jobs_of(state) { + if (elem(state, 0) != 'ok') { [] } + else { valid_jobs_loop(elem(state, 1), []) } +} + +fun valid_jobs_loop(entries, acc) { + if (length(entries) == 0) { acc } + else { + r = normalize_job(hd(entries)) + next = if (elem(r, 0) == 'ok') { list_append(acc, elem(r, 1)) } else { acc } + valid_jobs_loop(tl(entries), next) + } } +# Crash-safe write: file_atomic_write writes .tmp. and +# rename(2)s it over the old file, so a crash mid-write leaves either +# the old schedule or the new one — never a truncated file that the +# next read would treat as corrupt. Returns 'ok' | 'error'. +fun save_at(p, entries) { file_atomic_write(p, json_encode(entries)) } + # ------------------------------------------------------------ -# Add a new job. Returns the assigned id (string), or nil if expr -# doesn't parse. +# normalize_job(raw) → {'ok', job} | {'invalid', why} # ------------------------------------------------------------ +# Validates one hand-editable entry. The returned job is the raw map +# with its fields coerced to the types tick() relies on (extra keys +# are preserved), so a fire can write it back without losing data: +# id string or integer, [A-Za-z0-9_-]+ (it names pid/.out files) +# expr a schedule expression parse_expr accepts +# prompt a non-empty string +# created_at / last_run non-negative integer ms (numeric strings +# coerced; floats truncated; missing → 0) +# runs non-negative integer (missing / junk → 0 — cosmetic) +# paused true/"true" → 'true', anything else → 'false' +fun normalize_job(raw) { + if (is_map(raw) == 'false') { {'invalid', "entry is not a JSON object"} } + else { + id_v = map_get(raw, 'id') + id_s = if (id_v == nil) { "" } + else { if (typeof(id_v) == "string" || typeof(id_v) == "int") { to_string(id_v) } + else { "" } } + expr_v = map_get(raw, 'expr') + prompt_v = map_get(raw, 'prompt') + last_run = coerce_ms(map_get(raw, 'last_run')) + created = coerce_ms(map_get(raw, 'created_at')) + runs_c = coerce_ms(map_get(raw, 'runs')) + runs = if (runs_c == nil) { 0 } else { runs_c } + paused_v = map_get(raw, 'paused') + paused = if (paused_v == 'true' || to_string(paused_v) == "true") { 'true' } else { 'false' } + label = if (string_length(id_s) == 0) { "?" } else { id_s } + if (safe_id(id_s) == 'false') { + {'invalid', "job " ++ label ++ ": id must be a number or [A-Za-z0-9_-] string"} + } else { if (expr_v == nil || typeof(expr_v) != "string" || parse_expr(to_string(expr_v)) == nil) { + {'invalid', "job " ++ label ++ ": unparseable expr " ++ json_encode(expr_v)} + } else { if (prompt_v == nil || typeof(prompt_v) != "string" || + string_length(string_trim(to_string(prompt_v))) == 0) { + {'invalid', "job " ++ label ++ ": prompt must be a non-empty string"} + } else { if (last_run == nil) { + {'invalid', "job " ++ label ++ ": last_run must be a timestamp in ms, got " ++ + json_encode(map_get(raw, 'last_run'))} + } else { if (created == nil) { + {'invalid', "job " ++ label ++ ": created_at must be a timestamp in ms, got " ++ + json_encode(map_get(raw, 'created_at'))} + } else { + j1 = map_put(raw, 'id', id_s) + j2 = map_put(j1, 'expr', string_trim(to_string(expr_v))) + j3 = map_put(j2, 'last_run', last_run) + j4 = map_put(j3, 'created_at', created) + j5 = map_put(j4, 'runs', runs) + {'ok', map_put(j5, 'paused', paused)} + }}}}} + } +} + +# Non-negative integer from a JSON value: nil → 0 (field absent), int +# as-is, float truncated, all-digit string parsed; anything else (a +# negative, "yesterday", a list, true) → nil = invalid. +fun coerce_ms(v) { + if (v == nil) { 0 } + else { + t = typeof(v) + if (t == "int") { if (v >= 0) { v } else { nil } } + else { if (t == "float") { if (v >= 0) { to_int(v) } else { nil } } + else { if (t == "string") { + if (all_digits(v) == 'true') { to_int(v) } else { nil } + } else { nil } } } + } +} + +fun safe_id(s) { + if (string_length(s) == 0 || string_length(s) > 64) { 'false' } + else { safe_id_loop(s, 0) } +} + +fun safe_id_loop(s, i) { + if (i >= string_length(s)) { 'true' } + else { + c = codepoint_at(s, i) + ok = (c >= 48 && c <= 57) || (c >= 65 && c <= 90) || + (c >= 97 && c <= 122) || c == 95 || c == 45 + if (ok) { safe_id_loop(s, i + 1) } else { 'false' } + } +} + +# 'true' iff s is one or more ASCII digits and nothing else. +fun all_digits(s) { + if (string_length(s) == 0) { 'false' } else { all_digits_loop(s, 0) } +} + +fun all_digits_loop(s, i) { + if (i >= string_length(s)) { 'true' } + else { + c = codepoint_at(s, i) + if (c >= 48 && c <= 57) { all_digits_loop(s, i + 1) } else { 'false' } + } +} + +# ------------------------------------------------------------ +# Add a new job. +# ------------------------------------------------------------ +# add_checked(expr, prompt) → {'ok', id} | {'error', message}. Refuses +# to write — rather than silently replacing every job — when +# schedule.json exists but does not parse as an array: the old code +# treated a corrupt file as [] and overwrote it with just the new job. +fun add_checked(expr, prompt) { add_at(schedule_path(), expr, prompt) } + +# add(expr, prompt) → id string, or nil on any failure. The specific +# reason (bad expression, corrupt schedule.json, write failure) is +# printed here, since the /schedule caller only sees nil. fun add(expr, prompt) { - expr_str = string_trim(to_string(expr)) - # Validate: either a known interval OR a daily HH:MM form. - if (parse_expr(expr_str) == nil && daily_time_ms(expr_str) == nil) { nil } + r = add_checked(expr, prompt) + if (elem(r, 0) == 'ok') { elem(r, 1) } else { - jobs = list_all() - id = to_string(next_id(jobs, 0)) - job = %{ - id: id, - expr: expr_str, - prompt: to_string(prompt), - created_at: timestamp(), - # Seed last_run to NOW, not 0. With 0 (epoch), compute_next_fire - # returns a 1970 timestamp that is always < now, so the next 2s - # heartbeat fires the job immediately on creation (and a past-slot - # daily HH:MM fires right away) instead of after one interval. - last_run: timestamp(), - runs: 0, - paused: 'false' - } - save_all(list_append(jobs, job)) - id + print("\e[38;5;208m✗ " ++ to_string(elem(r, 1)) ++ "\e[0m") + nil } } +fun add_at(p, expr, prompt) { + expr_str = string_trim(to_string(expr)) + bad_expr = expr_error(expr_str) + prompt_str = to_string(prompt) + if (bad_expr != nil) { {'error', bad_expr} } + else { if (string_length(string_trim(prompt_str)) == 0) { + {'error', "schedule prompt must not be empty"} + } else { + state = read_state_at(p) + if (elem(state, 0) != 'ok') { + {'error', p ++ " " ++ to_string(elem(state, 1)) ++ + " — refusing to overwrite it (fix or move the file, then retry)"} + } else { + entries = elem(state, 1) + id = to_string(next_id(entries, 0)) + job = %{ + id: id, + expr: expr_str, + prompt: prompt_str, + created_at: timestamp(), + # Seed last_run to NOW, not 0. With 0 (epoch), compute_next_fire + # returns a 1970 timestamp that is always < now, so the next 2s + # heartbeat fires the job immediately on creation (and a past-slot + # daily HH:MM fires right away) instead of after one interval. + last_run: timestamp(), + runs: 0, + paused: 'false' + } + # Append to the RAW entries: entries this version can't + # validate are preserved on disk, not dropped by the rewrite. + if (save_at(p, list_append(entries, job)) == 'ok') { {'ok', id} } + else { {'error', "could not write " ++ p} } + } + }} +} + +# Next free numeric id over the raw entries (non-map / non-numeric ids +# are ignored, never crash). fun next_id(jobs, max_so_far) { if (length(jobs) == 0) { max_so_far + 1 } else { j = hd(jobs) - v = to_string(map_get(j, 'id')) - n = parse_int_simple(v) + n = if (is_map(j) == 'false') { 0 } + else { parse_int_simple(to_string(map_get(j, 'id'))) } new_max = if (n > max_so_far) { n } else { max_so_far } next_id(tl(jobs), new_max) } } +# Remove by id. Never writes a corrupt file (nothing to remove there). fun remove(id) { - target = to_string(id) - jobs = list_all() - kept = remove_loop(jobs, target, []) - if (length(kept) == length(jobs)) { 'false' } - else { save_all(kept) ; 'true' } + p = schedule_path() + state = read_state_at(p) + if (elem(state, 0) != 'ok') { 'false' } + else { + jobs = elem(state, 1) + kept = remove_loop(jobs, to_string(id), []) + if (length(kept) == length(jobs)) { 'false' } + else { save_at(p, kept) ; 'true' } + } } fun remove_loop(jobs, target, acc) { if (length(jobs) == 0) { acc } else { j = hd(jobs) - new_acc = if (to_string(map_get(j, 'id')) == target) { acc } - else { list_append(acc, j) } + hit = if (is_map(j) == 'false') { 'false' } + else { if (to_string(map_get(j, 'id')) == target) { 'true' } else { 'false' } } + new_acc = if (hit == 'true') { acc } else { list_append(acc, j) } remove_loop(tl(jobs), target, new_acc) } } @@ -146,8 +333,25 @@ fun remove_loop(jobs, target, acc) { # actually fired (the dirty-flag gate). Previously this rewrote on # every tick because tick_loop always rebuilt a same-length list. # ------------------------------------------------------------ +# This runs INSIDE main's heartbeat handler: a panic here kills the +# interactive session. Hence tick_at's contract — any file content +# (wrong shape, bad field types) is skipped + reported, never raised. fun tick(opts, count) { - jobs = list_all() + r = tick_at(schedule_path(), count) + warn_problems(opts, r) + map_get(r, 'status') +} + +# tick_at(path, count) → %{status, fired, problems} +# status 'noop' (no valid jobs) | 'ok' | 'corrupt' +# fired number of jobs dispatched this tick +# problems human-readable reasons for every skipped entry / a +# corrupt file (surfaced once per distinct problem set) +fun tick_at(p, count) { + state = read_state_at(p) + ok_state = if (elem(state, 0) == 'ok') { 'true' } else { 'false' } + entries = if (ok_state == 'true') { elem(state, 1) } else { [] } + valid = valid_jobs_loop(entries, []) # Prune .out files on a ~15-minute cadence (tick 1, then every 450 # ticks at the 2s default), NEVER per-tick: prune shells out, and # shell() runs on the CALLER's fiber — this is main's heartbeat @@ -155,26 +359,69 @@ fun tick(opts, count) { # tick interval and froze the UI behind a growing backlog # (2026-07-09). Still prunes even when the job list is empty or # corrupt, so .out files can't accumulate unbounded. - if (count % 450 == 1) { prune_old_out_files(jobs) } else { 'skip' } - if (length(jobs) == 0) { 'noop' } - else { + if (count % 450 == 1) { prune_old_out_files(valid) } else { 'skip' } + if (ok_state == 'false') { + %{status: 'corrupt', fired: 0, + problems: [p ++ " " ++ to_string(elem(state, 1)) ++ " — no scheduled jobs will run"]} + } else { now = timestamp() - r = tick_loop(jobs, now, [], 'false') + r = tick_loop(entries, now, [], 0, []) updated = elem(r, 0) - dirty = elem(r, 1) - if (dirty == 'true') { save_all(updated) } - 'ok' + fired = elem(r, 1) + problems = elem(r, 2) + if (fired > 0) { save_at(p, updated) } + %{status: (if (length(valid) == 0) { 'noop' } else { 'ok' }), + fired: fired, problems: problems} } } -fun tick_loop(jobs, now, acc, dirty) { - if (length(jobs) == 0) { {acc, dirty} } +# Walk the RAW entries: invalid ones are carried through untouched +# (and reported), valid ones get a fire check. A fired job is written +# back in its normalized form; an unfired one stays byte-identical. +fun tick_loop(entries, now, acc, fired, problems) { + if (length(entries) == 0) { {acc, fired, problems} } + else { + raw = hd(entries) + n = normalize_job(raw) + if (elem(n, 0) != 'ok') { + tick_loop(tl(entries), now, list_append(acc, raw), fired, + list_append(problems, to_string(elem(n, 1)) ++ " — skipped")) + } else { + result = maybe_fire(elem(n, 1), now) + did = elem(result, 1) + out = if (did == 'true') { elem(result, 0) } else { raw } + tick_loop(tl(entries), now, list_append(acc, out), + (if (did == 'true') { fired + 1 } else { fired }), problems) + } + } +} + +# One-time warning per distinct problem set: the heartbeat ticks every +# 2s, so re-printing would spam the prompt. The last-warned set lives +# on the heartbeat ETS table (main's session state); print_above keeps +# the pinned input line intact. A fixed file clears the latch, so a +# later breakage warns again. +fun warn_problems(opts, r) { + problems = map_get(r, 'problems') + table = if (opts == nil) { nil } else { map_get(opts, 'heartbeat_table') } + sig = if (problems == nil) { "" } else { json_encode(problems) } + if (table == nil) { 'skip' } + else { if (problems == nil || length(problems) == 0) { + ets_put(table, 'sched_warned', nil) + 'ok' + } else { if (ets_get(table, 'sched_warned') == sig) { 'skip' } + else { + ets_put(table, 'sched_warned', sig) + print_above("\e[38;5;208m⚠ schedule: " ++ problem_lines(problems, "") ++ "\e[0m") + 'warned' + }}} +} + +fun problem_lines(ps, acc) { + if (length(ps) == 0) { acc } else { - result = maybe_fire(hd(jobs), now) - new_job = elem(result, 0) - fired = elem(result, 1) - new_dirty = if (fired == 'true') { 'true' } else { dirty } - tick_loop(tl(jobs), now, list_append(acc, new_job), new_dirty) + sep = if (string_length(acc) == 0) { "" } else { "; " } + problem_lines(tl(ps), acc ++ sep ++ to_string(hd(ps))) } } @@ -212,15 +459,19 @@ fun maybe_fire(job, now) { # ------------------------------------------------------------ # compute_next_fire — when (in ms-since-epoch) the next fire is # scheduled for, given the expression and the prior fire's epoch ms. -# Two semantics: +# Three semantics: # * "daily HH:MM" — wallclock. Compute today's HH:MM slot in UTC. # If that's already past last_run, fire then. Otherwise wait # for tomorrow's slot. -# * intervals (30s/5m/2h/1d/hourly/daily) — last_run + interval_ms. +# * "hourly" — wallclock, the first top-of-the-hour (:00 UTC) after +# last_run. A missed hour (machine asleep) fires once on wake, +# then realigns to the next :00. +# * intervals (30s/5m/2h/1d/daily) — last_run + interval_ms. # Returns nil if the expression is unparseable. # ------------------------------------------------------------ fun compute_next_fire(expr, last_run, now) { - daily_ms_offset = daily_time_ms(expr) + trimmed = string_trim(to_string(expr)) + daily_ms_offset = daily_time_ms(trimmed) if (daily_ms_offset != nil) { day_ms = 86400000 today_midnight = (now / day_ms) * day_ms @@ -230,52 +481,84 @@ fun compute_next_fire(expr, last_run, now) { if (today_slot > last_run) { today_slot } else { today_slot + day_ms } } + else { if (trimmed == "hourly") { + hour_ms = 3600000 + (last_run / hour_ms + 1) * hour_ms + } else { - interval_ms = parse_expr(expr) + interval_ms = parse_expr(trimmed) if (interval_ms == nil) { nil } else { last_run + interval_ms } - } + }} } # parse_expr — return the interval in MILLISECONDS for supported -# expression forms. Returns nil for unparseable, including the -# "daily HH:MM" form (callers route those through daily_time_ms + -# compute_next_fire's wallclock branch). "daily" alone is the 24h -# interval; "daily HH:MM" returns 86400000 here too so add()'s -# validation accepts both shapes uniformly. +# expression forms, nil for anything else. "hourly" and "daily HH:MM" +# are wallclock-aligned (see compute_next_fire) but still report +# their nominal period here so validation accepts every shape +# uniformly. Strict: see expr_error for what is rejected and why. fun parse_expr(s) { - trimmed = string_trim(s) - if (string_length(trimmed) == 0) { nil } - else { if (trimmed == "hourly") { 3600000 } + if (expr_error(to_string(s)) == nil) { expr_period_ms(string_trim(to_string(s))) } + else { nil } +} + +fun expr_period_ms(trimmed) { + if (trimmed == "hourly") { 3600000 } else { if (trimmed == "daily") { 86400000 } else { if (daily_time_ms(trimmed) != nil) { 86400000 } - else { parse_interval(trimmed) }}}} + else { parse_interval(trimmed) }}} +} + +# expr_error(s) → nil when `s` is a valid schedule expression, else a +# one-line reason. The old parsers read the leading digits and ignored +# the rest, so "1.5h" ran hourly, "10x5m" every 10 minutes, "daily :" +# at 00:00 and "daily 9:5x" at 09:05 — all silently accepted. +fun expr_error(s) { + t = string_trim(to_string(s)) + q = "'" ++ t ++ "'" + hint = " — use e.g. 30s, 5m, 2h, 1d, hourly, daily or daily 09:00" + if (string_length(t) == 0) { "empty schedule expression" ++ hint } + else { if (t == "hourly" || t == "daily") { nil } + else { if (string_starts_with(t, "daily ") == 'true') { + if (daily_time_ms(t) != nil) { nil } + else { "invalid time in " ++ q ++ ": want daily HH:MM (UTC, 00:00-23:59, e.g. daily 9:05)" } + } + else { if (parse_interval(t) != nil) { nil } + else { + "invalid schedule " ++ q ++ ": an interval is a whole number plus s/m/h/d" ++ hint + }}}} } -# parse_interval — `30s`, `5m`, `2h`, `1d` → MILLISECONDS. -# Previously this returned seconds, which silently mismatched -# timestamp()'s ms units and made every interval fire ~3.6 seconds -# after creation (with back-pressure masking it). Fixed: ms throughout. +# parse_interval — `30s`, `5m`, `2h`, `1d` → MILLISECONDS, or nil. +# Strict: the count must be ONLY digits (no sign, decimal point, space +# or other junk), >= 1 and <= 9 digits (no int overflow), followed by +# exactly one unit letter. Previously this returned seconds, which +# silently mismatched timestamp()'s ms units; ms throughout now. fun parse_interval(s) { n = string_length(s) if (n < 2) { nil } else { suffix = string_sub(s, n - 1, 1) num_str = string_sub(s, 0, n - 1) - num = parse_int_simple(num_str) - if (num <= 0) { nil } + if (all_digits(num_str) == 'false' || string_length(num_str) > 9) { nil } else { - if (suffix == "s") { num * 1000 } - else { if (suffix == "m") { num * 60000 } - else { if (suffix == "h") { num * 3600000 } - else { if (suffix == "d") { num * 86400000 } - else { nil }}}} + num = parse_int_simple(num_str) + if (num <= 0) { nil } + else { + if (suffix == "s") { num * 1000 } + else { if (suffix == "m") { num * 60000 } + else { if (suffix == "h") { num * 3600000 } + else { if (suffix == "d") { num * 86400000 } + else { nil }}}} + } } } } # daily_time_ms — parse "daily HH:MM" -> ms since midnight, or nil -# if the expression isn't a daily-with-time form. Used by +# if the expression isn't a (valid) daily-with-time form. Strict: the +# hour is 1-2 digits, the minute exactly 2, nothing else — "daily :", +# "daily 9:5x", "daily 9:5" and "daily 24:00" are all nil. Used by # compute_next_fire to schedule wallclock-aligned fires. fun daily_time_ms(s) { if (string_starts_with(s, "daily ") == 'false') { nil } @@ -284,12 +567,17 @@ fun daily_time_ms(s) { parts = string_split(time_part, ":") if (length(parts) != 2) { nil } else { - h_str = string_trim(hd(parts)) - m_str = string_trim(hd(tl(parts))) - h = parse_int_simple(h_str) - m = parse_int_simple(m_str) - if (h < 0 || h > 23 || m < 0 || m > 59) { nil } - else { (h * 3600 + m * 60) * 1000 } + h_str = hd(parts) + m_str = hd(tl(parts)) + shape_ok = all_digits(h_str) == 'true' && all_digits(m_str) == 'true' && + string_length(h_str) <= 2 && string_length(m_str) == 2 + if (shape_ok == 'false') { nil } + else { + h = parse_int_simple(h_str) + m = parse_int_simple(m_str) + if (h < 0 || h > 23 || m < 0 || m > 59) { nil } + else { (h * 3600 + m * 60) * 1000 } + } } } } @@ -319,19 +607,29 @@ fun dispatch(job) { prompt = to_string(map_get(job, 'prompt')) ts = to_string(timestamp()) out_path = jobs_dir() ++ "/scheduled-" ++ id ++ "-" ++ ts ++ ".out" - bin = swarm_binary_path() - # Pidfile records "PID LSTART" (process start-time) so - # previous_fire_alive can detect a recycled PID — kill(pid,0) - # alone returns alive on EPERM, wedging the job in skipped_busy. - inner = - "nohup " ++ Util.shell_q(bin) ++ " --no-resume -p " ++ Util.shell_q(prompt) ++ - " > " ++ Util.shell_q(out_path) ++ " 2>&1 & SW_PID=$!; " ++ - "echo \"$SW_PID $(ps -o lstart= -p \"$SW_PID\" 2>/dev/null)\" > " ++ Util.shell_q(pid_file) - shell("bash -c " ++ Util.shell_q(inner)) + shell(dispatch_cmd(swarm_binary_path(), prompt, out_path, pid_file)) 'dispatched' } } +# The shell command dispatch() runs. Pure, so the safety prefix is +# unit-testable. SWARM_CODE_DENY_DANGEROUS=1 turns the dangerous-bash +# gate into a hard deny in the child — a headless run otherwise +# auto-approves every 'ask', and a job firing unattended (nobody at +# the terminal, possibly hours later) must not auto-run `rm -rf ~/…`. +# Same rule as Flows.build_task_cmd. +# Pidfile records "PID LSTART" (process start-time) so +# previous_fire_alive can detect a recycled PID — kill(pid,0) +# alone returns alive on EPERM, wedging the job in skipped_busy. +fun dispatch_cmd(bin, prompt, out_path, pid_file) { + inner = + "SWARM_CODE_DENY_DANGEROUS=1 nohup " ++ Util.shell_q(bin) ++ + " --no-resume -p " ++ Util.shell_q(prompt) ++ + " > " ++ Util.shell_q(out_path) ++ " 2>&1 & SW_PID=$!; " ++ + "echo \"$SW_PID $(ps -o lstart= -p \"$SW_PID\" 2>/dev/null)\" > " ++ Util.shell_q(pid_file) + "bash -c " ++ Util.shell_q(inner) +} + # Has the previous fire's child exited? Cheap kill(pid, 0) check via # the pid_alive builtin first — no shell() overhead on the common # (dead) path. pid_alive returns 'true' on EPERM too, so a recycled @@ -447,19 +745,23 @@ fun resume_job(id) { } fun set_paused(target_id, paused_val) { - jobs = list_all() - r = set_paused_loop(jobs, target_id, paused_val, [], 'false') - new_jobs = elem(r, 0) - found = elem(r, 1) - if (found == 'true') { save_all(new_jobs) ; 'true' } - else { 'false' } + p = schedule_path() + state = read_state_at(p) + if (elem(state, 0) != 'ok') { 'false' } + else { + r = set_paused_loop(elem(state, 1), target_id, paused_val, [], 'false') + new_jobs = elem(r, 0) + found = elem(r, 1) + if (found == 'true') { save_at(p, new_jobs) ; 'true' } + else { 'false' } + } } fun set_paused_loop(jobs, target_id, paused_val, acc, found) { if (length(jobs) == 0) { {acc, found} } else { j = hd(jobs) - if (to_string(map_get(j, 'id')) == target_id) { + if (is_map(j) == 'true' && to_string(map_get(j, 'id')) == target_id) { new_j = map_put(j, 'paused', paused_val) set_paused_loop(tl(jobs), target_id, paused_val, list_append(acc, new_j), 'true') } else { diff --git a/src/skills.sw b/src/skills.sw index 3d35c75..440e0ee 100644 --- a/src/skills.sw +++ b/src/skills.sw @@ -43,7 +43,7 @@ import Util export [ load, skills_dir, index_path, skill_dir, skill_file_path, save, recall, list_index, forget, - as_prompt_section, slugify + as_prompt_section, slugify, valid_slug ] fun skills_dir() { getenv("HOME") ++ "/.swarm-code/skills" } @@ -64,6 +64,14 @@ fun load() { # ------------------------------------------------------------ fun save(name, description, triggers, instructions) { slug = slugify(to_string(name)) + if (valid_slug(slug) == 'false') { + "error: skill name must contain letters or digits" + } else { + save_as(slug, name, description, triggers, instructions) + } +} + +fun save_as(slug, name, description, triggers, instructions) { dir = skill_dir(slug) file_mkdir(dir) body = @@ -82,7 +90,40 @@ fun save(name, description, triggers, instructions) { } } +# ------------------------------------------------------------ +# Slug validation — recall/forget take the slug straight from the +# model's tool call and splice it into a path. `../../../work/proj` +# made forget_skill delete work/proj/SKILL.md and recall_skill read it. +# A slug is ONE path component under skills_dir(): non-empty, no '/' or +# '\', no '..', no leading '.', no control characters. save() already +# slugifies names to [a-z0-9_]; hand-made skill directories with other +# safe names (dashes, capitals, spaces) keep working. +# ------------------------------------------------------------ +fun valid_slug(slug) { + s = if (slug == nil) { "" } else { to_string(slug) } + if (string_length(s) == 0 || string_length(s) > 128) { 'false' } + else { if (string_contains(s, "/") == 'true' || string_contains(s, "\\") == 'true') { 'false' } + else { if (string_contains(s, "..") == 'true' || string_starts_with(s, ".") == 'true') { 'false' } + else { slug_no_ctrl(s, 0) } } } +} + +fun slug_no_ctrl(s, i) { + if (i >= string_length(s)) { 'true' } + else { if (codepoint_at(s, i) < 32 || codepoint_at(s, i) == 127) { 'false' } + else { slug_no_ctrl(s, i + 1) } } +} + +fun invalid_slug_msg(slug) { + "error: invalid skill slug '" ++ to_string(slug) ++ + "' — use a slug from the skills index (one name, no '/', '\\', '..' or leading '.')" +} + fun recall(slug) { + if (valid_slug(slug) == 'false') { invalid_slug_msg(slug) } + else { recall_valid(slug) } +} + +fun recall_valid(slug) { p = skill_file_path(slug) if (file_exists(p) == 'false') { "error: no skill named '" ++ slug ++ "' — see /skills" @@ -103,6 +144,11 @@ fun list_index() { } fun forget(slug) { + if (valid_slug(slug) == 'false') { invalid_slug_msg(slug) } + else { forget_valid(slug) } +} + +fun forget_valid(slug) { p = skill_file_path(slug) if (file_exists(p) == 'false') { "error: no skill named '" ++ slug ++ "'" diff --git a/src/test_runner.sw b/src/test_runner.sw index 1992340..894523c 100644 --- a/src/test_runner.sw +++ b/src/test_runner.sw @@ -26,6 +26,13 @@ import UI import Memory import Tools import Mcp +import McpServer +import JsonCheck +import Flows +import Trajectory +import Log +import Skills +import SessionSearch import ToolGuardrails import Agent import Scheduler @@ -34,6 +41,12 @@ import MemVec import ToolExecutor import ToolRegistry import Background +import PathGuard +import Hooks +import ToolSchemas +import Util +import Prompts +import Browser fun main() { print("") @@ -124,6 +137,25 @@ fun main() { t_scheduler_next_id_empty(), t_scheduler_next_id_nonempty(), t_scheduler_jobs_dir_suffix(), + # --- agents / MCP / scheduler / persistence regressions --- + t_mcp_msg_kind_collision(), + t_mcp_health_structured(), + t_mcp_server_spec_envelope(), + t_json_check_strict(), + t_subagent_type_allowlists(), + t_subagent_result_capped(), + t_subagent_partial_keeps_work(), + t_flows_validate_shapes(), + t_flows_launch_quota(), + t_trajectory_redacts_valid_jsonl(), + t_log_redact_value_shapes(), + t_skill_slug_traversal_blocked(), + t_session_search_reindexes_changed(), + t_sched_wrong_shape_never_panics(), + t_sched_corrupt_refuses_write(), + t_sched_strict_exprs(), + t_sched_hourly_on_the_hour(), + t_sched_dispatch_denies_dangerous(), t_memory_embed_db_path(), t_memory_dir_suffix(), t_memory_slugify_spaces(), @@ -213,7 +245,89 @@ fun main() { t_ticker_phase_content(), t_ticker_phase_think_plus_tok(), t_stream_timeout_no_first_token(), - t_large_ctx_hint_once() + t_large_ctx_hint_once(), + # --- ESC ends the turn --- + t_interrupt_skips_rest_and_ends_turn(), + t_interrupt_not_on_normal_results(), + # --- read / edit correctness --- + t_read_js_is_text(), + t_read_line_count_and_empty(), + t_read_missing_and_binary(), + t_edit_crlf_file(), + # --- security: untrusted project settings --- + t_project_scope_strips_untrusted(), + t_project_scope_trusted_applies(), + t_project_ignored_keys(), + t_project_trusted_dir_match(), + # --- security: network-isolation gate --- + t_endpoint_gate_bypasses_refused(), + t_endpoint_gate_locals_allowed(), + t_endpoint_host_parse(), + t_endpoint_refusal_reason(), + # --- security: swarm-code control files --- + t_pathguard_control_files_blocked(), + t_pathguard_data_dirs_writable(), + t_pathguard_case_insensitive(), + # --- security: headless 'ask' --- + t_headless_ask_denied(), + t_headless_default_allowed_still_run(), + # --- tool-layer review fixes --- + t_bash_trailing_comment(), + t_bash_heredoc_last(), + t_bash_syntax_error_reaches_model(), + t_bg_trailing_comment(), + t_grep_invalid_regex_surfaces(), + t_grep_glob_default_path(), + t_file_watch_no_injection(), + t_file_watch_portable_mtime(), + t_wait_timeout_clamped(), + t_classifier_catches_bypasses(), + t_classifier_no_false_positives(), + t_command_tools_gated(), + t_denial_names_reason(), + t_sudo_tab_blocked(), + t_bg_sessions_isolated(), + t_pre_tool_hook_big_payload(), + t_configured_hook_big_payload(), + t_pre_tool_hook_fails_closed(), + t_hook_matcher_families(), + t_read_offset_past_64k(), + t_read_huge_file_window(), + t_edit_refuses_nul_and_huge(), + t_run_tests_quoting_timeout_output(), + t_edit_create_makes_parents(), + t_tilde_paths_expand(), + t_guardrail_counts_bash_exit_codes(), + t_schema_text_matches_code(), + # --- cut-off tool calls are never dispatched --- + t_json_args_well_formed(), + t_args_malformed_despite_lenient_decode(), + t_cut_turn_reason(), + t_cut_turn_calls_refused(), + t_sanitize_cut_tool_calls(), + # --- headless reports only this run's answer --- + t_headless_answer_this_run_only(), + # --- context budget scales with the window --- + t_context_budget_scales_with_window(), + # --- compaction keeps the live turn, never loses history --- + t_compact_split_keeps_live_user(), + t_compact_split_pair_boundary(), + t_compact_nothing_old_is_noop(), + t_compact_failed_summary_keeps_history(), + # --- a fatal 4xx keeps completed work; context overflow retries --- + t_context_overflow_detection(), + t_mech_trim_protects_live_user(), + # --- profile override: env precedence, kwargs, /model, tool format --- + t_override_env_beats_stale_override(), + t_profile_override_keeps_chat_template_kwargs(), + t_model_override_keeps_active_profile(), + t_system_prompt_follows_wire_format(), + # --- browser_screenshot goes through the write guard --- + t_browser_screenshot_path_guard(), + # --- swarm-code trust --- + t_trust_edits_settings(), + t_trust_refuses_corrupt_settings(), + t_json_pretty_round_trips() ] passed = sum_list(results, 0) @@ -576,11 +690,10 @@ fun t_drop_to_last_clean_user() { # F2: trailing_turn_incomplete keeps a COMPLETE tool turn (every tool_call # answered) and flags a PARTIAL one (a tool_call left unanswered) — the robust -# detection that doesn't depend on json_decode. (NB: the args-malformed sub-check -# is only a weak backup: sw's json_decode is lenient and recovers truncated JSON -# into a partial map rather than nil, so a mid-string-truncated tool_call is -# caught by F4's finish_reason/marker path and this partial-tool-set check, not -# by json_decode==nil.) +# detection that doesn't depend on json_decode. (NB: sw's json_decode is lenient +# and recovers truncated JSON into a partial map rather than nil; the +# args-malformed sub-check uses the strict Util.json_args_well_formed for that — +# see t_args_malformed_despite_lenient_decode.) fun t_trailing_turn_incomplete_detection() { complete = [%{role: 'user', content: "u"}, %{role: 'assistant', content: "", tool_calls: [%{id: "1", arguments: "{}"}]}, @@ -607,6 +720,127 @@ fun t_mcp_unconfigured() { check("mcp: unconfigured -> no schemas, no prompt section", ok) } +# 'true' iff every element of a list of 'true'/'false' atoms is 'true'. +# (ag_ prefix: helpers for the agents/MCP/scheduler regression block.) +fun ag_all(lst) { + if (length(lst) == 0) { 'true' } + else { if (hd(lst) != 'true') { 'false' } else { ag_all(tl(lst)) } } +} + +fun ag_is(a, b) { if (a == b) { 'true' } else { 'false' } } + +# A server→client request whose id COLLIDES with our in-flight call id +# (the reviewer's {"id":100,"method":"roots/list"}) used to be taken as +# the response → "MCP response carried no result". Anything carrying +# `method` is a request/notification, never our reply; ping is answered +# with {} and other requests with -32601. +fun t_mcp_msg_kind_collision() { + req = json_decode("{\"jsonrpc\":\"2.0\",\"id\":100,\"method\":\"roots/list\"}") + ping = json_decode("{\"jsonrpc\":\"2.0\",\"id\":\"p1\",\"method\":\"ping\"}") + note = json_decode("{\"jsonrpc\":\"2.0\",\"method\":\"notifications/progress\"}") + resp = json_decode("{\"jsonrpc\":\"2.0\",\"id\":100,\"result\":{}}") + sresp = json_decode("{\"jsonrpc\":\"2.0\",\"id\":\"100\",\"error\":{\"code\":1}}") + stale = json_decode("{\"jsonrpc\":\"2.0\",\"id\":99,\"result\":{}}") + ping_reply = Mcp.mcp_server_request_reply(ping) + roots_reply = Mcp.mcp_server_request_reply(req) + rerr = map_get(roots_reply, 'error') + ok = ag_all([ + ag_is(Mcp.mcp_msg_kind(req, 100), 'request'), + ag_is(Mcp.mcp_msg_kind(note, 100), 'notification'), + ag_is(Mcp.mcp_msg_kind(resp, 100), 'response'), + ag_is(Mcp.mcp_msg_kind(sresp, 100), 'response'), + ag_is(Mcp.mcp_msg_kind(stale, 100), 'other'), + ag_is(Mcp.mcp_msg_kind(json_decode("[1]"), 100), 'other'), + ag_is(Mcp.mcp_msg_kind(nil, 100), 'other'), + ag_is(json_encode(ping_reply), "{\"jsonrpc\":\"2.0\",\"id\":\"p1\",\"result\":{}}"), + ag_is(map_get(roots_reply, 'id'), 100), + ag_is(map_get(rerr, 'code'), -32601)]) + check("mcp: colliding-id server request is not our response; ping -> {}, others -> -32601", ok) +} + +# Health bookkeeping keys on the structured call status, never the text: +# a SUCCESSFUL result mentioning "connection lost" used to mark the +# server failed (forced reconnect, then "not running" for 60s), and +# "did not respond" in output counted as a timeout strike. EOF / write +# failure ('lost') must fail the server at once. +fun t_mcp_health_structured() { + ok_text = Mcp.mcp_format_result(json_decode( + "{\"id\":1,\"result\":{\"content\":[{\"type\":\"text\",\"text\":\"db connection lost; peer did not respond\"}]}}")) + is_err = Mcp.mcp_format_result(json_decode( + "{\"id\":1,\"result\":{\"content\":[{\"type\":\"text\",\"text\":\"boom\"}],\"isError\":true}}")) + rpc_err = Mcp.mcp_format_result(json_decode("{\"id\":1,\"error\":{\"code\":-1,\"message\":\"bad\"}}")) + t = ets_new() + ets_put(t, "s/status", "ok") + Mcp.mcp_note_result(t, "s", elem(ok_text, 0)) + Mcp.mcp_note_result(t, "s", elem(ok_text, 0)) + after_ok = ets_get(t, "s/status") + Mcp.mcp_note_result(t, "s", 'timeout') + after_one_timeout = ets_get(t, "s/status") + Mcp.mcp_note_result(t, "s", 'error') + Mcp.mcp_note_result(t, "s", 'timeout') + after_reset_timeout = ets_get(t, "s/status") + Mcp.mcp_note_result(t, "s", 'timeout') + after_two_timeouts = ets_get(t, "s/status") + ets_put(t, "s/status", "ok") + Mcp.mcp_note_result(t, "s", 'lost') + after_lost = ets_get(t, "s/status") + ets_drop(t) + ok = ag_all([ + ag_is(elem(ok_text, 0), 'ok'), + ag_is(elem(ok_text, 1), "db connection lost; peer did not respond"), + ag_is(elem(is_err, 0), 'error'), + ag_is(elem(rpc_err, 0), 'error'), + ag_is(after_ok, "ok"), + ag_is(after_one_timeout, "ok"), + ag_is(after_reset_timeout, "ok"), + ag_is(after_two_timeouts, "failed"), + ag_is(after_lost, "failed")]) + check("mcp: health follows call status (ok/error/timeout/lost), never result text", ok) +} + +# --mcp-server spec conformance: ping used to be -32601 (spec: empty +# result), an unknown tool -32601 (spec: -32602 Invalid params), and +# neither "jsonrpc" nor the id type was validated (object ids echoed). +fun ag_rpc(line) { + out = McpServer.dispatch(json_decode(line), + %{settings: map_new(), execution_context: "mcp_server"}) + if (out == nil) { nil } else { json_decode(to_string(out)) } +} + +fun ag_rpc_code(r) { + if (r == nil) { nil } + else { e = map_get(r, 'error') ; if (e == nil) { nil } else { map_get(e, 'code') } } +} + +fun t_mcp_server_spec_envelope() { + ping = McpServer.dispatch(json_decode("{\"jsonrpc\":\"2.0\",\"id\":7,\"method\":\"ping\"}"), + %{settings: map_new(), execution_context: "mcp_server"}) + unknown_tool = ag_rpc("{\"jsonrpc\":\"2.0\",\"id\":2,\"method\":\"tools/call\",\"params\":{\"name\":\"nope\",\"arguments\":{}}}") + bad_args = ag_rpc("{\"jsonrpc\":\"2.0\",\"id\":3,\"method\":\"tools/call\",\"params\":{\"name\":\"read\",\"arguments\":[1]}}") + unknown_method = ag_rpc("{\"jsonrpc\":\"2.0\",\"id\":4,\"method\":\"nope/x\"}") + wrong_ver = ag_rpc("{\"jsonrpc\":\"1.0\",\"id\":5,\"method\":\"ping\"}") + no_ver = ag_rpc("{\"id\":6,\"method\":\"ping\"}") + obj_id = ag_rpc("{\"jsonrpc\":\"2.0\",\"id\":{\"a\":1},\"method\":\"ping\"}") + null_id = ag_rpc("{\"jsonrpc\":\"2.0\",\"id\":null,\"method\":\"ping\"}") + str_id = ag_rpc("{\"jsonrpc\":\"2.0\",\"id\":\"abc\",\"method\":\"ping\"}") + note = ag_rpc("{\"jsonrpc\":\"2.0\",\"method\":\"notifications/initialized\"}") + ok = ag_all([ + ag_is(to_string(ping), "{\"jsonrpc\":\"2.0\",\"id\":7,\"result\":{}}"), + ag_is(ag_rpc_code(unknown_tool), -32602), + ag_is(ag_rpc_code(bad_args), -32602), + ag_is(ag_rpc_code(unknown_method), -32601), + ag_is(ag_rpc_code(wrong_ver), -32600), + ag_is(map_get(wrong_ver, 'id'), 5), + ag_is(ag_rpc_code(no_ver), -32600), + ag_is(ag_rpc_code(obj_id), -32600), + ag_is(map_get(obj_id, 'id'), nil), + ag_is(ag_rpc_code(null_id), -32600), + ag_is(map_get(str_id, 'id'), "abc"), + ag_is(ag_rpc_code(str_id), nil), + ag_is(note, nil)]) + check("mcp-server: ping -> {}, unknown tool -> -32602, jsonrpc/id validated (-32600)", ok) +} + # remember saved frontmatter but an empty body: the schema named the # content field `body` while do_remember read `content`, so the # mismatch silently dropped every memory's content. Guard: a remember @@ -884,6 +1118,232 @@ fun t_subagent_blocked_tool() { check("subagent_blocked: blocks task/remember, allows read/bash", ok) } +# A journal still being written by another instance when this one +# booted was marked indexed and never refreshed (only presence in meta +# was checked). Now size+mtime are recorded and a changed journal is +# reindexed at the next init — here: a line appended after indexing. +fun t_session_search_reindexes_changed() { + d = ag_tmp("sessidx") + file_delete(d) + file_mkdir(d) + j = d ++ "/journal-100.jsonl" + file_write(j, "{\"role\":\"user\",\"content\":\"alphaword first turn\"}\n") + SessionSearch.init_at(d) + h1 = SessionSearch.search_at(d, "alphaword", 10) + file_append(j, "{\"role\":\"assistant\",\"content\":\"zuluword appended later\"}\n") + SessionSearch.init_at(d) + h2 = SessionSearch.search_at(d, "zuluword", 10) + h3 = SessionSearch.search_at(d, "alphaword", 10) + SessionSearch.init_at(d) + h4 = SessionSearch.search_at(d, "alphaword", 10) + file_delete(j) + file_delete(d ++ "/index.db") + exec_argv("rmdir", [d]) + ok = ag_all([ + ag_is(length(h1), 1), + ag_is(length(h2), 1), + ag_is(length(h3), 1), + ag_is(length(h4), 1)]) + check("session search: a journal that changed after indexing is reindexed", ok) +} + +# forget_skill {"slug":"../../../work/proj"} deleted work/proj/SKILL.md +# and recall_skill read it: the slug was spliced into a path unchecked. +# Plant a SKILL.md outside skills_dir and aim a traversal slug at it. +fun t_skill_slug_traversal_blocked() { + d = ag_tmp("skilltrav") + file_delete(d) + file_mkdir(d) + target = d ++ "/SKILL.md" + file_write(target, "TRAVERSAL-TARGET-CONTENT") + slug = ag_repeat("../", 16, "") ++ string_sub(d, 1, string_length(d) - 1) + rec = Skills.recall(slug) + fgt = Skills.forget(slug) + survived = file_exists(target) + file_delete(target) + exec_argv("rmdir", [d]) + ok = ag_all([ + ag_is(string_contains(rec, "TRAVERSAL-TARGET-CONTENT"), 'false'), + string_starts_with(rec, "error: invalid skill slug"), + string_starts_with(fgt, "error: invalid skill slug"), + ag_is(survived, 'true'), + ag_is(Skills.valid_slug("deploy_mally_otp"), 'true'), + ag_is(Skills.valid_slug("ship-openear-dmg"), 'true'), + ag_is(Skills.valid_slug("a/b"), 'false'), + ag_is(Skills.valid_slug("a\\b"), 'false'), + ag_is(Skills.valid_slug(".."), 'false'), + ag_is(Skills.valid_slug(".hidden"), 'false'), + ag_is(Skills.valid_slug(""), 'false'), + ag_is(Skills.valid_slug(nil), 'false'), + ag_is(Skills.valid_slug("x\ny"), 'false'), + # slugify("") is "" — would have written skills_dir()//SKILL.md + string_starts_with(Skills.save("", "d", "t", "i"), "error:")]) + check("skills: traversal slugs rejected by recall/forget (target untouched)", ok) +} + +# Trajectory export ran Log.redact over the ENCODED line: the blob layer +# swallowed the `n` of a `\n` escape → `\[REDACTED]` → invalid JSONL; +# and AWS_SECRET_ACCESS_KEY=…/…, PGPASSWORD=, postgres://user:pw@, +# {"password": "…"} (space after colon), YAML password:, PEM lines with +# '/' all leaked. Every exported line must be strict JSON and carry none +# of the secrets. +fun t_trajectory_redacts_valid_jsonl() { + out = ag_tmp("traj") + blob = "Z9x8Y7w6V5u4T3s2R1q0P9o8N7m6L5k4J3i2H1g0Z9x8" + h = [LLM.new_message_user("connect with PGPASSWORD=pgsecret1 psql, or postgres://admin:urlsecret2@db:5432/x"), + LLM.new_message_assistant("on it", [%{id: "c1", name: "bash", + arguments: "{\"command\":\"export AWS_SECRET_ACCESS_KEY=wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY\"}"}], nil), + LLM.new_message_tool("c1", "config:\n password: yamlsecret3\n" ++ + "{\"password\": \"jsonsecret4\", \"api_key\": \"apisecret5xyz\"}\n" ++ + "-----BEGIN RSA PRIVATE KEY-----\nMIIEpemline/abc+def123\nsecondpemline/Zz9\n-----END RSA PRIVATE KEY-----\n" ++ + "tail\n" ++ blob ++ "\n\"quoted\" and a \\ backslash"), + LLM.new_message_assistant("done", [], nil)] + Trajectory.export_current(out, h) + body = file_read(out) + file_delete(out) + lines = filter(string_split(to_string(body), "\n"), fn(l) { string_length(string_trim(l)) > 0 }) + secrets = ["pgsecret1", "urlsecret2", "wJalrXUtnFEMI", "K7MDENG", "yamlsecret3", "jsonsecret4", + "apisecret5xyz", "MIIEpemline", "secondpemline", "Z9x8Y7w6V5u4"] + ok = ag_all([ + ag_is(length(lines), 1), + ag_all(map(fn(l) { JsonCheck.valid(l) }, lines)), + ag_all(map(fn(l) { if (json_decode(l) == nil) { 'false' } else { 'true' } }, lines)), + ag_all(map(fn(x) { ag_is(string_contains(to_string(body), x), 'false') }, secrets)), + string_contains(to_string(body), "postgres://admin:[REDACTED]@db"), + string_contains(to_string(body), "BEGIN RSA PRIVATE KEY")]) + check("trajectory: export is valid JSONL and masks env/URL/JSON/YAML/PEM secrets", ok) +} + +# redact_value walks decoded values (what Log.event now encodes): nested +# maps/lists redacted, keys and non-strings kept; the encoded result is +# strict JSON even when a masked run sits right after a newline. +fun t_log_redact_value_shapes() { + v = %{type: "tool_call", n: 3, ok: 'true', + args: "{\"api_key\": \"topsecret-api-value\"}", + nested: [%{note: "line\nZ9x8Y7w6V5u4T3s2R1q0P9o8N7m6L5k4J3i2H1g0Z9x8"}, nil, 7]} + r = Log.redact_value(v) + enc = json_encode(r) + keep = Log.redact("max_tokens: 4096 max_token: 4096 sort_key: id; if password == x; /usr/lib/x86_64-linux-gnu/libc.so.6") + ok = ag_all([ + JsonCheck.valid(enc), + ag_is(string_contains(enc, "topsecret"), 'false'), + ag_is(string_contains(enc, "Z9x8Y7"), 'false'), + ag_is(map_get(r, 'n'), 3), + ag_is(map_get(r, 'ok'), 'true'), + ag_is(length(map_get(r, 'nested')), 3), + ag_is(keep, "max_tokens: 4096 max_token: 4096 sort_key: id; if password == x; /usr/lib/x86_64-linux-gnu/libc.so.6")]) + check("log: redact_value masks nested strings, keeps structure; benign text untouched", ok) +} + +# /flows with {"phases":"oops"} panicked the interactive session (hd on +# a string in init_phases). validate_workflow rejects every shape the +# run would walk, before the alt-screen opens or anything launches. +fun t_flows_validate_shapes() { + bad = ["{\"phases\":\"oops\"}", "[1,2]", "\"str\"", + "{\"phases\":[5]}", "{\"phases\":[{\"tasks\":\"x\"}]}", + "{\"phases\":[{\"tasks\":[7]}]}", "{\"phases\":[{\"tasks\":[{\"label\":\"a\"}]}]}", + "{\"phases\":[{\"tasks\":[{\"prompt\":{\"x\":1}}]}]}", + "{\"phases\":[{\"tasks\":[{\"prompt\":\"p\",\"model\":[1]}]}]}"] + good = ["{}", "{\"phases\":[]}", "{\"phases\":[{\"name\":\"P\"}]}", + "{\"phases\":[{\"tasks\":[{\"prompt\":\"p\",\"label\":3}]}]}"] + ok = ag_all([ + ag_all(map(fn(x) { ag_is(elem(Flows.validate_workflow(json_decode(x)), 0), 'error') }, bad)), + ag_all(map(fn(x) { ag_is(elem(Flows.validate_workflow(json_decode(x)), 0), 'ok') }, good)), + string_contains(to_string(elem(Flows.validate_workflow(json_decode("{\"phases\":\"oops\"}")), 1)), + "\"phases\" must be an array")]) + check("flows: malformed workflow shapes rejected with a reason (no panic)", ok) +} + +# /flows fan-out was unbounded (12 tasks → 12 children in 0.15s). +# launch_quota = free slots under the cap, bounded by what is queued. +fun t_flows_launch_quota() { + q = %{bg_task_id: nil, status: 'pending'} + r = %{bg_task_id: "bg-1", status: 'running'} + d = %{bg_task_id: "bg-2", status: 'done'} + e = %{bg_task_id: nil, status: 'error'} + twelve = map(fn(i) { q }, 1..12) + ok = ag_all([ + ag_is(Flows.launch_quota(twelve, 4), 4), + ag_is(Flows.launch_quota([r, r, r, q, q], 4), 1), + ag_is(Flows.launch_quota([r, r, r, r, q], 4), 0), + ag_is(Flows.launch_quota([d, d, r, q, q, q], 2), 1), + ag_is(Flows.launch_quota([d, e, q], 4), 1), + ag_is(Flows.launch_quota([d, e], 4), 0), + ag_is(Flows.flows_max_parallel(), 4)]) + check("flows: fan-out capped (default 4 in flight), rest queued", ok) +} + +# explore / bash subagent restrictions were prompt-only — an explore +# subagent ran bash and write. Now a policy (ToolRegistry contexts + +# subagent_blocked_for), with a clear refusal naming what IS allowed. +fun t_subagent_type_allowlists() { + ok = ag_all([ + ag_is(Agent.subagent_blocked_for("explore", "bash"), 'true'), + ag_is(Agent.subagent_blocked_for("explore", "write"), 'true'), + ag_is(Agent.subagent_blocked_for("explore", "edit"), 'true'), + ag_is(Agent.subagent_blocked_for("explore", "web_fetch"), 'true'), + ag_is(Agent.subagent_blocked_for("explore", "mcp__srv__tool"), 'true'), + ag_is(Agent.subagent_blocked_for("explore", "read"), 'false'), + ag_is(Agent.subagent_blocked_for("explore", "grep"), 'false'), + ag_is(Agent.subagent_blocked_for("explore", "glob"), 'false'), + ag_is(Agent.subagent_blocked_for("bash", "bash"), 'false'), + ag_is(Agent.subagent_blocked_for("bash", "write"), 'true'), + ag_is(Agent.subagent_blocked_for("general", "write"), 'false'), + ag_is(Agent.subagent_blocked_for("general", "task"), 'true'), + ag_is(ToolRegistry.allowed_in("subagent_explore", "bash"), 'false'), + ag_is(ToolRegistry.allowed_in("subagent_explore", "read"), 'true'), + ag_is(Agent.subagent_type_of(nil), "general"), + ag_is(Agent.subagent_type_of(5), "general"), + ag_is(Agent.subagent_type_of(" Explore "), "explore"), + ag_is(Agent.subagent_context("explore"), "subagent_explore"), + string_contains(Agent.subagent_block_msg("explore", "bash"), "read-only")]) + check("subagent: explore = read-only, bash = shell only (enforced, not prompt-only)", ok) +} + +# A 320KB subagent answer reached the parent uncapped. Capped like +# bash/MCP output (24000) with an explicit marker; head + tail kept; +# cut points never split a UTF-8 sequence. +fun t_subagent_result_capped() { + big = "HEAD-MARK " ++ ag_repeat("finding: details details details\n", 10000, "") ++ " TAIL-MARK" + capped = Agent.subagent_cap(big) + # One ASCII byte first puts every 'é' (2 bytes) at an ODD offset, so + # a naive cut at 16000 would split one; a boundary-safe cut is odd. + uni = "x" ++ ag_repeat("é", 20000, "") + ucap = Agent.subagent_cap(uni) + marker_at = string_index_of(ucap, "\n\n…[") + ok = ag_all([ + if (string_length(capped) <= 24400) { 'true' } else { 'false' }, + string_contains(capped, "bytes of the subagent's answer elided"), + string_starts_with(capped, "HEAD-MARK "), + string_ends_with(capped, " TAIL-MARK"), + ag_is(Agent.subagent_cap("short answer"), "short answer"), + ag_is(marker_at % 2, 1), + string_ends_with(ucap, "é")]) + check("subagent: answer capped at 24000 with marker, UTF-8 safe", ok) +} + +# Hitting max steps returned only "[subagent hit max steps…]" — the +# work was discarded. subagent_partial keeps the latest notes plus a +# digest of the most recent tool calls/results, headed by the reason. +fun t_subagent_partial_keeps_work() { + h = [LLM.new_message_system("sys"), LLM.new_message_user("task"), + LLM.new_message_assistant("looking at config", [%{id: "c1", name: "bash", arguments: "{\"command\":\"echo F1\"}"}], nil), + LLM.new_message_tool("c1", "[exit 0]\nFINDING_ONE"), + LLM.new_message_assistant("", [%{id: "c2", name: "read", arguments: "{\"path\":\"x\"}"}], nil), + LLM.new_message_tool("c2", "FINDING_TWO")] + p = Agent.subagent_partial(h, "hit the 15-step limit before a final answer") + empty = Agent.subagent_partial([LLM.new_message_user("t")], "LLM call failed — boom") + ok = ag_all([ + string_contains(p, "subagent stopped: hit the 15-step limit"), + string_contains(p, "looking at config"), + string_contains(p, "FINDING_ONE"), + string_contains(p, "FINDING_TWO"), + string_contains(p, "read {\"path\":\"x\"}"), + string_contains(empty, "LLM call failed — boom"), + string_contains(empty, "no findings yet")]) + check("subagent: abnormal stop returns partial work + reason", ok) +} + # Non-interactive entry points cannot answer an "ask" permission # decision. They must fail closed instead of silently running it. fun t_tool_executor_noninteractive_ask_denied() { @@ -1398,6 +1858,142 @@ fun t_scheduler_jobs_dir_suffix() { string_ends_with(d, "telemetry")) } +# A fresh, empty temp file path (mkstemp) — no shell() round-trip. +fun ag_tmp(prefix) { file_temp("/tmp/swarm-test-" ++ prefix ++ "-") } + +fun ag_repeat(s, n, acc) { if (n <= 0) { acc } else { ag_repeat(s, n - 1, acc ++ s) } } + +# json_decode guesses through damage ("[1,,2]" → [1, nil, 2], "[tru]" +# → [nil, nil, nil]), so writers need a strict well-formedness check +# before trusting a file enough to rewrite it. +fun t_json_check_strict() { + bad = ["[1, 2", "[1 2]", "{\"a\":1} x", "[{\"a\":1},]", "[", "{\"a\":}", "[tru]", + "[\"abc]", "{\"a\" 1}", "[1,,2]", "01", "1.", "\"a\\q\"", "", "nul"] + good = ["[]", "{}", " [1, -2.5e+3, 0, true, false, null, \"x\\\"\\u00e9\", {\"k\": [{}]}] ", + "\"é\"", "-0", "1E5"] + long_arr = "[" ++ ag_repeat("1,", 20000, "") ++ "1]" + ok = ag_all([ + ag_all(map(fn(x) { ag_is(JsonCheck.valid(x), 'false') }, bad)), + ag_all(map(fn(x) { ag_is(JsonCheck.valid(x), 'true') }, good)), + ag_is(JsonCheck.valid(long_arr), 'true')]) + check("json-check: strict RFC 8259 validity (rejects what json_decode guesses at)", ok) +} + +# schedule.json of the wrong SHAPE used to crash every interactive +# session ~2s after launch, inside main's heartbeat handler: +# {"jobs":[]} decoded to a map → hd() panic in prune_jobs_loop, and a +# job with "last_run":"yesterday" → "str" + int panic in +# compute_next_fire. tick_at must report, skip and never panic; valid +# entries still load (numeric strings coerced), invalid ones untouched. +fun t_sched_wrong_shape_never_panics() { + p = ag_tmp("sched-shape") + file_write(p, "{\"jobs\":[]}") + r1 = Scheduler.tick_at(p, 2) + far = "99999999999999" + body = "[5, \"str\", null, " ++ + "{\"id\":\"1\",\"expr\":\"1h\",\"prompt\":\"x\",\"last_run\":\"yesterday\"}," ++ + "{\"id\":\"2\",\"expr\":\"1h\",\"prompt\":\"y\",\"last_run\":\"" ++ far ++ "\"}," ++ + "{\"id\":\"../x\",\"expr\":\"1h\",\"prompt\":\"z\"}," ++ + "{\"id\":\"4\",\"expr\":\"1.5h\",\"prompt\":\"w\"}]" + file_write(p, body) + r2 = Scheduler.tick_at(p, 2) + disk_now = file_read(p) + valid = Scheduler.valid_jobs_of(Scheduler.read_state_at(p)) + file_delete(p) + ok = ag_all([ + ag_is(map_get(r1, 'status'), 'corrupt'), + ag_is(length(map_get(r1, 'problems')), 1), + ag_is(map_get(r2, 'status'), 'ok'), + ag_is(map_get(r2, 'fired'), 0), + ag_is(length(map_get(r2, 'problems')), 6), + ag_is(disk_now, body), + ag_is(length(valid), 1), + ag_is(map_get(hd(valid), 'last_run'), 99999999999999)]) + check("scheduler: wrong-shape schedule.json / bad fields are skipped, never panic", ok) +} + +# /schedule on a corrupt schedule.json used to treat it as [] and +# overwrite it with just the new job — silently deleting every other +# job. It must refuse (file byte-identical); entries it can't parse in +# an otherwise-valid array survive an add. +fun t_sched_corrupt_refuses_write() { + p = ag_tmp("sched-corrupt") + file_write(p, "[{\"id\":\"1\",\"expr\":\"1h\",\"prompt\":\"keep me\"} ,,, oops") + before = file_read(p) + r1 = Scheduler.add_at(p, "5m", "new job") + same1 = file_read(p) + file_write(p, "{\"jobs\":[{\"id\":\"1\"}]}") + r2 = Scheduler.add_at(p, "5m", "new job") + same2 = file_read(p) + file_write(p, "[7, {\"id\":\"3\",\"expr\":\"1h\",\"prompt\":\"old\"}]") + r3 = Scheduler.add_at(p, "5m", "new job") + grown = Scheduler.read_state_at(p) + file_delete(p) + r4 = Scheduler.add_at(p, "bogus", "x") + missing_after_bad = file_exists(p) + r5 = Scheduler.add_at(p, "2h", "fresh") + fresh = Scheduler.read_state_at(p) + file_delete(p) + entries = elem(grown, 1) + ok = ag_all([ + ag_is(elem(r1, 0), 'error'), + string_contains(to_string(elem(r1, 1)), "refusing to overwrite"), + ag_is(same1, before), + ag_is(elem(r2, 0), 'error'), + ag_is(same2, "{\"jobs\":[{\"id\":\"1\"}]}"), + ag_is(r3, {'ok', "4"}), + ag_is(length(entries), 3), + ag_is(hd(entries), 7), + ag_is(elem(r4, 0), 'error'), + ag_is(missing_after_bad, 'false'), + ag_is(r5, {'ok', "1"}), + ag_is(length(elem(fresh, 1)), 1)]) + check("scheduler: add refuses to overwrite a corrupt schedule.json, keeps unknown entries", ok) +} + +# Loose parsing accepted garbage by reading leading digits: 1.5h ran +# hourly, 10x5m every 10m, "daily :" at 00:00, "daily 9:5x" at 09:05. +fun t_sched_strict_exprs() { + rejects = ["1.5h", "10x5m", "5 m", "-1h", "+1h", "0m", "1234567890s", "5", "m", + "daily :", "daily 9:5x", "daily 9:5", "daily 24:00", "daily 12:60", + "daily 9", "daily :30", "hourlyx", "1h30m"] + accepts = ["30s", "5m", "2h", "1d", "hourly", "daily", "daily 9:05", "daily 23:59", " 10m "] + ok = ag_all([ + ag_all(map(fn(e) { ag_is(Scheduler.parse_expr(e), nil) }, rejects)), + ag_all(map(fn(e) { if (Scheduler.expr_error(e) == nil) { 'false' } else { 'true' } }, rejects)), + ag_all(map(fn(e) { if (Scheduler.parse_expr(e) == nil) { 'false' } else { 'true' } }, accepts)), + ag_is(Scheduler.daily_time_ms("daily 9:05"), (9 * 3600 + 5 * 60) * 1000), + ag_is(elem(Scheduler.add_at("/nonexistent-dir/s.json", "1.5h", "x"), 0), 'error')]) + check("scheduler: strict expression parsing rejects 1.5h / 10x5m / 'daily :' / 'daily 9:5x'", ok) +} + +# `hourly` is documented as "every hour on the hour" but ran 60 min +# after creation. Next fire is now the first :00 (UTC) after last_run. +fun t_sched_hourly_on_the_hour() { + hour = 3600000 + last = 1779635000000 # 15:03:20 UTC — mid-hour + nf = Scheduler.compute_next_fire("hourly", last, last + 1000) + on_hour = (last / hour + 1) * hour + nf2 = Scheduler.compute_next_fire("hourly", on_hour, on_hour + 5) + nf_1h = Scheduler.compute_next_fire("1h", last, last + 1000) + ok = ag_all([ + ag_is(nf, on_hour), + ag_is(nf % hour, 0), + if (nf > last && nf - last < hour) { 'true' } else { 'false' }, + ag_is(nf2, on_hour + hour), + ag_is(nf_1h, last + hour)]) + check("scheduler: hourly fires at the top of the hour, 1h stays relative", ok) +} + +# Scheduled children run unattended in headless mode, which +# auto-approves 'ask' — without SWARM_CODE_DENY_DANGEROUS=1 a job whose +# model ran `rm -rf ~/victim` deleted it. /flows already set it. +fun t_sched_dispatch_denies_dangerous() { + cmd = Scheduler.dispatch_cmd("/bin/swarm", "clean up", "/tmp/o.out", "/tmp/p.pid") + check("scheduler: dispatched jobs run with SWARM_CODE_DENY_DANGEROUS=1", + string_contains(cmd, "SWARM_CODE_DENY_DANGEROUS=1 nohup ")) +} + # ------------------------------------------------------------ # MEMORY — save/recall round-trip and path guards # ------------------------------------------------------------ @@ -1527,7 +2123,7 @@ fun t_bg_auto_after_ms_no_double_finalize() { backgrounded = bool_and3( string_contains(s, "[backgrounded"), string_contains(s, id), - string_contains(s, Background.log_path_for(id))) + string_contains(s, Background.log_path_for(table, id))) st = Background.wait_for_task(table, id, 6000) done_ok = if (st == 'done') { 'true' } else { 'false' } @@ -1577,7 +2173,7 @@ fun t_bg_stdin_devnull_no_hang() { # bg_kill must SIGTERM the WHOLE process group, not just the /bin/sh wrapper. # Launch a wrapper that forks a backgrounded sleep plus a foreground sleep; -# after kill_task + a short grace, pgrep -g must find zero survivors. +# after kill_task + a short grace, pgrep -g must find zero live survivors. fun t_bg_kill_group() { table = Background.init() id = Background.launch(table, "sh -c 'sleep 60 & sleep 60; wait'", "kill probe") @@ -1586,7 +2182,11 @@ fun t_bg_kill_group() { pid = ets_get(table, id ++ "/pid") Background.kill_task(table, id) sleep(600) - r = shell("pgrep -g " ++ to_string(pid) ++ " 2>/dev/null | wc -l | tr -d ' \n'") + # Count LIVE members only: a killed child whose parent is gone becomes a + # zombie until init reaps it, and container inits (Docker without + # --init, sandboxes) never do — zombies have already terminated. + r = shell("for p in $(pgrep -g " ++ to_string(pid) ++ " 2>/dev/null); do " ++ + "ps -o stat= -p $p 2>/dev/null; done | grep -vc '^ *Z' | tr -d ' \n'") count = string_trim(to_string(elem(r, 1))) check("bg_kill terminates the whole process group (0 survivors)", if (count == "0") { 'true' } else { 'false' }) @@ -2543,3 +3143,1148 @@ fun t_sf_four_backtick_fence() { eqs(Markdown.fence_info("````"), ""))) check("fence: 4-backtick opener needs >=4 closer; fence_info('````') is empty", ok) } + +# ESC on a tool ends the turn: the remaining calls of the batch each get a +# [skipped] result (so every tool_call id still has its tool message) and +# turn_interrupted() tells run_turn not to call the model again. +fun t_interrupt_skips_rest_and_ends_turn() { + base = [LLM.new_message_user("go"), + LLM.new_message_tool("a", "[stopped by user (ESC/Ctrl-C) — process group killed]")] + rest = [%{id: "b", name: "bash", arguments: "{}"}, %{id: "c", name: "read", arguments: "{}"}] + h = Agent.skip_remaining_tools(rest, base, %{}) + last2 = tl(tl(h)) + ids_ok = if (map_get(hd(last2), 'tool_call_id') == "b" && + map_get(hd(tl(last2)), 'tool_call_id') == "c") { 'true' } else { 'false' } + check("interrupt: remaining tool_calls get [skipped] results and the turn ends", + bool_and3(ids_ok, + if (length(h) == 4) { 'true' } else { 'false' }, + Agent.turn_interrupted(h, 3))) +} + +fun t_interrupt_not_on_normal_results() { + h = [LLM.new_message_tool("a", "[exit 0]\nok"), + LLM.new_message_tool("b", "[interrupted] tool 'web_fetch' was stopped by the user (ESC).")] + check("interrupt: detected for non-shell tools, not for normal results", + bool_and(if (Agent.turn_interrupted(take_first_n(h, 1), 1) == 'false') { 'true' } else { 'false' }, + Agent.turn_interrupted(h, 2))) +} + +fun take_first_n(lst, n) { + if (n <= 0 || length(lst) == 0) { [] } + else { [hd(lst) | take_first_n(tl(lst), n - 1)] } +} + +# libmagic calls .js/.ts "application/javascript"; read used to refuse them +# as binary. Binary is now "a NUL byte in the first 8KB". +fun t_read_js_is_text() { + p = "/tmp/swc_read_js.js" + file_write(p, "function f() {\n return 1;\n}\n") + r = Tools.exec_raw('read', %{path: p}, %{}) + file_delete(p) + check("read: a .js file is text (line-numbered, not 'binary')", + bool_and(string_starts_with(r, "1\tfunction f() {"), + if (string_contains(r, "binary") == 'false') { 'true' } else { 'false' })) +} + +# A trailing newline ends the last line (no phantom empty line); an empty +# file is reported as empty, not as binary. +fun t_read_line_count_and_empty() { + p = "/tmp/swc_read_lines.txt" + file_write(p, "a\nb\n") + r = Tools.exec_raw('read', %{path: p}, %{}) + e = "/tmp/swc_read_empty.txt" + file_write(e, "") + re = Tools.exec_raw('read', %{path: e}, %{}) + file_delete(p) + file_delete(e) + check("read: 2-line file reads as 2 lines; empty file says [empty file]", + bool_and(if (r == "1\ta\n2\tb") { 'true' } else { 'false' }, + string_starts_with(re, "[empty file]"))) +} + +fun t_read_missing_and_binary() { + m = Tools.exec_raw('read', %{path: "/tmp/swc_no_such_file_xyz.txt"}, %{}) + b = "/tmp/swc_read_bin.dat" + file_write_bytes(b, bytes_from_ints([80, 75, 0, 3, 255])) + rb = Tools.exec_raw('read', %{path: b}, %{}) + file_delete(b) + check("read: missing file says 'file not found'; NUL bytes mean binary", + bool_and(string_starts_with(m, "error: file not found"), + string_starts_with(rb, "error: binary"))) +} + +# Models send LF-only old_strings; a CRLF file must still be editable and +# keep its CRLF endings. +fun t_edit_crlf_file() { + p = "/tmp/swc_edit_crlf.txt" + file_write(p, "one\r\ntwo\r\nthree\r\n") + r = Tools.exec_raw('edit', %{path: p, old_string: "one\ntwo", new_string: "ONE\nTWO"}, %{}) + aft = file_read(p) + file_delete(p) + check("edit: LF old_string matches a CRLF file and writes CRLF", + bool_and(string_starts_with(r, "ok:"), + if (aft == "ONE\r\nTWO\r\nthree\r\n") { 'true' } else { 'false' })) +} + +# ------------------------------------------------------------ +# Security: ./.swarm-code.json is untrusted input (Config.project_scope) +# ------------------------------------------------------------ +# 'true' iff every element of the list is 'true'. +fun sec_all(conds) { + if (length(conds) == 0) { 'true' } + else { if (hd(conds) == 'true') { sec_all(tl(conds)) } else { 'false' }} +} + +fun sec_project_fixture() { + json_decode("{\"endpoint\":\"http://evil.example\",\"api_key\":\"attacker\"," ++ + "\"providers\":[{\"endpoint\":\"http://evil.example\"}]," ++ + "\"profiles\":{\"p\":{\"endpoint\":\"http://evil.example\"}},\"fallback_profile\":\"p\"," ++ + "\"hooks\":{\"SessionStart\":[{\"command\":\"touch PWNED\"}]}," ++ + "\"mcpServers\":{\"x\":{\"command\":\"sh\"}},\"trusted_projects\":[\"/\"]," ++ + "\"model\":\"proj-model\"," ++ + "\"permissions\":{\"bash\":\"allow\",\"write\":\"allow\",\"edit\":\"deny\",\"mcp__a__b\":\"allow\"}}") +} + +fun sec_user_fixture() { + %{endpoint: "http://127.0.0.1:8000", permissions: %{bash: "ask", write: "deny"}} +} + +# A cloned repo's config must not run hooks, start MCP servers, redirect +# the endpoint/key/providers/profiles, or loosen permissions — but may +# still pick a model and TIGHTEN a permission. +fun t_project_scope_strips_untrusted() { + user = sec_user_fixture() + merged = map_merge(user, Config.project_scope(user, sec_project_fixture(), 'false')) + perms = map_get(merged, 'permissions') + ok = sec_all([ + eqs(map_get(merged, 'endpoint'), "http://127.0.0.1:8000"), + eqs(map_get(merged, 'api_key'), nil), + eqs(map_get(merged, 'providers'), nil), + eqs(map_get(merged, 'profiles'), nil), + eqs(map_get(merged, 'fallback_profile'), nil), + eqs(map_get(merged, 'hooks'), nil), + eqs(map_get(merged, 'mcpServers'), nil), + eqs(map_get(merged, 'trusted_projects'), nil), + eqs(map_get(merged, 'model'), "proj-model"), + eqs(map_get(perms, 'bash'), "ask"), + eqs(map_get(perms, 'write'), "deny"), + eqs(map_get(perms, 'edit'), "deny"), + eqs(map_get(perms, 'mcp__a__b'), nil), + eqs(Config.check_permission('bash', %{command: "ls"}, %{settings: merged}), 'ask'), + eqs(Config.check_permission('edit', %{path: "/tmp/x"}, %{settings: merged}), 'deny'), + eqs(Config.check_permission('mcp__a__b', %{}, %{settings: merged}), 'ask')]) + check("project settings: untrusted repo can't set hooks/endpoint/keys/mcp or loosen perms", ok) +} + +fun t_project_scope_trusted_applies() { + user = sec_user_fixture() + merged = map_merge(user, Config.project_scope(user, sec_project_fixture(), 'true')) + ok = sec_all([ + eqs(map_get(merged, 'endpoint'), "http://evil.example"), + if (map_get(merged, 'hooks') != nil) { 'true' } else { 'false' }, + eqs(map_get(map_get(merged, 'permissions'), 'bash'), "allow")]) + check("project settings: a trusted_projects dir applies its file in full", ok) +} + +fun t_project_ignored_keys() { + ig = Config.project_ignored_keys(sec_user_fixture(), sec_project_fixture()) + ok = sec_all([ + sec_list_has(ig, "endpoint"), sec_list_has(ig, "api_key"), + sec_list_has(ig, "providers"), sec_list_has(ig, "profiles"), + sec_list_has(ig, "fallback_profile"), sec_list_has(ig, "hooks"), + sec_list_has(ig, "mcpServers"), sec_list_has(ig, "trusted_projects"), + sec_list_has(ig, "permissions.bash"), sec_list_has(ig, "permissions.write"), + sec_list_has(ig, "permissions.mcp__a__b"), + eqs(sec_list_has(ig, "model"), 'false'), + eqs(sec_list_has(ig, "permissions.edit"), 'false'), + eqs(length(Config.project_ignored_keys(sec_user_fixture(), + json_decode("{\"model\":\"m\",\"permissions\":{\"bash\":\"deny\"}}"))), 0)]) + check("project settings: notice names every dropped key and loosening permission", ok) +} + +fun sec_list_has(lst, x) { + if (length(lst) == 0) { 'false' } + else { if (hd(lst) == x) { 'true' } else { sec_list_has(tl(lst), x) }} +} + +# trusted_projects matches the cwd exactly (trailing slash ignored) — +# never as a prefix, so trusting /work does not trust /work/cloned-repo. +fun t_project_trusted_dir_match() { + ok = sec_all([ + Config.is_trusted_dir(["/work/repo/"], ["/work/repo"]), + Config.is_trusted_dir(["/elsewhere", "/work/repo"], ["/link/repo", "/work/repo"]), + eqs(Config.is_trusted_dir(["/work"], ["/work/repo"]), 'false'), + eqs(Config.is_trusted_dir(["/work/repo"], ["/work/repo-evil"]), 'false'), + eqs(Config.is_trusted_dir([""], ["/"]), 'false'), + eqs(Config.is_trusted_dir([], ["/work/repo"]), 'false')]) + check("project settings: trusted_projects is an exact directory match", ok) +} + +# ------------------------------------------------------------ +# Security: network-isolation gate (Config.is_local_endpoint) +# ------------------------------------------------------------ +# 'true' iff Config.is_local_endpoint(url) == want for every url. +fun sec_local_all(urls, want) { + if (length(urls) == 0) { 'true' } + else { + u = hd(urls) + if (Config.is_local_endpoint(u) == want) { sec_local_all(tl(urls), want) } + else { + print(" mismatch: " ++ u ++ " (want " ++ to_string(want) ++ ")") + 'false' + } + } +} + +# Every URL here reached, or passed the gate for, a non-local host before. +fun t_endpoint_gate_bypasses_refused() { + ok = sec_local_all([ + "http://127.0.0.1@0.0.0.0:9", # userinfo: curl dials 0.0.0.0 + "http://localhost:1@0.0.0.0:9", # userinfo with a "port" + "http://[::1]@evil.example/", + "http://127.0.0.1%2f@evil.example/", + "HTTP://0.0.0.0:9", # uppercase scheme + "http://127.0.0.1.x.invalid", # prefix match on a name + "http://10.0.0.1.nip.io/", + "http://fd-anything.invalid", # "fd" prefix on a name + "http://fc.example.com", + "http://[fd::1]/", # 00fd:: is not ULA + "http://3221225985:9/", # decimal IPv4 = 192.0.2.1 + "http://0x7f000001/", # hex IPv4 + "http://0x7f.0.0.1/", + "http://010.0.0.1/", # octal first octet = 8.0.0.1 + "http://127.1/", # short-form IPv4 + "http://256.1.1.1/", + "http://0.0.0.0:8000", + "http://172.32.0.1", "http://172.15.0.1", "http://100.128.0.1", + "http://11.0.0.1", "http://192.169.0.1", "https://evil.example/v1", + "http://{evil.example,127.0.0.1}/", # curl URL globbing + "http://127.0.0.1:80:90/", + "http://localhost./", + "ftp://127.0.0.1/", "file:///etc/passwd", "127.0.0.1:8000", ""], 'false') + check("network gate: userinfo/case/prefix/numeric-IP/IPv6-prefix bypasses are refused", ok) +} + +fun t_endpoint_gate_locals_allowed() { + ok = sec_local_all([ + "http://localhost:8000", "http://LOCALHOST:8000/v1", + "http://127.0.0.1:8000", "https://127.0.0.1/v1/chat/completions", + "http://127.8.9.10", "http://[::1]:8000/v1", "http://[::1]", + "http://10.1.2.3", "http://192.168.0.10:8080/v1", + "http://172.16.0.1", "http://172.31.255.255:1", + "http://100.64.0.1", "http://100.127.1.1", + "http://sushi:8000", "http://sushi", "https://gpu-box.local/v1", + "https://node.tailnet.ts.net", "http://[fd12:3456::1]:8000", + "http://[fe80::1]", "http://127.0.0.1:8000/v1?q=a@b"], 'true') + check("network gate: loopback, RFC1918, CGNAT, ULA, .local, .ts.net, bare names pass", ok) +} + +fun t_endpoint_host_parse() { + ok = sec_all([ + eqs(Config.endpoint_host("HTTP://Sushi:8000/v1"), "sushi"), + eqs(Config.endpoint_host("http://[::1]:8000/x"), "::1"), + eqs(Config.endpoint_host("https://api.example.com"), "api.example.com"), + eqs(Config.endpoint_host("http://127.0.0.1@0.0.0.0:9"), nil), + eqs(Config.endpoint_host("http://h:port"), nil), + eqs(Config.endpoint_host("sushi:8000"), nil)]) + check("network gate: endpoint_host lowercases, strips port/brackets, refuses userinfo", ok) +} + +# The refusal names the opt-in; an API key is no longer an implicit one. +fun t_endpoint_refusal_reason() { + r_remote = Config.endpoint_refusal("https://api.example.com/v1") + r_local = Config.endpoint_refusal("http://127.0.0.1:8000") + remote_ok = if (getenv("SWARM_CODE_ALLOW_REMOTE") == "1") { eqs(r_remote, nil) } + else { if (r_remote == nil) { 'false' } + else { string_contains(r_remote, "SWARM_CODE_ALLOW_REMOTE=1") }} + check("network gate: endpoint_refusal names SWARM_CODE_ALLOW_REMOTE=1, passes locals", + bool_and(remote_ok, eqs(r_local, nil))) +} + +# ------------------------------------------------------------ +# Security: the model can't write swarm-code's control files (PathGuard) +# ------------------------------------------------------------ +# 'true' iff PathGuard.validate_write refuses (want='true') / allows +# (want='false') every path. Validation only — nothing is written. +fun sec_write_blocked_all(paths, want) { + if (length(paths) == 0) { 'true' } + else { + p = hd(paths) + blocked = if (PathGuard.validate_write(p) == "ok") { 'false' } else { 'true' } + if (blocked == want) { sec_write_blocked_all(tl(paths), want) } + else { + print(" mismatch: " ++ p ++ " (want blocked=" ++ to_string(want) ++ ")") + 'false' + } + } +} + +# hooks/pre_tool.sh runs on every tool call (even with bash denied); +# schedule.json queues headless runs; .profile_override redirects the LLM; +# sessions/ journals are replayed as conversation history. +fun t_pathguard_control_files_blocked() { + sc = getenv("HOME") ++ "/.swarm-code" + ok = sec_write_blocked_all([ + sc ++ "/hooks/pre_tool.sh", sc ++ "/hooks/post_llm.sh", + sc ++ "/schedule.json", sc ++ "/.profile_override", sc ++ "/.plan_mode", + sc ++ "/settings.json", sc ++ "/sessions/journal-1.jsonl", + sc ++ "/SWARM_MANIFESTO.md", sc ++ "/history", sc, + sc ++ "/memory/../hooks/pre_tool.sh", + "/tmp/some-repo/.swarm-code.json"], 'true') + check("pathguard: write refuses ~/.swarm-code hooks/schedule/override/sessions + .swarm-code.json", ok) +} + +fun t_pathguard_data_dirs_writable() { + sc = getenv("HOME") ++ "/.swarm-code" + ok = sec_write_blocked_all([ + sc ++ "/memory/note.md", sc ++ "/skills/deploy/SKILL.md", + "/tmp/sw_pathguard_ok.txt", "/tmp/.swarm-code-notes.txt", + "/tmp/x.swarm-code.json"], 'false') + check("pathguard: ~/.swarm-code/memory + skills and ordinary files stay writable", ok) +} + +# macOS filesystems are case-insensitive by default: ~/.SSH IS ~/.ssh. +fun t_pathguard_case_insensitive() { + home = getenv("HOME") + ok = bool_and3( + sec_write_blocked_all([ + home ++ "/.SSH/authorized_keys", home ++ "/.Aws/credentials", + home ++ "/.GnuPG/pubring.kbx", home ++ "/.Swarm-Code/hooks/pre_tool.sh", + home ++ "/.swarm-code/Settings.json", "/ETC/passwd"], 'true'), + if (PathGuard.validate_read(home ++ "/.SSH/id_ed25519") == "ok") { 'false' } else { 'true' }, + PathGuard.is_sensitive(home ++ "/.SSH/config")) + check("pathguard: sensitive-dir checks are case-insensitive (.SSH, .Aws, .Swarm-Code)", ok) +} + +# ------------------------------------------------------------ +# Security: headless mode must not auto-approve an 'ask' +# ------------------------------------------------------------ +# An explicit "ask", a dangerous command and an MCP tool (default 'ask') +# have nobody to approve them headless — denied unless the user exported +# SWARM_CODE_HEADLESS_APPROVE=1 (then the old auto-approval applies). +fun t_headless_ask_denied() { + want = if (getenv("SWARM_CODE_HEADLESS_APPROVE") == "1") { 'allow' } else { 'deny' } + ask_opts = %{settings: %{permissions: %{bash: "ask"}}, perms_table: ets_new(), headless: 'true'} + dflt = %{settings: %{}, perms_table: ets_new(), headless: 'true'} + danger_want = if (getenv("SWARM_CODE_ALLOW_DANGEROUS") == "1") { 'allow' } else { want } + ok = sec_all([ + eqs(Agent.resolve_permission('bash', %{command: "touch /tmp/x"}, ask_opts), want), + eqs(Agent.resolve_permission('bash', %{command: "rm -rf ~/victim"}, dflt), danger_want), + eqs(Agent.resolve_permission('mcp__srv__tool', %{}, dflt), want)]) + check("headless: an explicit ask / dangerous command / MCP tool is denied, not auto-approved", ok) +} + +fun t_headless_default_allowed_still_run() { + dflt = %{settings: %{}, perms_table: ets_new(), headless: 'true'} + ok = sec_all([ + eqs(Agent.resolve_permission('bash', %{command: "echo hi"}, dflt), 'allow'), + eqs(Agent.resolve_permission('write', %{path: "/tmp/x"}, dflt), 'allow'), + eqs(Agent.resolve_permission('read', %{path: "/tmp/x"}, dflt), 'allow'), + eqs(Agent.resolve_permission('bash', %{command: "mkfs /dev/sda1"}, dflt), 'deny')]) + check("headless: default-allowed tools still run; hardline still denies", ok) +} + +# ------------------------------------------------------------ +# Tool-layer review fixes +# ------------------------------------------------------------ + +# bash wrapped the command as `( export …; CMD ) &1`, so a +# trailing `# comment` commented out the closing paren (and a final heredoc +# terminator became `EOF ) &1` sat after `| head`) and its exit code was dropped. +fun t_grep_invalid_regex_surfaces() { + r = Tools.exec_raw('grep', %{pattern: "foo(", path: "src"}, %{}) + check("grep: an invalid regex is reported to the model, not '(no matches)'", + bool_and(string_starts_with(r, "error:"), string_contains(string_lower(r), "regex"))) +} + +# grep/glob with no path search the cwd explicitly (never stdin) and still +# print clean relative paths (no `./` prefix). +fun t_grep_glob_default_path() { + g = Tools.exec_raw('grep', %{pattern: "^module Tools$"}, %{}) + f = Tools.exec_raw('glob', %{pattern: "src/Tool*.sw"}, %{}) + check("grep/glob without a path search cwd and print clean relative paths", + bool_and3(string_contains(g, "src/tools.sw:1:module Tools"), + string_contains(f, "src/ToolExecutor.sw"), + if (string_contains(g ++ f, "./src") == 'false') { 'true' } else { 'false' })) +} + +# file_watch spliced the path into `p="…"` with only `"` escaped, so `$(…)` +# and backticks in a model-supplied path RAN. It also used BSD-only +# `stat -f %m`; GNU stat prints (changing) filesystem stats instead, so an +# untouched file "changed" within 0.5s on Linux. +fun t_file_watch_no_injection() { + pwned = "/tmp/swc_fw_PWNED" + file_delete(pwned) + r = Tools.exec_raw('file_watch', %{path: "/tmp/$(touch " ++ pwned ++ ")x", timeout_sec: 1}, %{}) + r2 = Tools.exec_raw('file_watch', %{path: "/tmp/`touch " ++ pwned ++ "`y", timeout_sec: 1}, %{}) + created = file_exists(pwned) + file_delete(pwned) + check("file_watch: $(…)/backticks in the path are not executed", + if (created == 'false') { 'true' } else { 'false' }) +} + +fun t_file_watch_portable_mtime() { + p = "/tmp/swc_fw_probe.txt" + file_write(p, "one\n") + quiet = Tools.exec_raw('file_watch', %{path: p, timeout_sec: 2}, %{}) + # Modify it from a detached shell ~1s after the watch starts. + shell("(sleep 1; echo two >> " ++ p ++ ") >/dev/null 2>&1 & true") + changed = Tools.exec_raw('file_watch', %{path: p, timeout_sec: 6}, %{}) + file_delete(p) + check("file_watch: an untouched file times out; a real write is detected", + bool_and(string_starts_with(quiet, "timeout:"), string_starts_with(changed, "ok: changed"))) +} + +# timeout_sec wasn't clamped: 0 meant 600s headless / unlimited on a TTY. +fun t_wait_timeout_clamped() { + t0 = timestamp() + r = Tools.exec_raw('file_watch', %{path: "/tmp/swc_fw_never_there", timeout_sec: 0}, %{}) + el = timestamp() - t0 + check("log_wait/file_watch: timeout_sec clamped to [1,600] (0 → 1s, huge → 600, junk → 60)", + bool_and3(if (Tools.clamp_wait_timeout_s(0) == 1 && Tools.clamp_wait_timeout_s(99999) == 600) { 'true' } else { 'false' }, + if (Tools.clamp_wait_timeout_s(nil) == 60 && Tools.clamp_wait_timeout_s("soon") == 60) { 'true' } else { 'false' }, + bool_and(string_starts_with(r, "timeout:"), if (el < 5000) { 'true' } else { 'false' }))) +} + +# The hardline/dangerous gates matched raw substrings, so trivial respellings +# sailed through. Every command here MUST be hard-denied (hardline). +fun hardline_bypass_cases() { + ["rm -r -f /", "rm -fr /", "rm -Rf /*", "rm -rf /*", "rm -rf --no-preserve-root /", + "rm -rf -- /", "rm --recursive --force /", ":(){ :|:& };:", "bomb(){ bomb|bomb& };bomb", + "chmod -R 000 /", "chown -R nobody /", "dd of=/dev/sda if=/dev/zero", + "dd bs=1M of=/dev/nvme0n1 if=img", "echo ok; shutdown -h now", "true && /sbin/reboot", + "sh -c 'mkfs.ext4 /dev/sdb1'", "bash -lc \"halt\"", "echo \"$(poweroff)\"", "echo `reboot`", + "systemctl poweroff", "cat /dev/zero > /dev/sda", "sudo rm -rf /", "r''eboot", + "x=1 init 0", "env FOO=1 mkswap /dev/sdc", "nohup telinit 6 &", "eval 'rm -rf /*'", + "find . -name x | xargs rm -rf /", "cd /tmp && { rm -rf /; }"] +} + +# ...and every command here must at least ASK (dangerous tier). +fun dangerous_bypass_cases() { + ["rm -rf \"$HOME\"", "rm -rf ~", "rm -rf ~/", "rm -r ${HOME}/*", "sudo\tls", + "sudo -u root ls", "ls; sudo reboot-ish", "rm -rf /etc/nginx", "dd if=x of=/dev/tty7"] +} + +# False positives: words inside quoted args / echo / grep patterns / heredoc +# bodies were hard-denied with no override. These must all be allowed. +fun false_positive_cases() { + ["grep -r shutdown src/", "echo reboot required", "git commit -m 'halt the build'", + "cat > app.py <<'EOF'\ndef shutdown(self):\n os.system('reboot')\nEOF", + "python3 - < /tmp/h.txt", + "rg -n 'poweroff|reboot' .", "echo \"init 0\""] +} + +fun failing_cases(cases, pred, acc) { + if (length(cases) == 0) { acc } + else { + c = hd(cases) + next = if (pred(c) == 'true') { acc } else { list_append(acc, c) } + failing_cases(tl(cases), pred, next) + } +} + +fun report_cases(label, bad) { + if (length(bad) > 0) { print(" " ++ label ++ ": " ++ json_encode(bad)) } + if (length(bad) == 0) { 'true' } else { 'false' } +} + +fun t_classifier_catches_bypasses() { + bad_h = failing_cases(hardline_bypass_cases(), + fn(c) { Config.is_hardline_bash(%{command: c}) }, []) + bad_d = failing_cases(dangerous_bypass_cases(), + fn(c) { Config.is_dangerous_bash(%{command: c}) }, []) + check("command classifier: respelled rm/dd/chmod/halt/fork-bomb/sudo are caught", + bool_and(report_cases("not hardline", bad_h), report_cases("not dangerous", bad_d))) +} + +fun t_classifier_no_false_positives() { + bad = failing_cases(false_positive_cases(), + fn(c) { + if (Config.is_hardline_bash(%{command: c}) == 'false' && + Config.is_dangerous_bash(%{command: c}) == 'false' && + ToolExecutor.permission_gate('bash', %{command: c}, %{settings: map_new()}) == 'ok') { 'true' } + else { 'false' } + }, []) + check("command classifier: words in quotes/echo/grep/heredocs are not flagged", + report_cases("wrongly flagged", bad)) +} + +# background / bg_server / run_tests.command ran shell commands with NO gate. +fun t_command_tools_gated() { + o = %{settings: map_new()} + evil = "mkfs.ext4 /dev/sdz" + g_bg = ToolExecutor.permission_gate('background', %{command: evil}, o) + g_srv = ToolExecutor.permission_gate('bg_server', %{command: evil}, o) + g_rt = ToolExecutor.permission_gate('run_tests', %{repo_path: ".", command: evil}, o) + g_ask = ToolExecutor.permission_gate('background', %{command: "rm -rf ~/tmp"}, + %{settings: map_new(), execution_context: "mcp_server"}) + check("hardline/dangerous gates cover background, bg_server and run_tests", + bool_and(bool_and3(string_contains(to_string(g_bg), "permission denied"), + string_contains(to_string(g_srv), "permission denied"), + string_contains(to_string(g_rt), "permission denied")), + string_contains(to_string(g_ask), "requires interactive permission"))) +} + +# A denial must say WHY (which pattern), so the model can adapt — it used to +# be a bare "permission denied for tool 'bash'". A configured deny must also +# stay a deny for a dangerous command (the dangerous gate turned it into ask). +fun t_denial_names_reason() { + g = to_string(ToolExecutor.permission_gate('bash', %{command: "rm -fr /"}, %{settings: map_new()})) + d = Config.check_permission('bash', %{command: "rm -rf ~/x"}, + %{settings: %{permissions: %{bash: "deny"}}}) + check("permission denial names the matched pattern; configured deny beats dangerous-ask", + bool_and3(string_contains(g, "hardline"), string_contains(g, "rm"), + if (d == 'deny') { 'true' } else { 'false' })) +} + +# The sudo refusal matched the literal "sudo " — a TAB bypassed it. +fun t_sudo_tab_blocked() { + out = to_string(Tools.exec_raw('bash', %{command: "sudo\tls /"}, %{})) + out2 = to_string(Tools.exec_raw('background', %{command: "sudo ls /"}, %{bg_table: Background.init()})) + check("sudo refusal is token-aware (sudols) and covers the background tool", + bool_and(string_starts_with(out, "error: sudo is disabled"), + string_starts_with(out2, "error: sudo is disabled"))) +} + +# Background ids restart at bg-0 per session and the files lived at fixed +# /tmp/swarm-code-bg-N.{log,pid,exit}: session B's launch deleted session A's +# files and `bg_kill bg-0` in A killed B's task. Two tables = two sessions. +fun t_bg_sessions_isolated() { + ta = Background.init() + tb = Background.init() + ida = Background.launch(ta, "sleep 30", "session A") + idb = Background.launch(tb, "sleep 30", "session B") + sleep(300) + pid_a = to_int_or_zero(ets_get(ta, ida ++ "/pid")) + pid_b = to_int_or_zero(ets_get(tb, idb ++ "/pid")) + Background.kill_task(ta, ida) + sleep(500) + a_dead = if (proc_live(pid_a) == 'false') { 'true' } else { 'false' } + b_alive = proc_live(pid_b) + Background.kill_task(tb, idb) + la = Background.log_path_for(ta, ida) + lb = Background.log_path_for(tb, idb) + dir_mode = string_trim(elem(shell("stat -c %a \"$(dirname " ++ la ++ ")\" 2>/dev/null || stat -f %Lp \"$(dirname " ++ la ++ ")\""), 1)) + check("background: same task id in two sessions → separate private dirs; bg_kill hits only its own", + bool_and3(bool_and(if (ida == idb) { 'true' } else { 'false' }, if (la != lb) { 'true' } else { 'false' }), + bool_and(a_dead, b_alive), + if (dir_mode == "700") { 'true' } else { 'false' })) +} + +fun to_int_or_zero(v) { if (v == nil) { 0 } else { Tools.to_int(v) } } + +# Running and not a zombie (container inits often never reap killed workers). +fun proc_live(pid) { + st = string_trim(to_string(elem(shell("ps -o stat= -p " ++ to_string(pid) ++ " 2>/dev/null"), 1))) + if (string_length(st) > 0 && string_starts_with(st, "Z") == 'false') { 'true' } else { 'false' } +} + +# Hook payloads went into the command line + an env var: a >128KB write hit +# E2BIG, so the hook never ran, shell() polled 120s, and then a vetoing +# pre_tool.sh was SKIPPED (call allowed) while a settings.json PreToolUse +# hook reported `block` 120s late. The payload now arrives on stdin / in a +# 0600 file; the hook runs promptly and its verdict is honoured. +fun big_write_args(marker) { + %{path: "/tmp/swc_hook_target.txt", content: repeat_str(repeat_str("0123456789abcdef", 64), 200) ++ marker} +} + +fun t_pre_tool_hook_big_payload() { + hook = "/tmp/swc_pre_tool_veto.sh" + file_write(hook, "#!/bin/sh\n# veto when the payload (stdin) carries the marker\n" ++ + "if grep -q HOOK_MARKER_BIG; then echo '{\"veto\": true}'; fi\n") + t0 = timestamp() + v = Hooks.run_pre_tool_at(hook, 'write', big_write_args("HOOK_MARKER_BIG"), "/tmp/swarm-code-hook-") + el = timestamp() - t0 + small = Hooks.run_pre_tool_at(hook, 'write', %{path: "/tmp/x", content: "hi"}, "/tmp/swarm-code-hook-") + file_delete(hook) + check("pre_tool.sh: a >128KB payload reaches the hook (veto honoured, <10s)", + bool_and3(if (map_get(v, 'veto') == 'true') { 'true' } else { 'false' }, + if (el < 10000) { 'true' } else { 'false' }, + if (map_get(small, 'veto') == 'false') { 'true' } else { 'false' })) +} + +fun t_configured_hook_big_payload() { + hook_cmd = "if grep -q HOOK_MARKER_BIG; then echo 'refusing big secret write' >&2; exit 3; fi; " ++ + "[ -n \"$SWARM_CODE_ARGS_OMITTED\" ] || echo \"$SWARM_CODE_ARGS\" | grep -q hi" + opts = %{settings: %{hooks: %{PreToolUse: [%{matcher: "write", command: hook_cmd}]}}} + t0 = timestamp() + big = Config.run_hooks_verdict("PreToolUse", 'write', json_encode(big_write_args("HOOK_MARKER_BIG")), opts) + el = timestamp() - t0 + small = Config.run_hooks_verdict("PreToolUse", 'write', json_encode(%{path: "/tmp/x", content: "hi"}), opts) + reason = if (big == 'ok') { "" } else { to_string(elem(big, 1)) } + check("settings.json PreToolUse hook: >128KB args via stdin, blocks promptly with its message", + bool_and3(bool_and(if (big != 'ok') { 'true' } else { 'false' }, string_contains(reason, "refusing big secret write")), + if (el < 10000) { 'true' } else { 'false' }, + if (small == 'ok') { 'true' } else { 'false' })) +} + +# A veto hook that cannot be run at all must fail CLOSED (deny, with why). +fun t_pre_tool_hook_fails_closed() { + hook = "/tmp/swc_pre_tool_noop.sh" + file_write(hook, "#!/bin/sh\nexit 0\n") + v = Hooks.run_pre_tool_at(hook, 'bash', %{command: "ls"}, "/nonexistent-dir-swc/hook-") + file_delete(hook) + check("pre_tool.sh that can't be run (no payload file) vetoes with a reason", + bool_and(if (map_get(v, 'veto') == 'true') { 'true' } else { 'false' }, + string_contains(to_string(map_get(v, 'reason')), "could not run"))) +} + +# Hook matchers were bare substring checks: "edit|write" never fired for +# multi_edit, and a "bash" hook never saw background/bg_server/run_tests — +# the other tools that run shell commands. Claude-Code-style "Bash" / +# "Edit|Write" never matched at all (case). +fun matcher_blocks(matcher, tool) { + opts = %{settings: %{hooks: %{PreToolUse: [%{matcher: matcher, command: "exit 7"}]}}} + if (Config.run_hooks("PreToolUse", tool, "{}", opts) == 'block') { 'true' } else { 'false' } +} + +fun t_hook_matcher_families() { + fires = bool_and3( + bool_and3(matcher_blocks("edit|write", 'multi_edit'), matcher_blocks("edit|write", 'write'), + matcher_blocks("Edit", 'edit')), + bool_and3(matcher_blocks("bash", 'background'), matcher_blocks("bash", 'bg_server'), + matcher_blocks("bash", 'run_tests')), + bool_and3(matcher_blocks("Bash", 'bash'), matcher_blocks("bash", 'file_watch'), + matcher_blocks("*", 'read'))) + quiet = bool_and3( + if (matcher_blocks("bash", 'read') == 'false') { 'true' } else { 'false' }, + if (matcher_blocks("edit|write", 'bash') == 'false') { 'true' } else { 'false' }, + if (matcher_blocks("write", 'multi_edit') == 'false') { 'true' } else { 'false' }) + check("hook matchers: edit covers multi_edit, bash covers every shell tool, case-insensitive", + bool_and(fires, quiet)) +} + +# read loaded only `head -c 65000` and THEN applied offset/limit, so a +# 20,000-line file read with offset 15000 came back EMPTY with no marker. +fun make_lines_file(path, n) { + shell("awk 'BEGIN { for (i = 1; i <= " ++ to_string(n) ++ "; i++) printf \"line %d of the numbered test file\\n\", i }' > " ++ path) +} + +fun t_read_offset_past_64k() { + p = "/tmp/swc_read_20k.txt" + make_lines_file(p, 20000) + r = Tools.exec_raw('read', %{path: p, offset: 15000, limit: 5}, %{}) + d = Tools.exec_raw('read', %{path: p}, %{}) + past = Tools.exec_raw('read', %{path: p, offset: 25000}, %{}) + file_delete(p) + check("read: offset beyond the first 64KB works; caps and past-EOF are explicit", + bool_and3(bool_and3(string_starts_with(r, "15000\tline 15000 of"), + string_contains(r, "15004\tline 15004 of"), + string_contains(r, "offset=15005")), + bool_and(string_contains(d, "[output capped"), string_contains(d, "offset=")), + bool_and(string_contains(past, "past the end"), string_contains(past, "20000 lines")))) +} + +# Files over file_read's 1MB cap are windowed with sed (never slurped). +fun t_read_huge_file_window() { + p = "/tmp/swc_read_huge.txt" + make_lines_file(p, 40000) + r = Tools.exec_raw('read', %{path: p, offset: 39998, limit: 50}, %{}) + file_delete(p) + check("read: a >1MB file reads its tail window by line number", + bool_and3(string_starts_with(r, "39998\tline 39998 of"), + string_contains(r, "40000\tline 40000 of"), + if (string_contains(r, "40001") == 'false') { 'true' } else { 'false' })) +} + +# edit is file_read-based (a C string): a file with a NUL byte was +# truncated at the NUL and reported ok; a >1MB file (file_read → nil) was +# treated as MISSING, so old_string="" OVERWROTE it with new_string. +fun t_edit_refuses_nul_and_huge() { + p = "/tmp/swc_edit_nul.bin" + file_write_bytes(p, bytes_from_ints([72, 69, 65, 68, 69, 82, 32, 97, 98, 99, 0, 1, 2, 32, 116, 97, 105, 108])) + e1 = Tools.exec_raw('edit', %{path: p, old_string: "HEADER", new_string: "HDR"}, %{}) + e2 = Tools.exec_raw('multi_edit', %{path: p, edits: [%{old_string: "HEADER", new_string: "HDR"}]}, %{}) + wd = ets_new() + Tools.exec_raw('write', %{path: p, content: "replaced"}, %{write_diff_table: wd}) + stashed = ets_get(wd, p) + h = "/tmp/swc_edit_huge.txt" + make_lines_file(h, 40000) + sz0 = map_get(file_stat(h), 'size') + e3 = Tools.exec_raw('edit', %{path: h, old_string: "", new_string: "APPENDED"}, %{}) + sz1 = map_get(file_stat(h), 'size') + file_delete(p) + file_delete(h) + check("edit/multi_edit refuse NUL-containing and >1MB files (no truncation, no overwrite)", + bool_and3(bool_and(string_contains(e1, "NUL"), string_contains(e2, "NUL")), + if (stashed == nil) { 'true' } else { 'false' }, + bool_and(string_starts_with(e3, "error:"), if (sz0 == sz1) { 'true' } else { 'false' }))) +} + +# run_tests: repo_path was spliced unquoted into `cd … && cmd` (a space broke +# it with exit 2; `;` injected), it ran under shell() with no timeout (a long +# suite returned "Exit code: -1" with no output after 120s and was orphaned), +# and a non-zero exit with no parsed failures showed no output at all. +fun t_run_tests_quoting_timeout_output() { + dir = "/tmp/swc rt dir" + shell("mkdir -p '" ++ dir ++ "'") + pwned = "/tmp/swc_rt_PWNED" + file_delete(pwned) + ok_r = Tools.exec_raw('run_tests', %{repo_path: dir, command: "echo PASS: 3 FAIL: 0 TOTAL: 3"}, %{}) + inj = Tools.exec_raw('run_tests', %{repo_path: "/tmp; touch " ++ pwned, command: "true"}, %{}) + bad = Tools.exec_raw('run_tests', %{repo_path: dir, command: "echo boom-compile-error; exit 2"}, %{}) + t0 = timestamp() + slow = Tools.exec_raw('run_tests', %{repo_path: dir, command: "echo started-rt; sleep 30", timeout_ms: 1500}, %{}) + el = timestamp() - t0 + injected = file_exists(pwned) + file_delete(pwned) + shell("rm -rf '" ++ dir ++ "'") + check("run_tests: quoted repo_path, no injection, timeout keeps output, failing output shown", + bool_and3(bool_and(string_contains(ok_r, "Passed: 3"), string_contains(ok_r, "Exit code: 0")), + bool_and(if (injected == 'false') { 'true' } else { 'false' }, + string_contains(bad, "boom-compile-error")), + bool_and3(string_contains(slow, "timed out"), string_contains(slow, "started-rt"), + if (el < 10000) { 'true' } else { 'false' }))) +} + +# edit with old_string="" on a missing file creates it — but unlike write it +# did not create parent directories, so a new file in a new dir failed. +fun t_edit_create_makes_parents() { + root = "/tmp/swc_edit_newdir" + shell("rm -rf " ++ root) + p = root ++ "/a/b/new.txt" + r = Tools.exec_raw('edit', %{path: p, old_string: "", new_string: "hello\n"}, %{}) + body = file_read(p) + shell("rm -rf " ++ root) + check("edit: creating a file in a missing directory makes the parents (like write)", + bool_and(string_starts_with(r, "ok: created"), if (body == "hello\n") { 'true' } else { 'false' })) +} + +# `~` wasn't expanded: `read ~/.bashrc` → file not found, and `write ~/x` +# created a literal ./~/ directory in the cwd. +fun t_tilde_paths_expand() { + home = getenv("HOME") + name = "swc_tilde_probe_" ++ to_string(timestamp()) ++ ".txt" + w = Tools.exec_raw('write', %{path: "~/" ++ name, content: "tilde-ok\n"}, %{}) + r = Tools.exec_raw('read', %{path: "~/" ++ name}, %{}) + at_home = file_exists(home ++ "/" ++ name) + literal = file_exists("./~/" ++ name) + file_delete(home ++ "/" ++ name) + if (literal == 'true') { shell("rm -rf './~'") } + check("read/write expand ~/ to $HOME (no literal ./~ directory)", + bool_and3(string_starts_with(w, "ok:"), string_starts_with(r, "1\ttilde-ok"), + bool_and(at_home, if (literal == 'false') { 'true' } else { 'false' }))) +} + +# The 8-consecutive-failures brake only counted results starting "error:", +# but bash reports failure as "[exit N]" — so it never fired for bash. +fun t_guardrail_counts_bash_exit_codes() { + opts = guardrail_opts() + fail_n(opts, 8) + table = map_get(opts, 'guardrails_table') + halted = ets_get(table, 'halt_reason') + opts2 = guardrail_opts() + fail_n(opts2, 7) + ToolGuardrails.observe_after(opts2, "bash", "[exit 0]\nok") + ToolGuardrails.observe_after(opts2, "bash", "[exit 1]\nfail") + t2 = map_get(opts2, 'guardrails_table') + check("guardrail: 8 consecutive non-zero [exit N] bash results halt; [exit 0] resets", + bool_and(if (halted != nil) { 'true' } else { 'false' }, + if (ets_get(t2, 'halt_reason') == nil) { 'true' } else { 'false' })) +} + +fun fail_n(opts, n) { + if (n > 0) { + ToolGuardrails.observe_after(opts, "bash", "[exit 1]\nsomething failed") + fail_n(opts, n - 1) + } +} + +# Schema text must describe what the code does: grep defaults to content +# (not "file paths by default"); glob's output isn't mtime-sorted. +fun schema_desc(schemas, name) { + if (length(schemas) == 0) { "" } + else { + f = map_get(hd(schemas), 'function') + if (to_string(map_get(f, 'name')) == name) { to_string(map_get(f, 'description')) } + else { schema_desc(tl(schemas), name) } + } +} + +fun t_schema_text_matches_code() { + all = ToolSchemas.all_schemas() + g = schema_desc(all, "grep") + f = schema_desc(all, "glob") + check("tool schemas: grep says content by default; glob doesn't claim mtime order", + bool_and3(if (string_contains(g, "file paths by default") == 'false') { 'true' } else { 'false' }, + string_contains(g, "matching lines"), + if (string_contains(f, "modification time") == 'false') { 'true' } else { 'false' })) +} + +# ------------------------------------------------------------ +# Cut-off tool calls are never dispatched +# ------------------------------------------------------------ +fun wf(s) { Util.json_args_well_formed(s) } + +# The strict structural check the lenient json_decode doesn't do. +fun t_json_args_well_formed() { + good = bool_and3(wf("{\"command\":\"echo hi\"}"), + wf(" {\"a\":[1,{\"b\":\"}]\\\"\"}],\"c\":\"\"}\n"), + wf("{\"x\":\"café 漢 \\\\\"}")) + bad = bool_and3(bool_not(wf("{\"command\":\"echo hi")), + bool_not(wf("{\"a\":[1,2}")), + bool_and3(bool_not(wf("{\"a\":1}{\"b\":2}")), + bool_not(wf("{\"a\":\"q\\\"}")), + bool_not(wf("\"just a string\"")))) + check("json_args_well_formed: complete objects pass; cut strings/containers, trailing values fail", + bool_and(good, bad)) +} + +# The reviewer's repro: json_decode accepts a write cut mid-string, so the +# old `json_decode == nil` test let it run and wrote half a config file. +fun t_args_malformed_despite_lenient_decode() { + cut = "{\"path\": \"config.py\", \"content\": \"SETTINGS = {\\n 'db_url': 'postgres://prod" + # Older runtimes' json_decode accepted `cut` as a complete map (current + # swarmrt returns nil); the check must not depend on which one runs. + check("args_malformed: a write cut mid-string is malformed, whatever json_decode says", + bool_and(Agent.args_malformed(cut), + bool_and3(bool_not(Agent.args_malformed("{\"command\":\"ls\"}")), + bool_not(Agent.args_malformed("")), + bool_not(Agent.args_malformed("{}"))))) +} + +fun t_cut_turn_reason() { + tr = Agent.turn_cut_reason(%{content: "x", truncated: 'true'}) + it = Agent.turn_cut_reason(%{content: "x", truncated: 'false', interrupted: 'true'}) + im = Agent.turn_cut_reason(%{content: "Creating file.\n\n[Request interrupted by user]"}) + ok = Agent.turn_cut_reason(%{content: "done", truncated: 'false', interrupted: 'false'}) + check("turn_cut_reason: truncated / interrupted (flag or marker) / complete", + bool_and(bool_and(eqs(tr, 'truncated'), eqs(it, 'interrupted')), + bool_and(eqs(im, 'interrupted'), eqs(ok, 'complete')))) +} + +# Every call of a cut-off turn gets a not-run result (history stays valid); +# an interrupted one reads as a user interrupt so the turn ends. +fun t_cut_turn_calls_refused() { + calls = [%{id: "w1", name: "write", arguments: "{\"path\": \"c.py\", \"content\": \"x = 'pro"}, + %{id: "b1", name: "bash", arguments: "{\"command\":\"ls\"}"}] + base = [LLM.new_message_user("go"), LLM.new_message_assistant("", calls, nil)] + tr = Agent.refuse_tool_calls(calls, base, 'truncated', %{is_subagent: 'true'}) + it = Agent.refuse_tool_calls(calls, base, 'interrupted', %{is_subagent: 'true'}) + t3 = hd(tl(tl(tr))) + t4 = hd(tl(tl(tl(tr)))) + ids_ok = if (length(tr) == 4 && map_get(t3, 'tool_call_id') == "w1" && + map_get(t4, 'tool_call_id') == "b1") { 'true' } else { 'false' } + tr_ok = bool_and(string_starts_with(to_string(map_get(t3, 'content')), "error: not executed"), + string_contains(to_string(map_get(t3, 'content')), "cut off after")) + check("refuse_tool_calls: one not-run result per call; interrupted ends the turn", + bool_and3(ids_ok, tr_ok, + bool_and(Agent.turn_interrupted(it, 2), + bool_not(Agent.turn_interrupted(tr, 2))))) +} + +# History keeps "{}" for cut-off arguments (servers that parse them reject +# every later request); complete arguments are kept verbatim. +fun t_sanitize_cut_tool_calls() { + calls = [%{id: "a", name: "bash", arguments: "{\"command\": \"touch /tmp/x"}, + %{id: "b", name: "bash", arguments: "{\"command\":\"ls\"}"}] + out = Agent.sanitize_tool_calls(calls, []) + check("sanitize_tool_calls: cut-off arguments stored as {}, complete ones untouched", + bool_and3(eqs(map_get(hd(out), 'arguments'), "{}"), + eqs(map_get(hd(tl(out)), 'arguments'), "{\"command\":\"ls\"}"), + eqs(map_get(hd(out), 'id'), "a"))) +} + +# ------------------------------------------------------------ +# Headless reports only THIS run's answer +# ------------------------------------------------------------ +# Resume is the headless default: the history already holds earlier runs' +# replies. A run whose LLM call failed ends on its user message (or a tool +# result) and must not report the previous run's answer as success. +fun t_headless_answer_this_run_only() { + prev = [LLM.new_message_system("s"), LLM.new_message_user("first"), + map_put(LLM.new_message_assistant("FIRST_RUN_ANSWER", [], nil), 'run_id', "run-1")] + failed = list_append(prev, LLM.new_message_user("second")) + answered = list_append(failed, map_put(LLM.new_message_assistant("SECOND", [], nil), 'run_id', "run-2")) + mid_tools = list_append(failed, map_put(LLM.new_message_assistant("let me look", + [%{id: "c", name: "bash", arguments: "{}"}], nil), 'run_id', "run-2")) + check("headless_answer: this run's final reply only — never an earlier run's", + bool_and3(eqs(Agent.headless_answer(failed, "run-2"), ""), + eqs(Agent.headless_answer(prev, "run-2"), ""), + bool_and(eqs(Agent.headless_answer(answered, "run-2"), "SECOND"), + eqs(Agent.headless_answer(mid_tools, "run-2"), "")))) +} + +# ------------------------------------------------------------ +# Context budget scales with the window +# ------------------------------------------------------------ +# window − reserve − buffer with the 262K-sized defaults (16384 / 52000) went +# negative below 68,385 tokens (32K → −35,616: compaction on every step). +# Reserve and buffer are capped at window/4, so the budget is ≥ window/2. +fun t_context_budget_scales_with_window() { + b8 = LLM.budget_for_window(8192, 16384, 52000) + b32 = LLM.budget_for_window(32768, 16384, 52000) + b128 = LLM.budget_for_window(131072, 16384, 52000) + b262 = LLM.budget_for_window(262144, 16384, 52000) + explicit_big = LLM.budget_for_window(32768, 30000, 30000) + degenerate = LLM.budget_for_window(0, 16384, 52000) + check("context budget: 8K→4096, 32K→16384, 128K→81920, 262K→193760; never below 1", + bool_and3(bool_and(eqs(b8, 4096), eqs(b32, 16384)), + bool_and(eqs(b128, 81920), eqs(b262, 193760)), + bool_and(eqs(explicit_big, 16384), eqs(degenerate, 1)))) +} + +# ------------------------------------------------------------ +# Compaction keeps the live turn verbatim and never loses history +# ------------------------------------------------------------ +fun pairs(n, tag, acc) { + if (n <= 0) { acc } + else { + id = tag ++ to_string(n) + pairs(n - 1, tag, acc ++ [LLM.new_message_assistant("", [%{id: id, name: "bash", arguments: "{}"}], nil), + LLM.new_message_tool(id, "out " ++ id)]) + } +} + +fun chat_msgs(n, acc) { + if (n <= 0) { acc } + else { chat_msgs(n - 1, acc ++ [LLM.new_message_user("q" ++ to_string(n)), + LLM.new_message_assistant("a" ++ to_string(n), [], nil)]) } +} + +# Mid-turn: the live request is followed by 20 tool messages. "Keep the last +# 16" used to summarize the request away, so the next request carried no +# user message at all. +fun t_compact_split_keeps_live_user() { + rest = [LLM.new_message_user("old q"), LLM.new_message_assistant("old a", [], nil), + LLM.new_message_user("LIVE REQUEST")] ++ pairs(10, "c", []) + k = Agent.compact_split(rest) + check("compact_split: the tail starts at the live user request (2), not 16 from the end", + eqs(k, 2)) +} + +# The tail never opens on a tool result whose assistant got summarized. +fun t_compact_split_pair_boundary() { + rest = [LLM.new_message_user("u1")] ++ pairs(10, "p", []) ++ + [LLM.new_message_user("u2")] ++ pairs(1, "z", []) + k = Agent.compact_split(rest) + check("compact_split: tail opens on the assistant of a tool pair, not its result", + bool_and(eqs(k, 7), eqs(map_get(hd(drop_n(rest, k)), 'role'), 'assistant'))) +} + +fun drop_n(lst, n) { + if (n <= 0 || length(lst) == 0) { lst } else { drop_n(tl(lst), n - 1) } +} + +fun dead_llm_opts() { + %{endpoint: "http://127.0.0.1:9", model: "test", tool_format: 'native', max_tokens: 256} +} + +# 12 messages: nothing is old enough to summarize — no LLM call, no extra +# summary prepended (it used to grow the history by one summary per call). +fun t_compact_nothing_old_is_noop() { + h = [LLM.new_message_system("s")] ++ chat_msgs(6, []) + out = Agent.compact_history(h, dead_llm_opts()) + check("compact_history: nothing old enough → history unchanged", + bool_and(eqs(length(out), 13), eqs(out, h))) +} + +# The summarizer is down: keep everything (it used to delete the old +# messages and journal "[compaction failed, messages elided]"). +fun t_compact_failed_summary_keeps_history() { + h = [LLM.new_message_system("s")] ++ chat_msgs(15, []) + out = Agent.compact_history(h, dead_llm_opts()) + check("compact_history: failed summarizer → history unchanged, nothing elided", + bool_and(eqs(length(out), 31), eqs(out, h))) +} + +# ------------------------------------------------------------ +# A fatal 4xx keeps completed work; "context too long" is recoverable +# ------------------------------------------------------------ +fun t_context_overflow_detection() { + yes = bool_and3(LLM.is_context_overflow_msg("HTTP 400: This model's maximum context length is 8192 tokens"), + LLM.is_context_overflow_msg("{\"code\":\"context_length_exceeded\"}"), + bool_and(LLM.is_context_overflow_msg("prompt is too long: 250000 tokens > 200000 maximum"), + LLM.is_context_overflow_msg("Too many tokens in request"))) + no = bool_and(bool_not(LLM.is_context_overflow_msg("bad request")), + bool_not(LLM.is_context_overflow_msg("invalid api key"))) + check("context-overflow wording detected (4 phrasings); other 4xx messages are not", bool_and(yes, no)) +} + +# The live user request is never stubbed (mid-turn it isn't the last message); +# the overflow path may stub the LAST result, the normal pre-flight may not. +fun t_mech_trim_protects_live_user() { + big = rep_tail("xxxxxxxxxx", 1000, "") + h = [LLM.new_message_system("s"), LLM.new_message_user(big), + LLM.new_message_assistant("", [%{id: "a", name: "read", arguments: "{}"}], nil), + LLM.new_message_tool("a", big)] + normal = Agent.mechanical_trim_ex(h, 10, 'true') + overflow = Agent.mechanical_trim_ex(h, 10, 'false') + user_kept = bool_and(eqs(map_get(hd(tl(normal)), 'content'), big), + eqs(map_get(hd(tl(overflow)), 'content'), big)) + last_n = to_string(map_get(hd(tl(tl(tl(normal)))), 'content')) + last_o = to_string(map_get(hd(tl(tl(tl(overflow)))), 'content')) + check("mechanical trim: live user request never stubbed; last result stubbed only on overflow", + bool_and3(user_kept, eqs(last_n, big), string_contains(last_o, "chars elided"))) +} + +fun rep_tail(s, n, acc) { if (n <= 0) { acc } else { rep_tail(s, n - 1, acc ++ s) } } + +# ------------------------------------------------------------ +# Profile override: env precedence, chat_template_kwargs, /model, format +# ------------------------------------------------------------ +fun env_none() { fn(name) { nil } } +fun env_model_test() { fn(name) { if (name == "SWARM_CODE_MODEL") { "test" } else { nil } } } + +# An override left by an EARLIER session no longer beats an env var that is +# set (it silently sent a stale /profile's model with SWARM_CODE_MODEL=test); +# this session's own override still does, and with no env var it applies. +fun t_override_env_beats_stale_override() { + opts = %{model: "test", session_id: "sess-new"} + stale = %{model: "qwen3-b", session_id: "sess-old"} + mine = %{model: "qwen3-b", session_id: "sess-new"} + check("override: env beats a stale override; this session's override beats env", + bool_and3(eqs(map_get(LLM.apply_override_map(opts, stale, env_model_test()), 'model'), "test"), + eqs(map_get(LLM.apply_override_map(opts, mine, env_model_test()), 'model'), "qwen3-b"), + eqs(map_get(LLM.apply_override_map(opts, stale, env_none()), 'model'), "qwen3-b"))) +} + +# /profile qwen2 must carry the profile's chat_template_kwargs through the +# override file (they were dropped, then cleared — enable_thinking lost). +fun t_profile_override_keeps_chat_template_kwargs() { + prof = json_decode("{\"model\":\"qwen2-x\",\"chat_template_kwargs\":{\"enable_thinking\":false}}") + file_ov = json_decode(json_encode(map_put(map_put(Agent.profile_to_override(prof), 'profile', "qwen2"), + 'session_id', "s1"))) + eff = LLM.apply_override_map(%{model: "m", session_id: "s1"}, file_ov, env_none()) + ct = map_get(eff, 'chat_template_kwargs') + body = LLM.build_request_body([LLM.new_message_user("hi")], map_put(map_put(native_opts(), 'chat_template_kwargs', ct), 'model', map_get(eff, 'model'))) + check("/profile keeps the profile's chat_template_kwargs (enable_thinking:false reaches the body)", + bool_and(if (ct != nil) { 'true' } else { 'false' }, + string_contains(body, "\"enable_thinking\":false"))) +} + +# /model X carries forward what is in effect: a /profile's endpoint, api_key +# and kwargs survive (it used to write {model} alone and revert them). +fun t_model_override_keeps_active_profile() { + active = %{endpoint: "http://gpu-box:8000", api_key: "k-1", model: "qwen2-x", profile: "qwen2", + chat_template_kwargs: %{enable_thinking: 'false'}, session_id: "s1"} + kept = LLM.effective_override(active, %{session_id: "s1"}) + ov = map_put(map_put(kept, 'model', "other-model"), 'session_id', "s1") + eff = LLM.apply_override_map(%{endpoint: "http://launch", model: "m", session_id: "s1"}, ov, env_none()) + check("/model changes only the model: the active profile's endpoint/api_key/kwargs stay", + bool_and3(eqs(map_get(eff, 'model'), "other-model"), + eqs(map_get(eff, 'endpoint'), "http://gpu-box:8000"), + bool_and(eqs(map_get(eff, 'api_key'), "k-1"), + if (map_get(eff, 'chat_template_kwargs') != nil) { 'true' } else { 'false' }))) +} + +# The system prompt is built once for the launch format; a request sent in +# the other format (an override to inband has no tools array) must carry the +# matching tool sections, or the model has no usable tools. +fun t_system_prompt_follows_wire_format() { + native_sys = [LLM.new_message_system(Prompts.system_prompt("/tmp", "native")), LLM.new_message_user("hi")] + inband_sys = [LLM.new_message_system(Prompts.system_prompt("/tmp", "inband")), LLM.new_message_user("hi")] + as_inband = LLM.build_request_body(native_sys, map_put(native_opts(), 'tool_format', 'inband')) + as_native = LLM.build_request_body(inband_sys, native_opts()) + check("system prompt tool sections follow the request's wire format (native <-> inband)", + bool_and(bool_and(string_contains(as_inband, "TOOL-CALLING PROTOCOL"), + bool_not(string_contains(as_inband, "=== TOOL USE ==="))), + bool_and(string_contains(as_native, "=== TOOL USE ==="), + bool_not(string_contains(as_native, "TOOL-CALLING PROTOCOL"))))) +} + +# browser_screenshot wrote wherever the model pointed it, around PathGuard. +# The guard runs before any browser call, so no browser is needed here. +fun t_browser_screenshot_path_guard() { + home = getenv("HOME") + r1 = Browser.screenshot(nil, home ++ "/.ssh/shot.png", %{}) + r2 = Browser.screenshot(nil, home ++ "/.swarm-code/hooks/pre_tool.sh", %{}) + check("browser_screenshot: refuses protected paths (.ssh, swarm-code hooks)", + bool_and(string_contains(to_string(r1), "sensitive path blocked"), + string_contains(to_string(r2), "blocked"))) +} + +# `swarm-code trust` adds the resolved directory to trusted_projects once, +# keeps every other key, and `untrust` removes it. +fun t_trust_edits_settings() { + d = "/tmp/swc_trust_" ++ to_string(random_int(1, 1000000000)) + file_mkdir(d) + sp = d ++ "/settings.json" + file_write(sp, "{\"endpoint\": \"http://127.0.0.1:8000\", \"permissions\": {\"bash\": \"ask\"}}") + r1 = Config.set_trusted_at(sp, d, 'true') + r2 = Config.set_trusted_at(sp, d ++ "/", 'true') + after_add = json_decode(file_read(sp)) + r3 = Config.set_trusted_at(sp, d, 'false') + after_rm = json_decode(file_read(sp)) + shell("rm -rf " ++ Util.shell_q(d)) + tp = map_get(after_add, 'trusted_projects') + check("trust: adds the dir once (keeps other keys), untrust removes it", + sec_all([eqs(elem(r1, 0), 'ok'), eqs(elem(r1, 2), 'true'), eqs(elem(r2, 2), 'false'), + eqs(length(tp), 1), eqs(hd(tp), elem(r1, 1)), + eqs(map_get(after_add, 'endpoint'), "http://127.0.0.1:8000"), + eqs(map_get(map_get(after_add, 'permissions'), 'bash'), "ask"), + eqs(elem(r3, 2), 'true'), eqs(length(map_get(after_rm, 'trusted_projects')), 0)])) +} + +fun t_trust_refuses_corrupt_settings() { + d = "/tmp/swc_trust_bad_" ++ to_string(random_int(1, 1000000000)) + file_mkdir(d) + sp = d ++ "/settings.json" + file_write(sp, "{broken") + r = Config.set_trusted_at(sp, d, 'true') + kept = file_read(sp) + missing = Config.set_trusted_at(sp, d ++ "/no/such/dir", 'true') + shell("rm -rf " ++ Util.shell_q(d)) + check("trust: never overwrites an unparseable settings.json; missing dir is an error", + sec_all([eqs(elem(r, 0), 'error'), eqs(kept, "{broken"), eqs(elem(missing, 0), 'error')])) +} + +fun t_json_pretty_round_trips() { + v = json_decode("{\"a\": 1, \"b\": [true, null, \"x\\\"y\"], \"c\": {}, \"d\": []}") + out = Util.json_pretty(v) + check("json_pretty: indented, parses back to the same value", + bool_and(string_contains(out, "\n \"b\": [\n true,"), + eqs(json_encode(json_decode(out)), json_encode(v)))) +} diff --git a/src/tools.sw b/src/tools.sw index abdb465..53ee302 100644 --- a/src/tools.sw +++ b/src/tools.sw @@ -30,6 +30,7 @@ import Mcp import Util import TestRunner import PathGuard +import CommandGuard export [exec_raw, max_output_bytes] @@ -39,7 +40,7 @@ export [exec_raw, max_output_bytes] # turn). Each tool now gets a cap tuned to its role, mirroring claude-code's # per-tool limits. truncate_output keeps HEAD+TAIL so the failing tail of a # log survives the cut. max_output_bytes() stays as the conservative default -# (and the read_ceiling base) for any caller without a dedicated cap. +# for any caller without a dedicated cap. fun max_output_bytes() { 6000 } fun bash_output_cap() { 24000 } fun run_tests_raw_cap() { 16000 } @@ -143,7 +144,9 @@ fun all_tools() { # removed: it only SIGALRM'd the direct child (leaking grandchildren) and left # swarmrt's shell() polling for an exit file that never arrived → a multi-minute # wedge. These constants supply the per-tool budgets in SECONDS; callers pass -# `* 1000` to shell_managed. Timeouts mirror Claude Code: +# `* 1000` to shell_managed. Helper pipelines go through run_sh (stdin from +# /dev/null — a child must never read the terminal or the MCP server's +# JSON-RPC stream); bash runs Util.noninteractive_wrap. Timeouts mirror Claude Code: # bash: 120s default, 600s max, overridable via timeout_ms arg # git/search: 30s · web_fetch: 45s · read-probes: 5s # ------------------------------------------------------------ @@ -288,16 +291,9 @@ fun do_bash(args, opts) { if (cmd == nil) { "error: missing 'command' argument" } else { - # Sudo guard — outright block unless explicitly enabled. The - # model has no safe way to enter a sudo password, and the - # ambient setuid escalation risk is too high to gate behind - # only a prompt (which is_dangerous_bash already does). - # Set SWARM_CODE_ALLOW_SUDO=1 to opt back in. cmd_s = to_string(cmd) - sudo_allowed = getenv("SWARM_CODE_ALLOW_SUDO") - if (string_contains(cmd_s, "sudo ") == 'true' && sudo_allowed != "1") { - "error: sudo is disabled — set SWARM_CODE_ALLOW_SUDO=1 to enable (acknowledge that an agent running sudo is high-risk)" - } + refusal = sudo_refusal(cmd_s) + if (refusal != nil) { refusal } else { bg_table = map_get(opts, 'bg_table') run_bg = bash_run_in_background(args) @@ -321,6 +317,20 @@ fun do_bash(args, opts) { } } +# Sudo guard — outright block unless explicitly enabled. The model has no +# safe way to enter a sudo password, and the ambient setuid escalation risk +# is too high to gate behind only a prompt (which the dangerous-command gate +# already does). Token-aware (CommandGuard: `sudols`, `env sudo …`, +# `$(sudo …)` all count; the word "sudo" inside an echo'd string doesn't), +# and applied to every command-running tool, not just bash. +# Set SWARM_CODE_ALLOW_SUDO=1 to opt back in. Returns nil or the error. +fun sudo_refusal(cmd_s) { + if (getenv("SWARM_CODE_ALLOW_SUDO") == "1") { nil } + else { if (CommandGuard.uses_sudo(cmd_s) == 'true') { + "error: sudo is disabled — set SWARM_CODE_ALLOW_SUDO=1 to enable (acknowledge that an agent running sudo is high-risk)" + } else { nil }} +} + # Classic blocking path: run under shell_managed with the resolved timeout. # shell_managed runs the command in its own process group, enforces the # timeout in C, killpg's the WHOLE tree on timeout, and lets a user ESC stop @@ -376,12 +386,12 @@ fun bash_label(cmd_s) { # Explicit detach — launch and return a task id immediately, no wait. fun bash_launch_bg(bg_table, cmd_s) { - id = Background.launch(bg_table, cmd_s, bash_label(cmd_s)) + id = Background.launch_cmd(bg_table, noninteractive_wrap(cmd_s), cmd_s, bash_label(cmd_s)) # launch returns an "error: ..." STRING if shell_detached failed — surface # it directly instead of treating the error string as a task id. if (string_starts_with(to_string(id), "error") == 'true') { id } else { - "[backgrounded] task " ++ id ++ " — log: " ++ Background.log_path_for(id) ++ + "[backgrounded] task " ++ id ++ " — log: " ++ Background.log_path_for(bg_table, id) ++ "\n(bg_tail / bg_result / bg_kill to manage; you'll get a bg_done wake when it finishes)" } } @@ -391,7 +401,9 @@ fun bash_launch_bg(bg_table, cmd_s) { # foreground contract), ESC (kill the whole pgroup, return partial tail), # or budget exhausted (leave it running so the heartbeat's bg_done fires). fun bash_auto_bg(bg_table, cmd_s, after_ms) { - id = Background.launch(bg_table, cmd_s, bash_label(cmd_s)) + # Same wrapped script as the foreground path, so a command behaves the + # same whichever way it ends up running (CI=1 env, comments, heredocs). + id = Background.launch_cmd(bg_table, noninteractive_wrap(cmd_s), cmd_s, bash_label(cmd_s)) # launch returns an "error: ..." STRING if shell_detached failed — surface # it directly (no task exists to wait on, claim, or release). if (string_starts_with(to_string(id), "error") == 'true') { id } @@ -412,7 +424,7 @@ fun bash_auto_bg(bg_table, cmd_s, after_ms) { # Budget exhausted → hand off to the heartbeat: release the claim # so poll_loop can finalize it and fire bg_done on completion. Background.fg_release(bg_table, id) - path = Background.log_path_for(id) + path = Background.log_path_for(bg_table, id) "[backgrounded after " ++ to_string(after_ms / 1000) ++ "s] still running — task " ++ id ++ ", log: " ++ path ++ "\n\nlast output:\n" ++ Background.tail_log(bg_table, id, 20) ++ @@ -475,18 +487,16 @@ fun drain_stdin_keys() { }} } -# Wrap a user command in a subshell that: -# - exports CI=1 and friends BEFORE the command runs (so every -# child process, including npm/cargo/pip subcommands, sees them) -# - redirects stdin from /dev/null (closes the tty ghost) -# - merges stderr into stdout for capture -# Applied before the timeout wrapper so the alarm covers the whole -# pipeline including the `export` setup. +# Wrap a user command so that it: +# - runs with CI=1 and friends exported (so every child process, +# including npm/cargo/pip subcommands, sees them) +# - reads stdin from /dev/null (closes the tty ghost) +# - merges stderr into stdout for capture — INCLUDING sh's own syntax +# errors, so the model sees why a command didn't parse +# The command goes in on its own lines (see Util.noninteractive_wrap), so a +# trailing `# comment` or a final heredoc terminator can't eat the wrapper. fun noninteractive_wrap(user_cmd) { - "( export CI=1 DEBIAN_FRONTEND=noninteractive NO_COLOR=1 FORCE_COLOR=0 " ++ - "NPM_CONFIG_YES=true PIP_DISABLE_PIP_VERSION_CHECK=1 " ++ - "PYTHONUNBUFFERED=1; " ++ - user_cmd ++ " ) &1" + Util.noninteractive_wrap(user_cmd) } fun bash_max_lines() { 100 } @@ -500,8 +510,12 @@ fun do_read(args) { # limit is max lines to return. Both optional; defaults match CC. offset_raw = map_get(args, 'offset') limit_raw = map_get(args, 'limit') - offset = if (offset_raw == nil) { 1 } else { to_int(offset_raw) } - limit = if (limit_raw == nil) { 2000 } else { to_int(limit_raw) } + # parse_int_safe, not to_int: the to_int BUILTIN returns nil for junk + # like "abc", which then leaked into the line numbers as "nil". + offset_n = if (offset_raw == nil) { 1 } else { parse_int_safe(to_string(offset_raw), 1) } + limit_n = if (limit_raw == nil) { 2000 } else { parse_int_safe(to_string(limit_raw), 2000) } + offset = if (offset_n < 1) { 1 } else { offset_n } + limit = if (limit_n < 1) { 2000 } else { limit_n } if (p == nil) { "error: missing 'path' argument" } else { @@ -521,83 +535,164 @@ fun read_file_capped(path, offset, limit) { # and `head -c` would all hang, freezing the tool worker. Refuse before # touching the path. Every probe below is also timeout-guarded as a # backstop (a path on a dead NFS mount can hang even `test`). - rf_r = shell_managed( - "test -f " ++ pq ++ " && echo reg || echo nonreg 2>&1", read_probe_timeout_s() * 1000) + rf_r = run_sh( + "if test -f " ++ pq ++ "; then echo reg; elif test -e " ++ pq ++ + "; then echo nonreg; else echo missing; fi", read_probe_timeout_s() * 1000) rf = string_trim(elem(rf_r, 1)) if (elem(rf_r, 2) == 'true') { # Probe was killed by timeout/ESC (e.g. a dead NFS mount) — say so # rather than misreporting the path as a non-regular file. "error: read of " ++ to_string(path) ++ " was interrupted (timeout or ESC) " ++ "before the file could be classified — the path may be on an unresponsive filesystem." + } else { if (rf == "missing") { + "error: file not found — " ++ to_string(path) ++ + "\n(Check the path; glob finds files by name.)" } else { if (rf != "reg") { "error: not a regular file — " ++ to_string(path) ++ "\n(FIFOs, devices, sockets, and directories are refused to avoid a hang. " ++ "Use bash with an explicit, bounded reader if you really need to.)" } else { # Check if file is binary before reading — binary content poisons the - # model's context and causes empty responses. - file_type = elem(shell_managed("file --brief --mime-type " ++ pq ++ " 2>&1", read_probe_timeout_s() * 1000), 1) - is_text = string_starts_with(string_trim(file_type), "text") == 'true' - is_json = string_contains(file_type, "json") == 'true' - is_xml = string_contains(file_type, "xml") == 'true' - if (is_text == 'true' || is_json == 'true' || is_xml == 'true') { - # Size guard: file_read() pulls the WHOLE file into memory, so a - # multi-GB file would OOM the VM before truncate_output ever runs. - # Stat first; for anything large, read only a capped head via - # `head -c` instead of slurping the whole thing. - size_str = string_trim(elem(shell_managed("wc -c < " ++ pq ++ " 2>/dev/null", read_probe_timeout_s() * 1000), 1)) - size = parse_int_safe(size_str, 0) - read_ceiling = read_output_cap() - content = if (size > read_ceiling) { - head = elem(shell_managed("head -c " ++ to_string(read_ceiling) ++ " " ++ pq ++ " 2>&1", read_probe_timeout_s() * 1000), 1) - head ++ "\n...[file is " ++ size_str ++ " bytes — showing first " ++ - to_string(read_ceiling) ++ ". Use bash sed/grep for specific ranges.]" + # model's context and causes empty responses. Binary means a NUL byte + # in the first 8KB (git's heuristic). MIME types misclassify text — + # libmagic reports .js/.ts as application/javascript — and `file` is + # missing from many slim container images. + nul_r = run_sh("head -c 8192 " ++ pq ++ " | tr -dc '\\000' | wc -c", + read_probe_timeout_s() * 1000) + nul_count = parse_int_safe(string_trim(elem(nul_r, 1)), 0) + if (nul_count == 0) { + # Size guard: file_read() pulls the WHOLE file into memory (and + # refuses anything over 1MB), so only files within that cap are + # read in-process. Bigger ones are windowed by LINE with sed — + # the old code loaded just `head -c 65000` and THEN applied + # offset/limit, so offset 15000 of a 20,000-line file came back + # empty with no marker. Either way the OUTPUT is capped. + st = file_stat(to_string(path)) + size = if (st == nil) { 0 } else { map_get(st, 'size') } + if (size == 0) { + "[empty file] " ++ to_string(path) ++ " exists but has no content." + } else { if (size <= file_read_cap()) { + content = file_read(path) + if (content == nil) { + "error: could not read " ++ path + } else { if (string_length(content) != size) { + # A NUL byte past the first 8KB: file_read stopped at it. + read_window_streamed(path, pq, offset, limit, + "\n[note: this file contains NUL bytes after the first 8KB; they are not shown]") + } else { + read_window_from_content(path, content, offset, limit) + }} } else { - file_read(path) - } - if (content == nil) { - "error: could not read " ++ path + read_window_streamed(path, pq, offset, limit, "") + }} + } else { + # Best-effort label; `file` may not be installed. + ft_r = run_sh("file --brief --mime-type " ++ pq ++ " 2>/dev/null", + read_probe_timeout_s() * 1000) + ft = string_trim(to_string(elem(ft_r, 1))) + label = if (string_length(ft) == 0) { "binary" } else { "binary, " ++ ft } + hint = if (string_starts_with(ft, "image/") == 'true') { + "\nUse read_image to look at an image." } else { - sliced = slice_lines(content, offset, limit) - truncate_output(sliced, read_output_cap()) + "\nUse bash with hexdump, xxd, strings, or file to inspect binary files." } - } else { - "error: binary file (" ++ string_trim(file_type) ++ ") — " ++ to_string(path) ++ - "\nUse bash with hexdump, xxd, strings, or file to inspect binary files." + "error: " ++ label ++ " file — " ++ to_string(path) ++ hint } - } } -} - -# Slice [offset, offset+limit) lines from content (1-based offset) and -# prefix each line with its 1-based file line number + a tab — `N\t`, -# the cat -n shape claude-code's Read emits. The schema PROMISED line numbers -# (the edit/multi_edit prose tells the model to strip the leading number+tab -# before using a line as old_string); previously the body returned raw content -# with no anchors, so the model dropped to `od -c` to count offsets. The line -# number is the ABSOLUTE file line (start + position-in-window + 1), so anchors -# stay correct even when reading a windowed slice with offset > 1. -fun slice_lines(content, offset, limit) { - lines = string_split(content, "\n") + } } } +} + +# file_read's in-process cap (the runtime returns nil above it). +fun file_read_cap() { 1048576 } + +# Window [offset, offset+limit) of an in-memory file (1-based offset). +fun read_window_from_content(path, content, offset, limit) { + parts = string_split(content, "\n") + # A final newline ends the last line; it doesn't start an empty one + # (a 6-line file used to read back as 7 lines). + lines = if (string_ends_with(content, "\n") == 'true' && length(parts) > 1) { + take_first_lines(parts, length(parts) - 1, []) + } else { parts } total = length(lines) - start = if (offset < 1) { 0 } else { offset - 1 } - if (start >= total) { - "" + if (offset > total) { read_past_eof(path, offset, total) } + else { + window = take_first_lines(drop_first_n(lines, offset - 1), limit, []) + render_read_window(window, offset, total) + } +} + +# Window of a big file, selected by line number with sed (bounded: sed quits +# after the last wanted line, head -c caps the bytes). The first output line +# is the file's line count (wc -l, +1 when the last line has no newline). +fun read_window_streamed(path, pq, offset, limit, note) { + last = offset + limit - 1 + cmd = "n=$(wc -l < " ++ pq ++ "); [ -n \"$(tail -c 1 " ++ pq ++ ")\" ] && n=$((n+1)); echo \"$n\"; " ++ + "sed -n '" ++ to_string(offset) ++ "," ++ to_string(last) ++ "p;" ++ to_string(last) ++ "q' " ++ pq ++ + " | head -c " ++ to_string(read_output_cap() + 1) + r = run_sh(cmd, search_timeout_s() * 1000) + if (elem(r, 2) == 'true') { + "error: read of " ++ to_string(path) ++ " timed out after " ++ to_string(search_timeout_s()) ++ + "s (a very large file?) — use bash with sed -n 'A,Bp' for a specific range." } else { - window = take_first_lines(drop_first_n(lines, start), limit, []) - number_lines(window, start + 1, "") + out = elem(r, 1) + nl = string_index_of(out, "\n") + total = if (nl < 0) { parse_int_safe(string_trim(out), 0) } else { parse_int_safe(string_trim(string_sub(out, 0, nl)), 0) } + body = if (nl < 0) { "" } else { string_sub(out, nl + 1, string_length(out) - nl - 1) } + if (offset > total) { read_past_eof(path, offset, total) } + else { + parts = string_split(body, "\n") + lines = if (string_ends_with(body, "\n") == 'true' && length(parts) > 1) { + take_first_lines(parts, length(parts) - 1, []) + } else { parts } + render_read_window(lines, offset, total) ++ note + } } } -# Join lines, each prefixed with `\t`. lineno is the 1-based file line -# number of the FIRST window line; it increments per line. Matches cat -n / -# claude-code Read so the model has stable anchors for edit old_strings. -fun number_lines(lst, lineno, acc) { - if (length(lst) == 0) { acc } +fun read_past_eof(path, offset, total) { + "error: offset " ++ to_string(offset) ++ " is past the end of " ++ to_string(path) ++ + " (" ++ to_string(total) ++ " lines) — use a smaller offset." +} + +# Prefix each window line with its 1-based file line number + a tab — +# `N\t`, the cat -n shape claude-code's Read emits (the edit / +# multi_edit prose tells the model to strip it before using a line as +# old_string). Numbers are ABSOLUTE file lines, so anchors stay right for a +# windowed read. The output is capped at read_output_cap() on a LINE +# boundary, and every cut says where to continue: +# [output capped at N bytes — lines A-K of T shown; continue with offset=K+1] +# [lines A-B of T shown — continue with offset=B+1] (limit reached) +fun render_read_window(window, start, total) { + r = cap_numbered_lines(window, start, read_output_cap(), [], 0) + parts = elem(r, 0) + last = elem(r, 1) + capped = elem(r, 2) + body = Util.join_all(parts) + shown = to_string(start) ++ "-" ++ to_string(last) ++ " of " ++ to_string(total) + if (capped == 'true') { + body ++ "\n[output capped at " ++ to_string(read_output_cap()) ++ " bytes — lines " ++ shown ++ + " shown; continue with offset=" ++ to_string(last + 1) ++ "]" + } else { if (last < total) { + body ++ "\n[lines " ++ shown ++ " shown — continue with offset=" ++ to_string(last + 1) ++ "]" + } else { body }} +} + +# → {parts, last_line_number_included, capped}. A single line longer than +# the whole budget is cut (and flagged) rather than dropped. +fun cap_numbered_lines(lines, lineno, budget, acc, used) { + if (length(lines) == 0) { {acc, lineno - 1, 'false'} } else { - h = hd(lst) - sep = if (string_length(acc) == 0) { "" } else { "\n" } - line = to_string(lineno) ++ "\t" ++ h - number_lines(tl(lst), lineno + 1, acc ++ sep ++ line) + piece = to_string(lineno) ++ "\t" ++ hd(lines) + sep = if (length(acc) == 0) { "" } else { "\n" } + need = used + string_length(sep) + string_length(piece) + if (need > budget) { + if (length(acc) == 0) { + {[string_sub(piece, 0, budget) ++ "…[line " ++ to_string(lineno) ++ " is " ++ + to_string(string_length(hd(lines))) ++ " bytes — cut here; use bash (cut -c / head -c) for the rest]"], + lineno, 'true'} + } else { {acc, lineno - 1, 'true'} } + } else { + cap_numbered_lines(tl(lines), lineno + 1, budget, list_append(acc, sep ++ piece), need) + } } } @@ -664,7 +759,7 @@ fun do_write(args, opts) { # so the caller can show a real overwrite diff. Display-only # (headless suppresses the render). Skipped for files ≥64KB — # a large paste isn't worth diffing in the preview. - capture_write_prior(opts, to_string(path)) + capture_write_prior(opts, to_string(raw_path_arg(args)), to_string(path)) ensure_parent_dirs(to_string(path)) rc = file_write(path, content) if (rc == 'ok') { @@ -678,28 +773,32 @@ fun do_write(args, opts) { } # Stash an about-to-be-overwritten file's prior content into the -# 'write_diff_table' ETS (keyed by the raw path arg, matching -# resolve_path_key in agent.sw), so show_edit_diff can render a real +# 'write_diff_table' ETS (keyed by the RAW path arg, matching +# resolve_path_key in agent.sw — `path` is the ~-expanded file), so show_edit_diff can render a real # overwrite diff. No table (subagent / test opts) → no-op. Only for an # existing file under 64KB; a fresh create or a huge blob is left alone. -fun capture_write_prior(opts, path) { +fun capture_write_prior(opts, key, path) { wd = if (opts == nil) { nil } else { map_get(opts, 'write_diff_table') } if (wd == nil) { 'ok' } else { if (file_exists(path) == 'true') { prior = file_read(path) - if (prior != nil && string_length(prior) < 65536) { - ets_put(wd, path, prior) + # file_read stops at a NUL byte: a length that disagrees with the + # on-disk size means binary content — no (bogus) diff for that. + st = file_stat(path) + disk = if (st == nil) { 0 - 1 } else { map_get(st, 'size') } + if (prior != nil && string_length(prior) < 65536 && string_length(prior) == disk) { + ets_put(wd, key, prior) } else { # ≥64KB (or unreadable): clear any stale prior from an # earlier write to the same path so show_edit_diff doesn't # render a bogus diff against outdated content. - ets_delete(wd, path) + ets_delete(wd, key) } } else { # Fresh create: same stale-entry guard (the path may have been # written before, then deleted out-of-band). - ets_delete(wd, path) + ets_delete(wd, key) } } } @@ -754,11 +853,46 @@ fun do_edit_impl(path, old_s, new_s, replace_all) { else { do_edit_impl_inner(path, old_s, new_s, replace_all) } } +# Load a file for edit / multi_edit → {'ok', content} | {'missing', nil} | +# {'error', message}. file_read is C-string based with a 1MB cap: it stops at +# the first NUL byte (an edit then wrote back the truncated prefix and said +# "ok") and returns nil above 1MB (edit then took the file for MISSING, so +# old_string="" overwrote it with new_string). Comparing the loaded length +# with the on-disk size catches both — the whole file, not just a prefix. +fun load_for_edit(path) { + p = to_string(path) + st = file_stat(p) + if (st == nil) { {'missing', nil} } + else { if (map_get(st, 'is_dir') == 'true') { {'error', "error: " ++ p ++ " is a directory"} } + else { + size = map_get(st, 'size') + if (size == 0) { {'ok', ""} } + else { if (size > file_read_cap()) { + {'error', "error: " ++ p ++ " is " ++ to_string(size) ++ " bytes — too large to edit safely " ++ + "(edit/multi_edit handle files up to 1MB). Use bash (sed -i, python) for this file."} + } else { + c = file_read(p) + if (c == nil) { {'error', "error: could not read " ++ p} } + else { if (string_length(c) != size) { + {'error', "error: " ++ p ++ " contains NUL bytes (binary data) — edit works on text and would " ++ + "truncate it at the first NUL. The file was not modified; use bash/python for binary files."} + } else { {'ok', c} }} + }} + }} +} + fun do_edit_impl_inner(path, old_s, new_s, replace_all) { - original = file_read(path) + loaded = load_for_edit(path) + if (elem(loaded, 0) == 'error') { elem(loaded, 1) } + else { do_edit_loaded(path, elem(loaded, 1), old_s, new_s, replace_all) } +} + +fun do_edit_loaded(path, original, old_s, new_s, replace_all) { if (original == nil) { # Missing file. Empty old_string = create it with new_string. if (string_length(old_s) == 0) { + # Same as write: create missing parent directories first. + ensure_parent_dirs(to_string(path)) rc_c = file_write(path, new_s) if (rc_c == 'ok') { "ok: created " ++ path ++ " (" ++ to_string(string_length(new_s)) ++ " bytes)" @@ -788,6 +922,9 @@ fun do_edit_impl_inner(path, old_s, new_s, replace_all) { # Count occurrences in the ORIGINAL — checking the # post-replace buffer would falsely fire whenever # new_string contains old_string as a substring. + eol = crlf_adapt(original, old_s, new_s) + old_s = elem(eol, 0) + new_s = elem(eol, 1) occ = count_substrings(original, old_s) if (occ == 0) { # Exact miss. Before failing, retry against normalized @@ -823,6 +960,24 @@ fun do_edit_impl_inner(path, old_s, new_s, replace_all) { } } +# Models send LF-only strings, so a multi-line old_string never matched a +# CRLF file. When the file uses CRLF and the LF form isn't found, match and +# write with CRLF instead ({old, new}); an exact match always wins, so files +# with mixed line endings are left alone. +fun crlf_adapt(buffer, old_s, new_s) { + if (string_contains(buffer, "\r\n") == 'true' && + string_contains(old_s, "\n") == 'true' && + string_contains(old_s, "\r") == 'false' && + count_substrings(buffer, old_s) == 0) { + {to_crlf(old_s), to_crlf(new_s)} + } else { {old_s, new_s} } +} + +fun to_crlf(s) { + if (string_contains(s, "\r") == 'true') { s } + else { replace_all_occ(s, "\n", "\r\n") } +} + # ------------------------------------------------------------ # Edit normalization fallback (F7) # ------------------------------------------------------------ @@ -907,20 +1062,14 @@ fun join_chars(lst, acc) { # ------------------------------------------------------------ fun do_glob(args) { pattern = map_get(args, 'pattern') - path = map_get(args, 'path') + path = opt_path_arg(args, 'path') if (pattern == nil) { "error: missing 'pattern'" } else { - base = if (path == nil) { "." } else { to_string(path) } + base = if (path == nil || string_length(string_trim(to_string(path))) == 0) { "." } + else { to_string(path) } pat = to_string(pattern) pat_q = Util.shell_q(pat) - # When base is ".", DON'T pass it to rg — `rg --files . ` emits - # `./`-prefixed paths, and an anchored glob like `src/**/*.sw` - # never matches `./src/...`. With no path arg rg emits clean - # relative paths and anchored globs work. - base_part = if (base == "." || string_length(string_trim(base)) == 0) { - "" - } else { " " ++ Util.shell_q(base) } base_q = Util.shell_q(base) # ripgrep `--files -g` gives real glob semantics, but rg is # often unavailable to a plain /bin/sh (e.g. when it's only a @@ -942,20 +1091,29 @@ fun do_glob(args) { } else { "-name " ++ Util.shell_q(pat_find) } - # Pipe through `sed s|^\./||` to strip the `./` prefix `find` - # adds when invoked with `.` — matches ripgrep's clean output - # so downstream tools (Read) get the same path either way. + # The search root is ALWAYS passed explicitly (`.` by default): rg + # with no path searches its STDIN whenever stdin isn't a tty. Both + # rg and find then print `./`-prefixed paths for `.`, so one + # `sed s|^\./||` gives clean relative paths either way (anchored + # globs like `src/**/*.sw` still match — rg globs are relative to + # the search root). stderr (a bad glob) goes to a private file and + # is reported, instead of leaking onto the user's terminal. + errf = search_errfile() cmd = "if command -v rg >/dev/null 2>&1; then " ++ - "rg --files --hidden --no-messages -g " ++ pat_q ++ base_part ++ "; " ++ - "else find " ++ base_q ++ " -type f " ++ find_expr ++ " 2>/dev/null | sed 's|^\\./||'; fi | head -n 101" - result = shell_managed(cmd ++ " 2>&1", search_timeout_s() * 1000) - code = elem(result, 0) + "rg --files --hidden --no-messages -g " ++ pat_q ++ " " ++ base_q ++ "; " ++ + "else find " ++ base_q ++ " -type f " ++ find_expr ++ " 2>/dev/null; fi 2>" ++ errfile_redirect(errf) ++ + " | sed 's|^\\./||' | head -n 101" + result = run_sh(cmd, search_timeout_s() * 1000) out = elem(result, 1) interrupted = elem(result, 2) + err = take_errfile(errf) if (interrupted == 'true') { "[timed out after " ++ to_string(search_timeout_s()) ++ "s on " ++ base ++ " — narrow the path]" } else { - if (string_length(string_trim(out)) == 0) { "(no matches)" } + if (string_length(string_trim(out)) == 0) { + if (string_length(err) > 0) { "error: glob failed: " ++ err } + else { "(no matches)" } + } else { glob_cap_notice(out) } } } @@ -989,16 +1147,13 @@ fun do_grep(args) { pattern = map_get(args, 'pattern') if (pattern == nil) { "error: missing 'pattern'" } else { - path = map_get(args, 'path') + path = opt_path_arg(args, 'path') glob_arg = map_get(args, 'glob') mode = map_get(args, 'output_mode') hl = map_get(args, 'head_limit') - base = if (path == nil) { "." } else { to_string(path) } + base = if (path == nil || string_length(string_trim(to_string(path))) == 0) { "." } + else { to_string(path) } pat_q = Util.shell_q(to_string(pattern)) - # Omit a "." path so rg emits clean (non-`./`-prefixed) paths. - base_part = if (base == "." || string_length(string_trim(base)) == 0) { - "" - } else { " " ++ Util.shell_q(base) } base_q = Util.shell_q(base) # Optional glob filter (rg --glob / mirrors Claude Code's Grep). @@ -1026,41 +1181,124 @@ fun do_grep(args) { # `if/then/else` (not `rg || grep`) so an rg run that finds # nothing doesn't fall through and re-run grep over the whole # tree. --hidden so dotfiles aren't skipped; -e so a pattern - # starting with `-` isn't read as a flag. - grep_base = if (base == "." || string_length(string_trim(base)) == 0) { - " ." - } else { " " ++ base_q } - # Strip leading `./` so grep's `path:line:text` matches rg's - # `path:line:text` exactly — the model often hands those paths - # back to Read, which doesn't need (and shouldn't see) a `./`. + # starting with `-` isn't read as a flag. The fallback uses -E so + # the regex dialect matches rg's (ERE-style groups/alternation). + # + # The search path is ALWAYS explicit (`.` by default): with no path + # rg searches its STDIN when stdin isn't a tty — in --mcp-server + # mode that is the JSON-RPC stream, so it blocked for 30s and + # swallowed the client's next request. `sed s|^\./||` strips the + # `./` prefix that `.` adds so `path:line:text` stays clean (the + # model hands those paths back to read). + # + # stderr goes to a private file, not through the pipe: that's where + # an invalid regex (`foo(`) is reported, and the old `… | head 2>&1` + # sent it to the user's terminal while the model saw "(no matches)". + # --no-messages / -s keep unreadable-file noise out of it. + errf = search_errfile() cmd = "if command -v rg >/dev/null 2>&1; then " ++ "rg --color=never --hidden --no-messages" ++ rg_mode ++ glob_flag ++ - " -e " ++ pat_q ++ base_part ++ "; " ++ - "else grep" ++ grep_mode ++ " --color=never -e " ++ pat_q ++ grep_base ++ " 2>/dev/null | sed 's|^\\./||'; fi" ++ - " | head -n " ++ to_string(head_n) - result = shell_managed(cmd ++ " 2>&1", search_timeout_s() * 1000) - code = elem(result, 0) + " -e " ++ pat_q ++ " " ++ base_q ++ "; " ++ + "else grep -s -E" ++ grep_mode ++ " --color=never -e " ++ pat_q ++ " " ++ base_q ++ "; fi" ++ + " 2>" ++ errfile_redirect(errf) ++ + " | sed 's|^\\./||' | head -n " ++ to_string(head_n) + result = run_sh(cmd, search_timeout_s() * 1000) out = elem(result, 1) interrupted = elem(result, 2) + err = take_errfile(errf) if (interrupted == 'true') { "[timed out after " ++ to_string(search_timeout_s()) ++ "s on " ++ base ++ " — narrow the path]" } else { trimmed = truncate_output(out, grep_output_cap()) - if (string_length(string_trim(trimmed)) == 0) { "(no matches)" } else { trimmed } + if (string_length(string_trim(trimmed)) == 0) { + if (string_length(err) > 0) { + "error: grep failed: " ++ err ++ + "\n(For an invalid regex, escape literal ( ) [ ] { } . * + ? | with a backslash.)" + } else { "(no matches)" } + } else { + if (string_length(err) > 0) { trimmed ++ "\n[grep stderr: " ++ err ++ "]" } + else { trimmed } + } } } } +# Private temp file for a search pipeline's stderr (mkstemp → 0600, unique +# per call). nil if the temp dir is unusable — the command then discards +# stderr instead of leaking it onto the terminal. +fun search_errfile() { + file_temp(tmp_dir() ++ "/swarm-code-search-") +} + +fun errfile_redirect(errf) { + if (errf == nil) { "/dev/null" } else { Util.shell_q(errf) } +} + +# Read (capped, trimmed) and delete a search_errfile. "" when there was none. +fun take_errfile(errf) { + if (errf == nil) { "" } + else { + raw = file_read(errf) + file_delete(errf) + if (raw == nil) { "" } + else { string_trim(string_truncate(raw, 2000)) } + } +} + +fun tmp_dir() { + t = getenv("TMPDIR") + if (t == nil || string_length(to_string(t)) == 0) { "/tmp" } + else { + ts = to_string(t) + if (string_ends_with(ts, "/") == 'true' && string_length(ts) > 1) { + string_sub(ts, 0, string_length(ts) - 1) + } else { ts } + } +} + +# shell_managed for the harness's own helper pipelines: stdin is /dev/null +# (Util.no_stdin). shell_managed children otherwise inherit swarm-code's +# stdin — the terminal, or the MCP server's JSON-RPC pipe. +fun run_sh(cmd, timeout_ms) { + shell_managed(Util.no_stdin(cmd), timeout_ms) +} + # ------------------------------------------------------------ # helpers # ------------------------------------------------------------ -# Accept either 'path' or 'file_path' (Claude Code uses file_path). +# Accept either 'path' or 'file_path' (Claude Code uses file_path), with a +# leading `~` / `~/` expanded to $HOME (see expand_home). fun resolve_path_arg(args) { + p = raw_path_arg(args) + if (p == nil) { nil } else { expand_home(p) } +} + +fun raw_path_arg(args) { p = map_get(args, 'path') if (p != nil) { p } else { map_get(args, 'file_path') } } +# `~` / `~/…` → $HOME. Tool paths never pass through a shell unquoted, so +# nothing expanded them: `read ~/.bashrc` said "file not found" and +# `write ~/x` created a literal `./~/` directory in the cwd. (`~user` is +# left alone.) +fun expand_home(p) { + s = to_string(p) + home = getenv("HOME") + if (home == nil) { s } + else { if (s == "~") { to_string(home) } + else { if (string_starts_with(s, "~/") == 'true') { + to_string(home) ++ string_sub(s, 1, string_length(s) - 1) + } else { s }}} +} + +# An optional directory/file arg for search/wait tools: nil stays nil. +fun opt_path_arg(args, key) { + v = map_get(args, key) + if (v == nil) { nil } else { expand_home(v) } +} + # Truncate long output to prevent context blowup. # # Old version kept ONLY the head — exactly wrong for long builds where @@ -1265,12 +1503,14 @@ fun do_background(args, opts) { if (cmd == nil) { "error: background needs 'command'" } else { if (bg_table == nil) { "error: background system not initialized" } + else { if (sudo_refusal(to_string(cmd)) != nil) { sudo_refusal(to_string(cmd)) } else { label_str = if (label == nil) { to_string(cmd) } else { to_string(label) } - id = Background.launch(bg_table, to_string(cmd), label_str) - "launched " ++ id ++ ": " ++ label_str ++ - "\n(use bg_status and bg_result to check progress)" - } + id = Background.launch_cmd(bg_table, noninteractive_wrap(to_string(cmd)), to_string(cmd), label_str) + if (string_starts_with(to_string(id), "error") == 'true') { id } + else { "launched " ++ id ++ ": " ++ label_str ++ + "\n(use bg_status and bg_result to check progress)" } + }} } } @@ -1315,14 +1555,18 @@ fun do_bg_server(args, opts) { if (cmd == nil) { "error: bg_server needs 'command'" } else { if (bg_table == nil) { "error: background system not initialized" } + else { if (sudo_refusal(to_string(cmd)) != nil) { sudo_refusal(to_string(cmd)) } else { label_str = if (label == nil) { to_string(cmd) } else { to_string(label) } - id = Background.launch_server(bg_table, to_string(cmd), label_str) - log_file = Background.log_path_for(id) - "launched detached server " ++ id ++ ": " ++ label_str ++ - "\nlog: " ++ log_file ++ - "\n(use bg_tail to read log, bg_kill to stop)" - } + id = Background.launch_cmd(bg_table, noninteractive_wrap(to_string(cmd)), to_string(cmd), label_str) + if (string_starts_with(to_string(id), "error") == 'true') { id } + else { + log_file = Background.log_path_for(bg_table, id) + "launched detached server " ++ id ++ ": " ++ label_str ++ + "\nlog: " ++ log_file ++ + "\n(use bg_tail to read log, bg_kill to stop)" + } + }} } } @@ -1381,7 +1625,7 @@ fun do_web_search(args) { # the JSON "diag" field, but python itself can still emit warnings / # SSL chatter on stderr. Folding stderr into stdout would prepend # that text and break json_decode of an otherwise-good result. - result = shell_managed(cmd ++ " 2>/dev/null", fetch_timeout_s() * 1000) + result = run_sh(cmd ++ " 2>/dev/null", fetch_timeout_s() * 1000) code = elem(result, 0) out = elem(result, 1) interrupted = elem(result, 2) @@ -1564,26 +1808,39 @@ fun ddg_python_script() { # ------------------------------------------------------------ # run_tests parse and gate on test output from any framework # ------------------------------------------------------------ -# args: {"repo_path": "/path/to/repo", "command": "npm test" (optional)} +# args: {"repo_path": "/path/to/repo", "command": "npm test" (optional), +# "timeout_ms": 300000 (optional, 1000..600000)} fun do_run_tests(args) { repo = map_get(args, 'repo_path') cmd = map_get(args, 'command') if (repo == nil) { "error: run_tests needs 'repo_path'" } + else { if (cmd != nil && sudo_refusal(to_string(cmd)) != nil) { sudo_refusal(to_string(cmd)) } else { command = if (cmd == nil) { "" } else { to_string(cmd) } - result = TestRunner.run_tests(to_string(repo), command) + t_raw = map_get(args, 'timeout_ms') + t_s = if (t_raw == nil) { TestRunner.default_timeout_ms() / 1000 } else { resolve_bash_timeout_s(t_raw) } + result = TestRunner.run_tests_timed(expand_home(repo), command, t_s * 1000) fw = map_get(result, 'framework') passed = map_get(result, 'passed') failed = map_get(result, 'failed') total = map_get(result, 'total') exit_code = map_get(result, 'exit_code') raw = map_get(result, 'raw') - summary = "Framework: " ++ fw ++ "\n" ++ + timed_out = map_get(result, 'timed_out') + banner = if (timed_out == 'true') { + if (exit_code == 130) { "[stopped by user (ESC/Ctrl-C) — test process group killed]\n" } + else { "[timed out after " ++ to_string(t_s) ++ "s — test process group killed; pass a larger timeout_ms or a narrower command]\n" } + } else { "" } + summary = banner ++ + "Framework: " ++ fw ++ "\n" ++ "Passed: " ++ to_string(passed) ++ "\n" ++ "Failed: " ++ to_string(failed) ++ "\n" ++ "Total: " ++ to_string(total) ++ "\n" ++ "Exit code: " ++ to_string(exit_code) - if (failed > 0) { + # Show the output tail whenever something went wrong — failed tests, + # a non-zero exit with nothing parsed (compile error, missing tool), + # or a timeout — not only when the parser counted failures. + if (failed > 0 || exit_code != 0 || timed_out == 'true') { cap = run_tests_raw_cap() tail = if (string_length(raw) > cap) { string_sub(raw, string_length(raw) - cap, cap) @@ -1594,7 +1851,7 @@ fun do_run_tests(args) { } else { summary } - } + }} } # ------------------------------------------------------------ @@ -1604,7 +1861,7 @@ fun do_git_status(args) { cwd_arg = map_get(args, 'cwd') cwd_part = if (cwd_arg == nil) { "" } else { "-C " ++ Util.shell_q(to_string(cwd_arg)) ++ " " } cmd = git_noninteractive_env() ++ "git " ++ cwd_part ++ "status --porcelain --branch 2>&1 | head -n 100" - r = shell_managed(cmd, git_timeout_s() * 1000) + r = run_sh(cmd, git_timeout_s() * 1000) code = elem(r, 0) out = elem(r, 1) interrupted = elem(r, 2) @@ -1624,7 +1881,7 @@ fun do_git_diff(args) { cwd_part = if (cwd_arg == nil) { "" } else { "-C " ++ Util.shell_q(to_string(cwd_arg)) ++ " " } flag = if (staged == 'true') { "--staged " } else { "" } cmd = git_noninteractive_env() ++ "git " ++ cwd_part ++ "diff " ++ flag ++ "--no-color 2>&1" - r = shell_managed(cmd, git_timeout_s() * 1000) + r = run_sh(cmd, git_timeout_s() * 1000) code = elem(r, 0) out = elem(r, 1) interrupted = elem(r, 2) @@ -1659,7 +1916,7 @@ fun do_git_commit(args) { "{ git " ++ cwd_part ++ "add " ++ stage_list ++ " && git " ++ cwd_part ++ "commit -m " ++ Util.shell_q(to_string(msg)) ++ " && git " ++ cwd_part ++ "rev-parse --short HEAD ; } 2>&1" - r = shell_managed(cmd, git_timeout_s() * 1000) + r = run_sh(cmd, git_timeout_s() * 1000) code = elem(r, 0) out = elem(r, 1) interrupted = elem(r, 2) @@ -1696,7 +1953,7 @@ fun do_code_search(args) { if (pat == nil) { "error: code_search needs 'pattern'" } else { kind = map_get(args, 'kind') - path = map_get(args, 'path') + path = opt_path_arg(args, 'path') lang = map_get(args, 'lang') k = if (kind == nil) { "ref" } else { to_string(kind) } base = if (path == nil) { "." } else { to_string(path) } @@ -1729,7 +1986,7 @@ fun do_code_search(args) { Util.shell_q(rgx) ++ " " ++ base_q ++ " || grep -rn --color=never -E " ++ Util.shell_q(rgx) ++ " " ++ base_q ++ ") 2>&1 | head -n 80" - r = shell_managed(cmd, search_timeout_s() * 1000) + r = run_sh(cmd, search_timeout_s() * 1000) code = elem(r, 0) out = elem(r, 1) interrupted = elem(r, 2) @@ -1754,15 +2011,19 @@ fun do_log_wait(args, opts) { if (pat == nil) { "error: log_wait needs 'pattern'" } else { task_id = map_get(args, 'task_id') - path_arg = map_get(args, 'path') - timeout = map_get(args, 'timeout_sec') - timeout_n = if (timeout == nil) { 60 } else { parse_int_safe(to_string(timeout), 60) } + path_arg = opt_path_arg(args, 'path') + timeout_n = clamp_wait_timeout_s(map_get(args, 'timeout_sec')) - # Resolve log path: explicit path, or task_id's log file + # Resolve log path: explicit path, or task_id's log file (in this + # session's private background directory). + bg_table = map_get(opts, 'bg_table') log_path = if (path_arg != nil) { to_string(path_arg) } else { - if (task_id == nil) { "" } - else { "/tmp/swarm-code-" ++ to_string(task_id) ++ ".log" } + if (task_id == nil || bg_table == nil) { "" } + else { + lp = Background.log_path_for(bg_table, to_string(task_id)) + if (lp == nil) { "" } else { lp } + } } if (string_length(log_path) == 0) { "error: log_wait needs either 'task_id' or 'path'" @@ -1777,7 +2038,7 @@ fun do_log_wait(args, opts) { # wait, and so the `sleep` poll-loop's whole process group dies on # timeout. Timeout/interrupt surface via the interrupted flag now, # not exit 142. - r = shell_managed(inner, timeout_n * 1000) + r = run_sh(inner, timeout_n * 1000) code = elem(r, 0) interrupted = elem(r, 2) if (interrupted == 'true') { @@ -1791,31 +2052,52 @@ fun do_log_wait(args, opts) { } } +# log_wait / file_watch `timeout_sec`: default 60, clamped to [1, 600]. +# Unclamped, 0 reached shell_managed as "no timeout" — a 600s wedge headless, +# unbounded on a TTY. Non-numeric junk falls back to the default. +fun clamp_wait_timeout_s(raw) { + if (raw == nil) { 60 } + else { + n = parse_int_safe(to_string(raw), 60) + if (n < 1) { 1 } else { if (n > 600) { 600 } else { n } } + } +} + # ------------------------------------------------------------ -# file_watch — block until a file changes (mtime) or appears +# file_watch — block until a file changes (mtime/size), appears, or vanishes # ------------------------------------------------------------ # args: {"path": "/path/to/file", "timeout_sec": 60} -# Returns when the file's mtime changes or it appears, or on timeout. +# Returns when the file's signature changes, or on timeout. fun do_file_watch(args) { - path_arg = map_get(args, 'path') + path_arg = opt_path_arg(args, 'path') if (path_arg == nil) { "error: file_watch needs 'path'" } else { path = to_string(path_arg) - timeout = map_get(args, 'timeout_sec') - timeout_n = if (timeout == nil) { 60 } else { to_int(timeout) } + timeout_n = clamp_wait_timeout_s(map_get(args, 'timeout_sec')) - # Capture initial mtime, then poll every 0.5s until it changes. + # Capture an initial signature, then poll every 0.5s until it changes. # Run via shell_managed so the timeout is enforced in C and the poll # loop's process group is killed on timeout/ESC (no GNU `timeout` dep). + # + # The path is single-quoted with Util.shell_q: it comes from the model, + # and the old double-quote escaping ran `$(…)` / backticks in it. + # The signature is mtime (full resolution where GNU stat has it) + + # size, so two writes in the same second still differ. GNU and BSD + # stat disagree on flags — GNU `stat -f` is --file-system, which + # printed ever-changing free-space counters and fired "changed" in + # 0.5s on an untouched file — so pick the flavour once, up front. inner = - "p=" ++ shell_inner_quote(path) ++ "; " ++ - "initial=$(stat -f %m \"$p\" 2>/dev/null || echo missing); " ++ - "while true; do " ++ - " current=$(stat -f %m \"$p\" 2>/dev/null || echo missing); " ++ - " [ \"$current\" != \"$initial\" ] && echo \"changed: $initial -> $current\" && exit 0; " ++ - " sleep 0.5; " ++ + "p=" ++ Util.shell_q(path) ++ "\n" ++ + "if stat -c %Y / >/dev/null 2>&1; then " ++ + "sig() { stat -c '%y %s' \"$p\" 2>/dev/null || echo missing; }; " ++ + "else sig() { stat -f '%m %z' \"$p\" 2>/dev/null || echo missing; }; fi\n" ++ + "initial=$(sig)\n" ++ + "while :; do " ++ + "current=$(sig); " ++ + "if [ \"$current\" != \"$initial\" ]; then echo \"changed: $initial -> $current\"; exit 0; fi; " ++ + "sleep 0.5; " ++ "done" - r = shell_managed(inner, timeout_n * 1000) + r = run_sh(inner, timeout_n * 1000) code = elem(r, 0) out = string_trim(elem(r, 1)) interrupted = elem(r, 2) @@ -1866,7 +2148,7 @@ fun do_sw_check(args) { " >/dev/null 2>" ++ Util.shell_q(errf) ++ "; S=$?; " ++ "grep -v 'auto-imported\\|cannot open' " ++ Util.shell_q(errf) ++ "; rm -f " ++ Util.shell_q(errf) ++ "; exit $S" - r = shell_managed(cmd, 60 * 1000) + r = run_sh(cmd, 60 * 1000) code = elem(r, 0) out = string_trim(elem(r, 1)) interrupted = elem(r, 2) @@ -1891,19 +2173,13 @@ fun resolve_swc() { override = getenv("SWARM_CODE_SWC") if (override != nil) { to_string(override) } else { - r = shell_managed("command -v swc 2>/dev/null", read_probe_timeout_s() * 1000) + r = run_sh("command -v swc 2>/dev/null", read_probe_timeout_s() * 1000) found = string_trim(elem(r, 1)) if (string_length(found) > 0) { found } else { "../swarmrt/bin/swc" } } } -# Escape for use INSIDE a double-quoted shell string (bash). -fun shell_inner_quote(s) { - no_dq = string_replace(s, "\"", "\\\"") - "\"" ++ no_dq ++ "\"" -} - # Format the parsed results list as a readable string for the model. fun format_search_results(results, i, acc) { if (length(results) == 0) { @@ -1938,12 +2214,14 @@ fun do_multi_edit(args) { guard = PathGuard.validate_write(to_string(path)) if (guard != "ok") { guard } else { - original = file_read(path) - if (original == nil) { - "error: could not read " ++ path + loaded = load_for_edit(path) + tag = elem(loaded, 0) + if (tag == 'error') { elem(loaded, 1) } + else { if (tag == 'missing') { + "error: could not read " ++ path ++ " (no such file — use write to create it)" } else { - apply_edits(path, original, edits, 0, length(edits)) - } + apply_edits(path, elem(loaded, 1), edits, 0, length(edits)) + }} } } } @@ -1971,6 +2249,9 @@ fun apply_edits(path, buffer, edits, count, total) { # Count in the CURRENT buffer (pre-replace) — using # the post-replace check produced false positives when # new_string contained old_string as a substring. + eol = crlf_adapt(buffer, old_s, new_s) + old_s = elem(eol, 0) + new_s = elem(eol, 1) occ = count_substrings(buffer, old_s) if (occ == 0) { # Exact miss — retry against normalized copies (quotes / @@ -2084,7 +2365,7 @@ fun do_web_fetch(args, opts) { " -e 's/ / /g' -e 's/&/\\&/g'" ++ " -e 's/<//g'" ++ " -e 's/"/\"/g' | tr -s ' \\n' | head -c 30000" - result = shell_managed(strip_cmd ++ " 2>&1", fetch_timeout_s() * 1000) + result = run_sh(strip_cmd ++ " 2>&1", fetch_timeout_s() * 1000) text = elem(result, 1) file_delete(tmp_path) "fetched " ++ url ++ " (" ++ to_string(string_length(text)) ++ diff --git a/src/trajectory.sw b/src/trajectory.sw index b225cb5..8828716 100644 --- a/src/trajectory.sw +++ b/src/trajectory.sw @@ -36,11 +36,15 @@ import Log # journal (which already excludes the runtime system prompt by # design — see Agent.encode_journal). # -# Privacy: every exported line passes through Log.redact, which masks -# common secret shapes (sk-/AKIA/ghp_-style tokens, Bearer headers, -# *_key / *_token / *_secret fields, long letter+digit blobs). The -# heuristics are not exhaustive — still review exports manually before -# publishing them anywhere. +# Privacy: every string VALUE of an exported example passes through +# Log.redact_value BEFORE json_encode, which masks common secret shapes +# (sk-/AKIA/ghp_-style tokens, Bearer headers, PEM blocks, URL +# passwords, PASSWORD= / "password": / password: / *_key / *_token / +# *_secret fields, long letter+digit blobs). Redacting the encoded line +# instead produced invalid JSONL (a masked run could eat the letter of a +# \n escape) and missed secrets hidden behind escapes. The heuristics +# are not exhaustive — still review exports manually before publishing +# them anywhere. export [ export_all, export_current, @@ -105,7 +109,7 @@ fun export_one(journal_path, out_path) { else { wire = clean_messages(messages, []) example = %{messages: wire} - file_append(out_path, Log.redact(json_encode(example)) ++ "\n") + file_append(out_path, json_encode(Log.redact_value(example)) ++ "\n") 'true' } } @@ -122,7 +126,7 @@ fun export_current(out_path, history) { if (length(wire) < 2) { %{path: out_path, kept: 0, reason: "session too short"} } else { - file_write(out_path, Log.redact(json_encode(%{messages: wire})) ++ "\n") + file_write(out_path, json_encode(Log.redact_value(%{messages: wire})) ++ "\n") %{path: out_path, kept: 1} } } diff --git a/src/util.sw b/src/util.sw index 66f04fd..05e9a67 100644 --- a/src/util.sw +++ b/src/util.sw @@ -9,7 +9,7 @@ module Util # Agent/Config/Tools). Keeping these here means a bug fix lands # once, not 8 times. -export [shell_q] +export [shell_q, noninteractive_wrap, no_stdin, join_all, json_args_well_formed, json_pretty] # POSIX-safe single-quote wrap. Replaces `'` with `'\''` (close, # escape, reopen) so the result is always safe to splice into a @@ -17,3 +17,184 @@ export [shell_q] fun shell_q(s) { "'" ++ string_replace(s, "'", "'\\''") ++ "'" } + +# ------------------------------------------------------------ +# Shell-script wrappers for commands built from MODEL input +# ------------------------------------------------------------ +# The wrapped command is spliced in as its own LINES, never inline between +# parentheses: the old `( export …; CMD ) &1` form broke on any +# command ending in a `# comment` (it commented out the closing paren) or a +# heredoc (the terminator line became `EOF ) &1; export CI=1 DEBIAN_FRONTEND=noninteractive NO_COLOR=1 FORCE_COLOR=0 " ++ + "NPM_CONFIG_YES=true PIP_DISABLE_PIP_VERSION_CHECK=1 PYTHONUNBUFFERED=1\n" ++ + to_string(user_cmd) ++ "\n" +} + +# no_stdin: for the harness's OWN helper pipelines (grep/glob/git/probes). +# shell_managed children inherit swarm-code's stdin — a tool that reads stdin +# by accident (rg with no path) would block on the terminal or swallow the +# MCP server's next request. stderr is left alone: each caller decides. +fun no_stdin(cmd) { + "exec 0) { 'true' } else { 'false' } + if (in_str == 'true') { + # String content with no closing quote after it: unterminated. + if (more == 'false') { 'false' } + else { + escaped = if (trailing_backslashes(p, string_length(p) - 1, 0) % 2 == 1) { 'true' } else { 'false' } + jwf_parts(tl(parts), escaped, stack, started, done) + } + } else { + r = jwf_scan(p, 0, string_length(p), stack, started, done) + if (r == 'bad') { 'false' } + else { + st = elem(r, 0) + sd = elem(r, 1) + dn = elem(r, 2) + if (more == 'false') { dn } + else { + # A quote follows: it may only open a string INSIDE the + # top-level container. + if (sd == 'true' && dn == 'false') { jwf_parts(tl(parts), 'true', st, sd, dn) } + else { 'false' } + } + } + } +} + +fun trailing_backslashes(p, i, acc) { + if (i < 0) { acc } + else { if (string_sub(p, i, 1) == "\\") { trailing_backslashes(p, i - 1, acc + 1) } + else { acc } } +} + +# Walk one outside-string piece. Returns {stack, started, done} or 'bad'. +fun jwf_scan(p, i, n, stack, started, done) { + if (i >= n) { {stack, started, done} } + else { + c = string_sub(p, i, 1) + if (c == " " || c == "\n" || c == "\r" || c == "\t") { + jwf_scan(p, i + 1, n, stack, started, done) + } else { if (done == 'true') { 'bad' } + else { if (started == 'false') { + if (c == "{" || c == "[") { jwf_scan(p, i + 1, n, c, 'true', 'false') } + else { 'bad' } + } else { if (c == "{" || c == "[") { + jwf_scan(p, i + 1, n, stack ++ c, started, done) + } else { if (c == "}" || c == "]") { + sl = string_length(stack) + want = if (c == "}") { "{" } else { "[" } + if (sl == 0) { 'bad' } + else { if (string_sub(stack, sl - 1, 1) != want) { 'bad' } + else { + closed = if (sl == 1) { 'true' } else { 'false' } + jwf_scan(p, i + 1, n, string_sub(stack, 0, sl - 1), started, closed) + }} + } else { + jwf_scan(p, i + 1, n, stack, started, done) + }}}}} + } +} + +# ------------------------------------------------------------ +# Indented JSON for files people also edit by hand (settings.json). +# json_encode writes one line; this keeps the file readable after +# swarm-code rewrites it. Map key order is preserved. +# ------------------------------------------------------------ +fun json_pretty(v) { jp_val(v, "") ++ "\n" } + +fun jp_val(v, ind) { + if (is_map(v) == 'true') { + if (map_size(v) == 0) { "{}" } + else { "{\n" ++ jp_members(map_keys(v), map_values(v), ind ++ " ", "") ++ "\n" ++ ind ++ "}" } + } else { if (is_list(v) == 'true') { + if (length(v) == 0) { "[]" } + else { "[\n" ++ jp_items(v, ind ++ " ", "") ++ "\n" ++ ind ++ "]" } + } else { json_encode(v) } } +} + +fun jp_members(keys, vals, ind, acc) { + if (length(keys) == 0) { acc } + else { + sep = if (string_length(acc) == 0) { "" } else { ",\n" } + m = ind ++ json_encode(to_string(hd(keys))) ++ ": " ++ jp_val(hd(vals), ind) + jp_members(tl(keys), tl(vals), ind, acc ++ sep ++ m) + } +} + +fun jp_items(items, ind, acc) { + if (length(items) == 0) { acc } + else { + sep = if (string_length(acc) == 0) { "" } else { ",\n" } + jp_items(tl(items), ind, acc ++ sep ++ ind ++ jp_val(hd(items), ind)) + } +} diff --git a/tests/integration/agents_cases.sh b/tests/integration/agents_cases.sh new file mode 100644 index 0000000..0ef499d --- /dev/null +++ b/tests/integration/agents_cases.sh @@ -0,0 +1,542 @@ +# Integration cases for the multi-agent / MCP / scheduler / persistence +# surfaces. Sourced by tests/integration/run.sh — uses its helpers +# (new_case, start_mock, run_swarm, cleanup, final_json, req_has, +# req_count, pass, fail) and its $ROOT / $BIN / $CASE* variables. +# +# A1 MCP EOF — a server that dies mid-call is "connection +# lost" (not a timeout) and the next call +# takes the reconnect path +# A2 MCP id collision — a server→client request reusing our id is +# answered (-32601; ping → {}) and never taken +# as the response +# A3 MCP output text — a successful result that merely CONTAINS +# "connection lost" / "did not respond" never +# touches server health (no reconnect), and +# headless --json stdout stays one line +# A4 --mcp-server spec — ping → {}, unknown tool → -32602, bad +# "jsonrpc" / object or null ids → -32600, +# notifications get no reply +# A5 schedule.json shape — a wrong-shape file ({"jobs":[]}) or a job +# with "last_run":"yesterday" no longer +# crashes the interactive session; one warning +# A6 corrupt schedule — /schedule refuses to overwrite a damaged +# schedule.json (file byte-identical) +# A7 unattended jobs — a scheduled job's child runs with +# SWARM_CODE_DENY_DANGEROUS=1: `rm -rf ~/…` +# requested by its model is denied +# A8 subagent guardrail — a subagent's 8 failing reads stop the +# SUBAGENT (partial result + reason), never +# the parent's turn +# A9 explore is read-only — an explore subagent's bash / write calls +# are refused (nothing executes) +# A10 subagent max steps — hitting the step limit returns the work so +# far, not just a notice +# A11 subagent big answer — a 320KB answer reaches the parent capped +# A12 subagent LLM hang — a hung endpoint fails the subagent after +# the configured LLM timeout, with a reason +# A13 /flows malformed — {"phases":"oops"} and friends print an error; +# the session survives and stays responsive +# A14 /flows fan-out cap — 6 tasks with SWARM_CODE_FLOWS_MAX_PARALLEL=2 +# run at most 2 at a time, all complete +# +# A5-A7, A13, A14 drive the binary INTERACTIVELY (the scheduler runs off +# main's heartbeat; /flows is a slash command) through +# tests/integration/pty_run.py. +# +# Run standalone: tests/integration/run.sh (these run after T1..T10). + +FAKE_MCP="$ROOT/tests/integration/fake_mcp.py" +PTY_RUN="$ROOT/tests/integration/pty_run.py" + +# run_pty — run the binary interactively in a pty with +# the isolated env; transcript → $CASE/pty.txt, "ALIVE"/"DEAD n" → +# $PTY_STATUS. 90s watchdog, like run_swarm. +run_pty() { + PTY_STATUS="$( + cd "$WORK" || exit 97 + HOME="$CASE_HOME" \ + SWARM_CODE_EXECUTION_CONTEXT=main \ + SWARM_CODE_ENDPOINT="http://127.0.0.1:$PORT" \ + SWARM_CODE_MODEL=test \ + SWARM_CODE_TOOL_FORMAT=native \ + SWARM_CODE_PLAN=off \ + SWARM_CODE_BIN="$BIN" \ + SWARM_CODE_FLOWS_MAX_PARALLEL="${RUN_FLOWS_MAX:-}" \ + TERM=xterm SW_NO_TITLE=1 \ + perl -e 'alarm 90; exec @ARGV' python3 "$PTY_RUN" "$CASE/pty.txt" "$1" \ + "$BIN" --no-resume 2>"$CASE/pty.err" + )" +} + +# mcp_settings — user settings.json wiring the fake MCP server +# (its stdin log lands in $CASE/mcp.log). The explicit allow keeps the +# test independent of headless permission defaults for mcp__ tools. +mcp_settings() { + mkdir -p "$CASE_HOME/.swarm-code" + cat >"$CASE_HOME/.swarm-code/settings.json" </dev/null)" + echo "${n:-0}" +} + +# ------------------------------------------------------------ +# A1 — server exits on tools/call: EOF is a lost connection +# ------------------------------------------------------------ +a1() { + new_case a1 + mcp_settings die_on_call + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [ + {"type": "tool_calls", "calls": [{"id": "q1", "name": "mcp__fake__echo", "arguments": {"text": "one"}}]}, + {"type": "tool_calls", "calls": [{"id": "q2", "name": "mcp__fake__echo", "arguments": {"text": "two"}}]}, + {"type": "text", "content": "A1_DONE"} +]} +EOF + start_mock "$CASE/scenario.json" || { fail A1 "mock failed to start"; return; } + run_swarm -p "use the mcp tool" --no-resume --json + cleanup + local out; out="$(final_json)" + if [ "$RC" -ne 0 ]; then fail A1 "exit code $RC" + elif ! echo "$out" | grep -q "A1_DONE"; then fail A1 "final text missing: $out" + elif ! req_has 1 "connection lost"; then fail A1 "first call not reported as connection lost" + elif req_has 1 "did not respond"; then fail A1 "EOF misreported as a timeout" + elif [ "$(n_inits)" -ne 2 ]; then fail A1 "expected a reconnect (2 initializes), got $(n_inits)" + else pass A1; fi +} + +# ------------------------------------------------------------ +# A2 — server→client request with a colliding id + ping +# ------------------------------------------------------------ +a2() { + new_case a2 + mcp_settings collide + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [ + {"type": "tool_calls", "calls": [{"id": "q1", "name": "mcp__fake__echo", "arguments": {"text": "one"}}]}, + {"type": "text", "content": "A2_DONE"} +]} +EOF + start_mock "$CASE/scenario.json" || { fail A2 "mock failed to start"; return; } + run_swarm -p "use the mcp tool" --no-resume --json + cleanup + local out; out="$(final_json)" + if [ "$RC" -ne 0 ]; then fail A2 "exit code $RC" + elif ! echo "$out" | grep -q "A2_DONE"; then fail A2 "final text missing: $out" + elif req_has 1 "carried no result"; then fail A2 "server request taken as the response" + elif ! req_has 1 "ECHO:one roots=-32601 ping=ok"; then fail A2 "server requests not answered per spec" + else pass A2; fi +} + +# ------------------------------------------------------------ +# A3 — result text never drives health bookkeeping +# ------------------------------------------------------------ +a3() { + new_case a3 + mcp_settings ok + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [ + {"type": "tool_calls", "calls": [{"id": "q1", "name": "mcp__fake__echo", "arguments": {"text": "log: db connection lost at 12:00"}}]}, + {"type": "tool_calls", "calls": [{"id": "q2", "name": "mcp__fake__echo", "arguments": {"text": "peer did not respond (twice)"}}]}, + {"type": "tool_calls", "calls": [{"id": "q3", "name": "mcp__fake__echo", "arguments": {"text": "again: did not respond"}}]}, + {"type": "tool_calls", "calls": [{"id": "q4", "name": "mcp__fake__echo", "arguments": {"text": "third connection lost"}}]}, + {"type": "text", "content": "A3_DONE"} +]} +EOF + start_mock "$CASE/scenario.json" || { fail A3 "mock failed to start"; return; } + run_swarm -p "use the mcp tool" --no-resume --json + cleanup + local lines; lines="$(wc -l <"$CASE/stdout.txt" | tr -d ' ')" + if [ "$RC" -ne 0 ]; then fail A3 "exit code $RC" + elif ! final_json | grep -q "A3_DONE"; then fail A3 "final text missing: $(final_json)" + elif [ "$lines" -ne 1 ]; then fail A3 "stdout has $lines lines, want 1 JSON line" + elif ! req_has 4 "ECHO:third connection lost"; then fail A3 "4th call did not succeed" + elif req_has 4 "is not running"; then fail A3 "server marked failed by result text" + elif [ "$(n_inits)" -ne 1 ]; then fail A3 "result text forced a reconnect ($(n_inits) initializes)" + elif grep -q "reconnecting" "$CASE/stderr.txt"; then fail A3 "spurious reconnect notice" + else pass A3; fi +} + +# ------------------------------------------------------------ +# A4 — --mcp-server envelope + method conformance (no LLM) +# ------------------------------------------------------------ +a4() { + new_case a4 + ( + cd "$WORK" || exit 97 + printf '%s\n' \ + '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"a4","version":"1"}}}' \ + '{"jsonrpc":"2.0","method":"notifications/initialized"}' \ + '{"jsonrpc":"2.0","id":2,"method":"ping"}' \ + '{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"no_such_tool","arguments":{}}}' \ + '{"jsonrpc":"1.0","id":4,"method":"ping"}' \ + '{"id":5,"method":"ping"}' \ + '{"jsonrpc":"2.0","id":{"x":1},"method":"ping"}' \ + '{"jsonrpc":"2.0","id":"six","method":"no/such/method"}' | + HOME="$CASE_HOME" perl -e 'alarm 20; exec @ARGV' "$BIN" --mcp-server \ + >"$CASE/stdout.txt" 2>"$CASE/stderr.txt" + ) + RC=$? + if [ "$RC" -ne 0 ]; then fail A4 "MCP server exit code $RC"; return; fi + local why + why="$(python3 - "$CASE/stdout.txt" <<'PYEOF' +import json, sys +lines = [l for l in open(sys.argv[1]).read().splitlines() if l.strip()] +msgs = [] +for l in lines: + try: + msgs.append(json.loads(l)) + except ValueError: + print("non-JSON stdout line: %r" % l[:120]); sys.exit(0) +def by_id(i): + return [m for m in msgs if m.get("id") == i] +def code(m): + return (m.get("error") or {}).get("code") +checks = [ + (len(msgs) == 7, "want 7 responses (notification unanswered), got %d" % len(msgs)), + (all(m.get("jsonrpc") == "2.0" for m in msgs), "a response lacks jsonrpc 2.0"), + (by_id(2) and by_id(2)[0].get("result") == {}, "ping must return an empty result"), + (by_id(3) and code(by_id(3)[0]) == -32602, "unknown tool must be -32602"), + (by_id(4) and code(by_id(4)[0]) == -32600, "jsonrpc 1.0 must be -32600"), + (by_id(5) and code(by_id(5)[0]) == -32600, "missing jsonrpc must be -32600"), + (len([m for m in by_id(None) if code(m) == -32600]) == 1, "object id must be -32600 with id null"), + (not [m for m in msgs if isinstance(m.get("id"), dict)], "object id echoed back"), + (by_id("six") and code(by_id("six")[0]) == -32601, "unknown method must be -32601"), +] +bad = [why for ok, why in checks if not ok] +print(bad[0] if bad else "") +PYEOF +)" + if [ -n "$why" ]; then fail A4 "$why" + else pass A4; fi +} + +# ------------------------------------------------------------ +# A5 — wrong-shape schedule.json must not kill the session +# ------------------------------------------------------------ +a5() { + new_case a5 + local sched="$CASE_HOME/.swarm-code/schedule.json" + mkdir -p "$CASE_HOME/.swarm-code" + echo '{"responses": []}' >"$CASE/scenario.json" + # Survive 3+ heartbeat ticks, then prove main still answers input (a + # panicked main can leave the process itself lingering). + cat >"$CASE/pty.json" <<'EOF' +[{"wait": 7}, {"send": "/schedules\r"}, + {"until_out": "see /help for /schedule usage", "timeout": 8}] +EOF + start_mock "$CASE/scenario.json" || { fail A5 "mock failed to start"; return; } + printf '%s' '{"jobs":[]}' >"$sched" + run_pty "$CASE/pty.json" + local s1="$PTY_STATUS"; cp "$CASE/pty.txt" "$CASE/pty1.txt" + printf '%s' '[{"id":"1","expr":"1h","prompt":"x","last_run":"yesterday"}]' >"$sched" + cp "$sched" "$CASE/sched2.orig" + run_pty "$CASE/pty.json" + local s2="$PTY_STATUS" + cleanup + local warns; warns="$(grep -c "schedule: " "$CASE/pty1.txt")" + if [ "$s1" != "ALIVE" ] || grep -q "panic:" "$CASE/pty1.txt" || + ! grep -q "see /help for /schedule usage" "$CASE/pty1.txt"; then + fail A5 "session crashed/unresponsive with {\"jobs\":[]} ($s1)" + elif ! grep -q "must be a JSON array" "$CASE/pty1.txt"; then fail A5 "no warning for a non-array schedule.json" + elif [ "$warns" -ne 1 ]; then fail A5 "warning printed $warns times, want once" + elif [ "$s2" != "ALIVE" ] || grep -q "panic:" "$CASE/pty.txt" || + ! grep -q "see /help for /schedule usage" "$CASE/pty.txt"; then + fail A5 "session crashed/unresponsive on last_run:\"yesterday\" ($s2)" + elif ! grep -q "last_run must be a timestamp" "$CASE/pty.txt"; then fail A5 "no warning for the bad job" + elif ! cmp -s "$sched" "$CASE/sched2.orig"; then fail A5 "tick rewrote an entry it could not validate" + else pass A5; fi +} + +# ------------------------------------------------------------ +# A6 — /schedule must not replace a corrupt schedule.json +# ------------------------------------------------------------ +a6() { + new_case a6 + local sched="$CASE_HOME/.swarm-code/schedule.json" + mkdir -p "$CASE_HOME/.swarm-code" + printf '%s' '[{"id":"1","expr":"1d","prompt":"nightly report","last_run":0,"runs":3},{"id":"2",' >"$sched" + cp "$sched" "$CASE/sched.orig" + echo '{"responses": []}' >"$CASE/scenario.json" + cat >"$CASE/pty.json" <<'EOF' +[{"wait": 4}, {"send": "/schedule \"5m\" \"new job\"\r"}, + {"until_out": "refusing to overwrite", "timeout": 8}, {"wait": 1}] +EOF + start_mock "$CASE/scenario.json" || { fail A6 "mock failed to start"; return; } + run_pty "$CASE/pty.json" + cleanup + if [ "$PTY_STATUS" != "ALIVE" ] || grep -q "panic:" "$CASE/pty.txt"; then fail A6 "session died ($PTY_STATUS)" + elif ! cmp -s "$sched" "$CASE/sched.orig"; then fail A6 "corrupt schedule.json was overwritten: $(head -c 200 "$sched")" + elif ! grep -q "refusing to overwrite" "$CASE/pty.txt"; then fail A6 "no clear refusal message" + else pass A6; fi +} + +# ------------------------------------------------------------ +# A7 — a due job's child must not auto-approve dangerous bash +# ------------------------------------------------------------ +a7() { + new_case a7 + mkdir -p "$CASE_HOME/.swarm-code" "$CASE_HOME/victim" + touch "$CASE_HOME/victim/keep" + cat >"$CASE_HOME/.swarm-code/schedule.json" <<'EOF' +[{"id":"1","expr":"30s","prompt":"cleanup the victim dir","created_at":0,"last_run":0,"runs":0,"paused":false}] +EOF + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [ + {"type": "tool_calls", "calls": [{"id": "c1", "name": "bash", "arguments": {"command": "rm -rf ~/victim && echo gone"}}]}, + {"type": "text", "content": "CRON_DONE_A7"} +]} +EOF + start_mock "$CASE/scenario.json" || { fail A7 "mock failed to start"; return; } + cat >"$CASE/pty.json" < — byte length of that tool result +# in request #n (the parent's view of a task result). +tool_msg_len() { + python3 - "$REQLOG" "$1" "$2" <<'PYEOF2' +import json, sys +path, n, tcid = sys.argv[1], int(sys.argv[2]), sys.argv[3] +for line in open(path): + r = json.loads(line) + if r["n"] == n: + for m in r["body"].get("messages", []): + if m.get("role") == "tool" and m.get("tool_call_id") == tcid: + print(len(m.get("content") or "")); sys.exit(0) +print(-1) +PYEOF2 +} + +# ------------------------------------------------------------ +# A8 — subagent failure streak halts the subagent, not the parent +# ------------------------------------------------------------ +a8() { + new_case a8 + python3 - "$CASE/scenario.json" <<'PYEOF2' +import json, sys +reads = [{"id": "r%d" % i, "name": "read", "arguments": {"path": "missing_%d.txt" % i}} for i in range(8)] +json.dump({"responses": [ + {"type": "tool_calls", "calls": [{"id": "m1", "name": "task", "arguments": { + "description": "find config", "prompt": "locate the config file", "subagent_type": "explore"}}]}, + {"type": "tool_calls", "calls": reads}, + # Next request must be the PARENT's: the halted subagent never asks again. + {"type": "text", "content": "A8_MAIN_FINAL"}]}, open(sys.argv[1], "w")) +PYEOF2 + start_mock "$CASE/scenario.json" || { fail A8 "mock failed to start"; return; } + run_swarm -p "find the config" --no-resume --json + cleanup + local out; out="$(final_json)" + if [ "$RC" -ne 0 ]; then fail A8 "exit code $RC: $out" + elif grep -q "guardrail halt\]" "$CASE/stderr.txt"; then fail A8 "subagent's streak halted the PARENT turn" + elif ! echo "$out" | grep -q '"status":"ok"'; then fail A8 "parent did not finish ok: $out" + elif ! echo "$out" | grep -q "A8_MAIN_FINAL"; then fail A8 "halted subagent kept calling the LLM: $out" + elif ! req_has 2 "subagent stopped: guardrail halt"; then fail A8 "parent never got the subagent's halt result" + elif ! req_has 2 "missing_7.txt"; then fail A8 "partial result lacks the subagent's tool digest" + else pass A8; fi +} + +# ------------------------------------------------------------ +# A9 — explore subagent cannot run bash or write +# ------------------------------------------------------------ +a9() { + new_case a9 + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [ + {"type": "tool_calls", "calls": [{"id": "m1", "name": "task", "arguments": {"description": "survey", "prompt": "look around read-only", "subagent_type": "explore"}}]}, + {"type": "tool_calls", "calls": [{"id": "s1", "name": "bash", "arguments": {"command": "touch EXPLORE_RAN_BASH"}}]}, + {"type": "tool_calls", "calls": [{"id": "s2", "name": "write", "arguments": {"path": "explore_wrote.txt", "content": "x"}}]}, + {"type": "text", "content": "sub done"}, + {"type": "text", "content": "A9_MAIN_DONE"} +]} +EOF + start_mock "$CASE/scenario.json" || { fail A9 "mock failed to start"; return; } + run_swarm -p "survey the repo" --no-resume --json + cleanup + if [ -e "$WORK/EXPLORE_RAN_BASH" ]; then fail A9 "explore subagent ran bash" + elif [ -e "$WORK/explore_wrote.txt" ]; then fail A9 "explore subagent wrote a file" + elif [ "$RC" -ne 0 ]; then fail A9 "exit code $RC" + elif ! req_has 2 "not available to an explore (read-only) subagent"; then fail A9 "bash refusal not explained" + elif ! req_has 3 "not available to an explore (read-only) subagent"; then fail A9 "write refusal not explained" + elif ! final_json | grep -q "A9_MAIN_DONE"; then fail A9 "final text missing: $(final_json)" + else pass A9; fi +} + +# ------------------------------------------------------------ +# A10 — max steps returns the work done so far +# ------------------------------------------------------------ +a10() { + new_case a10 + python3 - "$CASE/scenario.json" <<'PYEOF2' +import json, sys +steps = [{"type": "tool_calls", "content": "", "calls": [{"id": "g%d" % i, "name": "bash", + "arguments": {"command": "echo FINDING_%d" % i}}]} for i in range(15)] +json.dump({"responses": [ + {"type": "tool_calls", "calls": [{"id": "m1", "name": "task", "arguments": { + "description": "dig", "prompt": "investigate", "subagent_type": "general"}}]}] + + steps + [{"type": "text", "content": "A10_MAIN_DONE"}]}, open(sys.argv[1], "w")) +PYEOF2 + start_mock "$CASE/scenario.json" || { fail A10 "mock failed to start"; return; } + run_swarm -p "investigate" --no-resume --json + cleanup + if [ "$RC" -ne 0 ]; then fail A10 "exit code $RC" + elif [ "$(req_count)" -ne 17 ]; then fail A10 "expected 17 requests (1 + 15 sub + 1), got $(req_count)" + elif ! req_has 16 "15-step limit"; then fail A10 "no max-steps notice" + elif ! req_has 16 "FINDING_14"; then fail A10 "subagent's findings discarded at max steps" + elif ! final_json | grep -q "A10_MAIN_DONE"; then fail A10 "final text missing" + else pass A10; fi +} + +# ------------------------------------------------------------ +# A11 — a huge subagent answer is capped before it reaches the parent +# ------------------------------------------------------------ +a11() { + new_case a11 + python3 - "$CASE/scenario.json" <<'PYEOF2' +import json, sys +big = "".join("finding %06d: details details details\n" % i for i in range(8000)) # ~320KB +json.dump({"responses": [ + {"type": "tool_calls", "calls": [{"id": "m1", "name": "task", "arguments": { + "description": "dig", "prompt": "investigate", "subagent_type": "general"}}]}, + # Small deltas, like a real server (one huge delta is cut by the + # runtime's SSE reader and would not reproduce the uncapped answer). + {"type": "text", "content": big, "chunk": 2000}, + {"type": "text", "content": "A11_MAIN_DONE"}]}, open(sys.argv[1], "w")) +PYEOF2 + start_mock "$CASE/scenario.json" || { fail A11 "mock failed to start"; return; } + run_swarm -p "investigate" --no-resume --json + cleanup + local n; n="$(tool_msg_len 2 m1)" + if [ "$RC" -ne 0 ]; then fail A11 "exit code $RC" + elif [ "$n" -lt 0 ]; then fail A11 "parent never received the task result" + elif [ "$n" -gt 26000 ]; then fail A11 "subagent answer reached the parent uncapped ($n bytes)" + elif ! req_has 2 "bytes of the subagent's answer elided"; then fail A11 "no truncation marker" + elif ! req_has 2 "finding 000000" || ! req_has 2 "finding 007999"; then fail A11 "head/tail of the answer lost" + else pass A11; fi +} + +# ------------------------------------------------------------ +# A12 — a hung LLM endpoint fails the subagent after the LLM timeout +# ------------------------------------------------------------ +a12() { + new_case a12 + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [ + {"type": "tool_calls", "calls": [{"id": "m1", "name": "task", "arguments": {"description": "dig", "prompt": "investigate", "subagent_type": "general"}}]}, + {"type": "text", "content": "never delivered", "delay": 15}, + {"type": "text", "content": "A12_MAIN_DONE"} +]} +EOF + start_mock "$CASE/scenario.json" || { fail A12 "mock failed to start"; return; } + SWARM_CODE_LLM_TIMEOUT_MS=3000 run_swarm -p "investigate" --no-resume --json + cleanup + # How long the parent was blocked: subagent request (#1) → the + # parent's next request (#2). Must track the 3s LLM timeout, not + # the old fixed 300s wait (nor the endpoint's 15s stall). + local gap + gap="$(python3 -c 'import json,sys; t={} +for l in open(sys.argv[1]): + r=json.loads(l); t[r["n"]]=r["t"] +print(int(t[2]-t[1]) if 1 in t and 2 in t else 999)' "$REQLOG")" + if [ "$RC" -ne 0 ]; then fail A12 "exit code $RC: $(final_json)" + elif [ "$gap" -gt 10 ]; then fail A12 "parent blocked ${gap}s on a hung subagent LLM call" + elif ! req_has 2 "no response from the LLM within 3s"; then fail A12 "failure reason not surfaced" + elif ! final_json | grep -q "A12_MAIN_DONE"; then fail A12 "final text missing: $(final_json)" + else pass A12; fi +} + +# ------------------------------------------------------------ +# A13 — a malformed /flows workflow must not kill the session +# ------------------------------------------------------------ +a13() { + new_case a13 + printf '%s' '{"phases":"oops"}' >"$WORK/bad1.json" + printf '%s' '{"phases":[{"name":"P","tasks":[7, {"prompt":"x"}]}]}' >"$WORK/bad2.json" + printf '%s' '{"phases":[{"tasks":[{"label":"a","prompt":"x"}' >"$WORK/bad3.json" + echo '{"responses": []}' >"$CASE/scenario.json" + cat >"$CASE/pty.json" <<'EOF' +[{"wait": 4}, + {"send": "/flows ./bad1.json\r"}, {"until_out": "must be an array", "timeout": 6}, + {"send": "/flows ./bad2.json\r"}, {"until_out": "task 1 must be an object", "timeout": 6}, + {"send": "/flows ./bad3.json\r"}, {"until_out": "invalid JSON in", "timeout": 6}, + {"send": "/schedules\r"}, {"until_out": "see /help for /schedule usage", "timeout": 6}] +EOF + start_mock "$CASE/scenario.json" || { fail A13 "mock failed to start"; return; } + run_pty "$CASE/pty.json" + cleanup + if [ "$PTY_STATUS" != "ALIVE" ] || grep -q "panic:" "$CASE/pty.txt"; then fail A13 "session died on a malformed workflow ($PTY_STATUS)" + elif ! grep -q '"phases" must be an array' "$CASE/pty.txt"; then fail A13 "no error for {\"phases\":\"oops\"}" + elif ! grep -q "task 1 must be an object" "$CASE/pty.txt"; then fail A13 "no error for a non-object task" + elif ! grep -q "invalid JSON in" "$CASE/pty.txt"; then fail A13 "truncated JSON accepted" + elif ! grep -q "see /help for /schedule usage" "$CASE/pty.txt"; then fail A13 "session unresponsive afterwards" + elif [ "$(req_count)" -ne 0 ]; then fail A13 "a malformed workflow launched tasks" + else pass A13; fi +} + +# ------------------------------------------------------------ +# A14 — /flows fan-out honours the concurrency cap +# ------------------------------------------------------------ +a14() { + new_case a14 + python3 - "$CASE/scenario.json" "$WORK/fan.json" <<'PYEOF2' +import json, sys +json.dump({"responses": [{"type": "text", "content": "child done", "delay": 3}] * 6}, + open(sys.argv[1], "w")) +json.dump({"name": "fan", "phases": [{"name": "P", "tasks": [ + {"label": "t%d" % i, "prompt": "child task %d" % i} for i in range(6)]}]}, + open(sys.argv[2], "w")) +PYEOF2 + cat >"$CASE/pty.json" < + +Speaks newline-delimited JSON-RPC 2.0 on stdin/stdout and exposes one tool, +`echo` ({"text": str} -> "ECHO:"). Every line received is appended to +$FAKE_MCP_LOG (prefixed "IN "), so a test can count handshakes (one +"initialize" per (re)connect) and inspect the client's replies. + +Modes: + ok well-behaved server + die_on_call exits as soon as a tools/call arrives (connection lost mid-call) + collide before answering tools/call , sends the client a + server->client request that REUSES (roots/list) plus a + ping, and waits (<=5s) for the client's replies; the tool + result then reports what came back: + "ECHO: roots= ping=" +""" +import json +import os +import select +import sys + +mode = sys.argv[1] if len(sys.argv) > 1 else "ok" +log = open(os.environ.get("FAKE_MCP_LOG", "/dev/null"), "a") + + +def out(obj): + sys.stdout.write(json.dumps(obj) + "\n") + sys.stdout.flush() + + +_buf = b"" + + +def read_line(timeout=None): + """One line from fd 0 via select+os.read — sys.stdin's own buffering + would hide already-read lines from select(). None on timeout / EOF.""" + global _buf + while b"\n" not in _buf: + r, _, _ = select.select([0], [], [], timeout) + if not r: + return None + chunk = os.read(0, 65536) + if not chunk: + return None + _buf += chunk + line, _buf = _buf.split(b"\n", 1) + text = line.decode("utf-8", "replace") + "\n" + log.write("IN " + text) + log.flush() + return text + + +def collect_replies(want_ids, timeout=5.0): + """Read client lines until every id in want_ids got a reply.""" + got = {} + while len(got) < len(want_ids): + line = read_line(timeout) + if not line: + break + try: + m = json.loads(line) + except ValueError: + continue + if isinstance(m, dict) and "method" not in m and m.get("id") in want_ids: + got[m["id"]] = m + return got + + +while True: + line = read_line() + if not line: + break + try: + m = json.loads(line) + except ValueError: + continue + mid = m.get("id") + meth = m.get("method") + if meth == "initialize": + out({"jsonrpc": "2.0", "id": mid, "result": { + "protocolVersion": "2025-06-18", "capabilities": {"tools": {}}, + "serverInfo": {"name": "fake", "version": "1"}}}) + elif meth == "tools/list": + out({"jsonrpc": "2.0", "id": mid, "result": {"tools": [{ + "name": "echo", "description": "echo text", + "inputSchema": {"type": "object", + "properties": {"text": {"type": "string"}}}}]}}) + elif meth == "tools/call": + text = str(m.get("params", {}).get("arguments", {}).get("text")) + if mode == "die_on_call": + sys.exit(0) + suffix = "" + if mode == "collide": + out({"jsonrpc": "2.0", "id": mid, "method": "roots/list"}) + out({"jsonrpc": "2.0", "id": "srv-ping-1", "method": "ping"}) + got = collect_replies([mid, "srv-ping-1"]) + roots = got.get(mid) + ping = got.get("srv-ping-1") + roots_s = str(roots["error"].get("code")) if roots and "error" in roots else "missing" + ping_s = "missing" if ping is None else ("ok" if ping.get("result") == {} else "bad") + suffix = " roots=%s ping=%s" % (roots_s, ping_s) + out({"jsonrpc": "2.0", "id": mid, "result": { + "content": [{"type": "text", "text": "ECHO:" + text + suffix}]}}) diff --git a/tests/integration/mock_llm.py b/tests/integration/mock_llm.py index 7784390..a5a9b77 100755 --- a/tests/integration/mock_llm.py +++ b/tests/integration/mock_llm.py @@ -9,11 +9,33 @@ {"type": "tool_calls", "calls": [ {"id": "call_1", "name": "bash", "arguments": {"command": "echo hello"}}]} - ]} - -Every request body is appended (one JSON object per line) to the log file, -so the harness can assert exactly what the binary sent — e.g. that a tool -result message came back after a tool_calls response. + ], + "silent": [ {"type": "text", "content": "summary"} ]} + +Streaming requests ("stream": true — every agent turn) consume +"responses". Non-streaming requests (compaction's summarizer, the daemon +pulse) consume "silent" when the scenario has that key, else they share +the "responses" queue (the original behavior). + +Response spec types: + text {"content": "..."} finish_reason "stop" + tool_calls {"calls": [{id, name, arguments}]} finish_reason "tool_calls" + arguments may be an object (JSON-encoded for you) or a raw + STRING sent verbatim — e.g. a truncated '{"command": "ech'. + Optional "content": prose streamed before the calls. + raw {"lines": ["data: {...}", ...]} each line sent + "\\n\\n" + http {"status": 400, "body": "...", "headers": {...}} +Common keys: "finish" overrides the finish_reason (e.g. "length"); +"delay_ms" / "delay" (seconds) sleep before answering, on the request's +own thread, so later requests are still served (a hung endpoint). A text +response may carry "chunk": to stream its content in n-byte deltas +(real servers send many small deltas; the default is two halves). + +Every request body is appended (one JSON object per line: n, kind, t, body) +to the log file, so the harness can assert exactly what the binary sent — +e.g. that a tool result message came back after a tool_calls response. +n counts per queue: stream requests 0,1,2…; silent requests 0,1,2… when +the scenario has a "silent" key. If more requests arrive than there are scripted responses, a plain text "MOCK-EXHAUSTED" response is served (so a looping binary terminates @@ -30,50 +52,57 @@ import json import sys import threading +import time from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer +USAGE = {"prompt_tokens": 10, "completion_tokens": 5, "total_tokens": 15} + + +def ev(delta, finish=None, usage=None): + e = {"choices": [{"index": 0, "delta": delta, "finish_reason": finish}]} + if usage: + e["usage"] = usage + return e -def sse_text_events(content): - """Chunked OpenAI SSE frames for a plain assistant text response.""" - mid = max(1, len(content) // 2) - parts = [content[:mid], content[mid:]] if len(content) > 1 else [content] - events = [{"choices": [{"index": 0, - "delta": {"role": "assistant", "content": ""}, - "finish_reason": None}]}] + +def sse_text_events(spec): + """Chunked OpenAI SSE frames for a plain assistant text response — + two halves, or `chunk`-sized deltas when the spec gives one.""" + content = spec.get("content", "") + chunk = int(spec.get("chunk") or 0) + if chunk > 0: + parts = [content[i:i + chunk] for i in range(0, len(content), chunk)] or [""] + else: + mid = max(1, len(content) // 2) + parts = [content[:mid], content[mid:]] if len(content) > 1 else [content] + events = [ev({"role": "assistant", "content": ""})] for p in parts: - events.append({"choices": [{"index": 0, "delta": {"content": p}, - "finish_reason": None}]}) - events.append({"choices": [{"index": 0, "delta": {}, - "finish_reason": "stop"}], - "usage": {"prompt_tokens": 10, "completion_tokens": 5, - "total_tokens": 15}}) + events.append(ev({"content": p})) + events.append(ev({}, spec.get("finish", "stop"), USAGE)) return events -def sse_tool_call_events(calls): +def sse_tool_call_events(spec): """Chunked SSE frames for a tool_calls response. The first frame for each call carries index/id/name; the arguments JSON is split across two later frames to exercise the client's fragment reassembly.""" events = [] - for i, call in enumerate(calls): - args = json.dumps(call.get("arguments", {})) + if spec.get("content"): + events.append(ev({"role": "assistant", "content": spec["content"]})) + for i, call in enumerate(spec["calls"]): + args = call.get("arguments", {}) + args = args if isinstance(args, str) else json.dumps(args) mid = max(1, len(args) // 2) frags = [args[:mid], args[mid:]] if len(args) > 1 else [args] - events.append({"choices": [{"index": 0, "delta": { + events.append(ev({ "role": "assistant", "tool_calls": [{"index": i, "id": call["id"], "type": "function", "function": {"name": call["name"], - "arguments": ""}}]}, - "finish_reason": None}]}) + "arguments": ""}}]})) for frag in frags: - events.append({"choices": [{"index": 0, "delta": { - "tool_calls": [{"index": i, - "function": {"arguments": frag}}]}, - "finish_reason": None}]}) - events.append({"choices": [{"index": 0, "delta": {}, - "finish_reason": "tool_calls"}], - "usage": {"prompt_tokens": 10, "completion_tokens": 5, - "total_tokens": 15}}) + events.append(ev({"tool_calls": [{"index": i, + "function": {"arguments": frag}}]})) + events.append(ev({}, spec.get("finish", "tool_calls"), USAGE)) return events @@ -81,17 +110,25 @@ class MockHandler(BaseHTTPRequestHandler): protocol_version = "HTTP/1.1" scenario = None log_path = None - counter = [0] + counters = {"stream": 0, "silent": 0} lock = threading.Lock() def log_message(self, fmt, *args): # silence default stderr access log pass + def send_body(self, status, ctype, payload, headers=None): + self.send_response(status) + for k, v in (headers or {}).items(): + self.send_header(k, v) + self.send_header("Content-Type", ctype) + self.send_header("Content-Length", str(len(payload))) + self.send_header("Connection", "close") + self.end_headers() + self.wfile.write(payload) + def do_POST(self): if not self.path.endswith("/chat/completions"): - self.send_response(404) - self.send_header("Content-Length", "0") - self.end_headers() + self.send_body(404, "text/plain", b"") return length = int(self.headers.get("Content-Length", 0)) @@ -100,40 +137,59 @@ def do_POST(self): body = json.loads(raw) except ValueError: body = {"_raw": raw} + stream = body.get("stream") in (True, "true") + kind = "stream" if stream else "silent" + # Without a "silent" queue every request shares "responses" (and + # one counter), exactly like the original mock. + queue = kind if "silent" in self.scenario else "stream" with self.lock: - n = self.counter[0] - self.counter[0] += 1 + n = self.counters[queue] + self.counters[queue] += 1 with open(self.log_path, "a") as f: - f.write(json.dumps({"n": n, "body": body}) + "\n") - - responses = self.scenario.get("responses", []) - if n < len(responses): - spec = responses[n] - else: - spec = {"type": "text", "content": "MOCK-EXHAUSTED"} + f.write(json.dumps({"n": n, "kind": kind, "t": time.time(), + "body": body}) + "\n") + + responses = self.scenario.get( + "silent" if queue == "silent" else "responses", []) + spec = responses[n] if n < len(responses) else \ + {"type": "text", "content": "MOCK-EXHAUSTED"} + if spec.get("delay_ms"): + time.sleep(spec["delay_ms"] / 1000.0) + if spec.get("delay"): + time.sleep(float(spec["delay"])) + t = spec.get("type", "text") + + if t == "http": + self.send_body(spec.get("status", 500), + spec.get("ctype", "application/json"), + spec.get("body", "").encode(), spec.get("headers")) + return + if not stream: + # Non-streaming chat.completion (compaction / pulse). + payload = json.dumps({ + "choices": [{"index": 0, "finish_reason": "stop", + "message": {"role": "assistant", + "content": spec.get("content", "")}}], + "usage": USAGE}).encode() + self.send_body(200, "application/json", payload) + return - if spec.get("type") == "tool_calls": - events = sse_tool_call_events(spec["calls"]) + if t == "raw": + payload = b"".join(l.encode() + b"\n\n" for l in spec["lines"]) else: - events = sse_text_events(spec.get("content", "")) - - # Compact separators are load-bearing: the swarmrt SSE extractor - # matches `"content":"` / `"arguments":"` with no space after the - # colon, exactly like real OpenAI-compatible servers emit. - payload = b"" - for ev in events: - payload += (b"data: " - + json.dumps(ev, separators=(",", ":")).encode() - + b"\n\n") - payload += b"data: [DONE]\n\n" - - self.send_response(200) - self.send_header("Content-Type", "text/event-stream") - self.send_header("Content-Length", str(len(payload))) - self.send_header("Connection", "close") - self.end_headers() - self.wfile.write(payload) + events = sse_tool_call_events(spec) if t == "tool_calls" \ + else sse_text_events(spec) + # Compact separators are load-bearing: the swarmrt SSE extractor + # matches `"content":"` / `"arguments":"` with no space after the + # colon, exactly like real OpenAI-compatible servers emit. + payload = b"" + for e in events: + payload += (b"data: " + + json.dumps(e, separators=(",", ":")).encode() + + b"\n\n") + payload += b"data: [DONE]\n\n" + self.send_body(200, "text/event-stream", payload) def main(): diff --git a/tests/integration/pty_run.py b/tests/integration/pty_run.py new file mode 100644 index 0000000..29b9cd0 --- /dev/null +++ b/tests/integration/pty_run.py @@ -0,0 +1,94 @@ +#!/usr/bin/env python3 +"""Drive the binary INTERACTIVELY under a pseudo-terminal for E2E tests. + + pty_run.py + +The environment is inherited (the caller sets HOME, SWARM_CODE_* ...), the +working directory is the caller's. Script steps, run in order, while the +child's output is continuously drained (so it never blocks on a full pty): + + {"wait": secs} pump output for secs + {"send": "text\\r"} type into the terminal + {"until": path, "contains": substr, "timeout": s} wait for file content + {"until_out": substr, "timeout": s} wait for terminal output + +At the end the ANSI-stripped transcript is written to , and +one status line is printed: "ALIVE" if the child is still running (it is +then killed), else "DEAD ". Interactive-only code paths (the +heartbeat-driven scheduler, slash commands) are only reachable this way. +""" +import json +import os +import pty +import re +import select +import signal +import sys +import time + +out_path, script_path, argv = sys.argv[1], sys.argv[2], sys.argv[3:] +script = json.load(open(script_path)) + +pid, fd = pty.fork() +if pid == 0: + signal.signal(signal.SIGPIPE, signal.SIG_DFL) + os.execvp(argv[0], argv) + +buf = b"" + + +def pump(secs): + """Drain output for up to secs; False once the child's side closed.""" + global buf + end = time.time() + secs + while time.time() < end: + r, _, _ = select.select([fd], [], [], max(0.0, min(0.2, end - time.time()))) + if r: + try: + data = os.read(fd, 65536) + except OSError: + return False + if not data: + return False + buf += data + return True + + +def text(): + return re.sub(rb"\x1b\[[0-9;?]*[A-Za-z]", b"", buf).decode("utf-8", "replace") + + +def file_has(path, needle): + try: + return needle in open(path, errors="replace").read() + except OSError: + return False + + +for step in script: + if "wait" in step: + pump(step["wait"]) + elif "send" in step: + os.write(fd, step["send"].encode()) + pump(0.3) + elif "until" in step: + end = time.time() + step.get("timeout", 20) + while time.time() < end and not file_has(step["until"], step.get("contains", "")): + pump(0.3) + elif "until_out" in step: + end = time.time() + step.get("timeout", 20) + while time.time() < end and step["until_out"] not in text(): + if not pump(0.3): + break + +try: + wpid, status = os.waitpid(pid, os.WNOHANG) +except ChildProcessError: + wpid, status = pid, -1 +open(out_path, "w").write(text()) +if wpid == 0: + print("ALIVE") + os.kill(pid, signal.SIGKILL) + os.waitpid(pid, 0) +else: + print("DEAD %d" % status) diff --git a/tests/integration/run.sh b/tests/integration/run.sh index 1699c1d..59b65c5 100755 --- a/tests/integration/run.sh +++ b/tests/integration/run.sh @@ -18,6 +18,40 @@ # T6 MCP safety boundary — MCP server cannot bypass hardline policy # T7 hook rewrite safety — rewritten args are checked before execution # T8 council boundary — read-only panel cannot execute shell commands +# T9 clean stdout — headless stdout is only the JSON line / answer +# T10 stale PWD — the real cwd, not $PWD, reaches the system prompt +# T11 untrusted project — ./.swarm-code.json can't run hooks, redirect +# the endpoint or loosen permissions unless the +# user lists the dir in trusted_projects +# T12 network gate — userinfo / uppercase-scheme / api-key bypasses +# are refused at startup; non-local providers[] +# and .profile_override endpoints at dial time +# T13 control files — write/edit can't touch ~/.swarm-code/hooks, +# schedule.json, .profile_override, or a case- +# variant .SSH; memory/ stays writable +# T14 headless 'ask' — an explicit "ask" or a dangerous command is +# denied headless unless HEADLESS_APPROVE=1 +# T15 request-body dump — never to /tmp; only SWARM_CODE_DEBUG=1, 0600 +# A* agents/MCP/scheduler — see tests/integration/agents_cases.sh +# T16 grep without a path — MCP server's stdin (the JSON-RPC stream) is never +# read by a tool subprocess; the next request survives +# T17 command-tool gate — `background` gets the same hardline gate as bash +# (denial names the pattern); a command merely +# MENTIONING reboot/halt still runs +# T18 truncated tool call — finish_reason=length: the cut-off write never runs +# T19 malformed args — cut mid-string, no finish_reason: strict check stops it +# T20 interrupted stream — tool calls of an ESC-interrupted stream never run +# T21 stale headless ok — a failed resumed run never reports the prior answer +# T22 small window — SWARM_CODE_MAX_TOKENS=32768 keeps a positive budget +# T23 /compact safety — no-op when nothing is old, merges summaries, 503 keeps all +# T24 mid-turn compaction — the live user request survives compaction +# T25 fatal 4xx — completed tool pairs survive; context overflow retries once +# T26 escapes round-trip — "
" / "\u003c" in args and prose, native + inband +# T27 profile override — env beats a stale override; kwargs kept; /model keeps profile +# T28 -p prompt parsing — -p --json "x", either order, "-x"/"-- x" prompts, stdin +# T29 swarm-code trust — the notice names it; after it, the repo's hook runs +# +# `run.sh t4 t11` or INTEG_ONLY="t4 t11" runs just those tests (default: all). # # Exit code: 0 iff every test passes. @@ -89,17 +123,39 @@ new_case() { # run_swarm — run the binary headless with the isolated env, # 90s watchdog (LLM retry backoff can stack up on a broken path). # Captures stdout/stderr into $CASE, sets RC. +# RUN_ENDPOINT endpoint URL to export (default: the mock); "-" exports +# none, so settings.json decides +# RUN_ENV extra VAR=value words, applied last — a bash array +# (RUN_ENV=(A=1 B=2)) or one string (RUN_ENV="A=1 B=2"); +# values must not contain spaces +# RUN_UNSET "VAR ..." removes defaults (e.g. SWARM_CODE_MODEL) +# RUN_STDIN file fed to stdin (default /dev/null) +# Opt-in knobs a developer may have exported are cleared first so the +# security cases below always see the defaults. +RUN_ENV=() run_swarm() { ( cd "$WORK" || exit 97 - HOME="$CASE_HOME" \ - SWARM_CODE_EXECUTION_CONTEXT="${RUN_EXECUTION_CONTEXT:-main}" \ - SWARM_CODE_ENDPOINT="http://127.0.0.1:$PORT" \ - SWARM_CODE_MODEL=test \ - SWARM_CODE_TOOL_FORMAT=native \ - SWARM_CODE_PLAN=off \ - SWARM_CODE_NO_RESUME=0 \ - "$BIN" "$@" "$CASE/stdout.txt" 2>"$CASE/stderr.txt" + unset SWARM_CODE_ENDPOINT SWARM_CODE_API_KEY SWARM_CODE_ALLOW_REMOTE \ + SWARM_CODE_PROVIDERS_JSON SWARM_CODE_FALLBACK_ENDPOINT \ + SWARM_CODE_HEADLESS_APPROVE SWARM_CODE_DEBUG + endpoint="${RUN_ENDPOINT:-http://127.0.0.1:$PORT}" + [ "$endpoint" = "-" ] && endpoint="" + export HOME="$CASE_HOME" \ + SWARM_CODE_EXECUTION_CONTEXT="${RUN_EXECUTION_CONTEXT:-main}" \ + SWARM_CODE_MODEL=test \ + SWARM_CODE_TOOL_FORMAT=native \ + SWARM_CODE_PLAN=off \ + SWARM_CODE_NO_RESUME=0 \ + PWD="${RUN_PWD:-$PWD}" + if [ -n "$endpoint" ]; then export SWARM_CODE_ENDPOINT="$endpoint"; fi + set -f # values like providers JSON contain [ ] — never glob them + # shellcheck disable=SC2068 + for kv in ${RUN_ENV[@]+${RUN_ENV[@]}}; do export "$kv"; done + # shellcheck disable=SC2086 + if [ -n "${RUN_UNSET:-}" ]; then unset $RUN_UNSET; fi + set +f + "$BIN" "$@" <"${RUN_STDIN:-/dev/null}" >"$CASE/stdout.txt" 2>"$CASE/stderr.txt" ) & local pid=$! ( sleep 90; kill -9 "$pid" 2>/dev/null ) & @@ -113,21 +169,63 @@ run_swarm() { # final_json — last {"status":...} line the binary printed. final_json() { grep '"status"' "$CASE/stdout.txt" | tail -1; } -# req_has — assert request #n to the mock contains the -# substring anywhere in its messages payload. Exit 0/1. +# req_has — assert streaming request #n (an agent turn) to +# the mock contains the substring anywhere in its messages payload. Exit 0/1. req_has() { python3 - "$REQLOG" "$1" "$2" <<'PYEOF' import json, sys path, n, needle = sys.argv[1], int(sys.argv[2]), sys.argv[3] for line in open(path): r = json.loads(line) - if r["n"] == n: + if r["n"] == n and r.get("kind", "stream") == "stream": sys.exit(0 if needle in json.dumps(r["body"].get("messages", [])) else 1) sys.exit(1) PYEOF } +# req_count — requests of every kind; stream_count / silent_count split +# agent turns from non-streaming ones (compaction's summarizer). req_count() { wc -l <"$REQLOG" | tr -d ' '; } +stream_count() { grep -c '"kind": "stream"' "$REQLOG"; } +silent_count() { grep -c '"kind": "silent"' "$REQLOG"; } + +# req_field — JSON of top-level field of streaming request #n. +req_field() { + python3 - "$REQLOG" "$1" "$2" <<'PYEOF' +import json, sys +path, n, key = sys.argv[1], int(sys.argv[2]), sys.argv[3] +for line in open(path): + r = json.loads(line) + if r["n"] == n and r.get("kind", "stream") == "stream": + print(json.dumps(r["body"].get(key), sort_keys=True)) +PYEOF +} + +# journal_file — the session journal .active points at (for resume checks). +journal_file() { cat "$CASE_HOME/.swarm-code/sessions/.active" 2>/dev/null; } + +# seed_journal [summary_text] — pre-write a resumable session: +# optional earlier-compaction summary, then n user/assistant pairs +# ("q1".."qN" / "a1".."aN"), and point .active at it. +seed_journal() { + local dir="$CASE_HOME/.swarm-code/sessions" + mkdir -p "$dir" + python3 - "$dir/journal-1000.jsonl" "$1" "${2:-}" <<'PYEOF' +import json, sys +path, n, summary = sys.argv[1], int(sys.argv[2]), sys.argv[3] +with open(path, "w") as f: + if summary: + f.write(json.dumps({"role": "assistant", + "content": "Summary of earlier conversation: " + summary}) + "\n") + for i in range(1, n + 1): + f.write(json.dumps({"role": "user", "content": "q%d" % i}) + "\n") + f.write(json.dumps({"role": "assistant", "content": "a%d" % i}) + "\n") +PYEOF + printf '%s' "$dir/journal-1000.jsonl" >"$dir/.active" +} + +# jcount — journal lines containing the substring. +jcount() { grep -c -- "$1" "$(journal_file)"; } # ------------------------------------------------------------ # T1 — plain prompt, final JSON line carries the scripted text @@ -345,17 +443,895 @@ EOF } # ------------------------------------------------------------ +# T9 — headless stdout carries only the result: with --json, exactly one +# JSON line even after a tool call; without it (stdout piped), exactly +# the final answer. The transcript goes to stderr. +# ------------------------------------------------------------ +t9() { + new_case t9 + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [ + {"type": "tool_calls", "calls": [ + {"id": "call_t9", "name": "bash", + "arguments": {"command": "echo transcript-t9"}}]}, + {"type": "text", "content": "RESULT_T9 **done**"} +]} +EOF + start_mock "$CASE/scenario.json" || { fail T9 "mock failed to start"; return; } + run_swarm -p "run it" --no-resume --json + cleanup + local lines; lines="$(wc -l <"$CASE/stdout.txt" | tr -d ' ')" + if [ "$RC" -ne 0 ]; then fail T9 "json: exit code $RC" + elif [ "$lines" -ne 1 ]; then fail T9 "json: stdout has $lines lines, want 1" + elif ! python3 -c 'import json,sys; d=json.load(open(sys.argv[1])); sys.exit(0 if d["summary"]=="RESULT_T9 **done**" else 1)' "$CASE/stdout.txt" + then fail T9 "json: stdout is not the result object: $(head -c 300 "$CASE/stdout.txt")" + elif ! grep -q "transcript-t9" "$CASE/stderr.txt"; then fail T9 "json: transcript missing from stderr" + else + start_mock "$CASE/scenario.json" || { fail T9 "mock failed to restart"; return; } + run_swarm -p "run it" --no-resume + cleanup + if [ "$RC" -ne 0 ]; then fail T9 "plain: exit code $RC" + elif [ "$(cat "$CASE/stdout.txt")" != "RESULT_T9 **done**" ]; then + fail T9 "plain: stdout is not just the answer: $(head -c 300 "$CASE/stdout.txt")" + else pass T9; fi + fi +} + +# ------------------------------------------------------------ +# T10 — a stale $PWD (launcher chdir'd without updating it) must not +# become the working directory the model is told about. +# ------------------------------------------------------------ +t10() { + new_case t10 + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [{"type": "text", "content": "CWD_OK_T10"}]} +EOF + start_mock "$CASE/scenario.json" || { fail T10 "mock failed to start"; return; } + RUN_PWD=/ run_swarm -p "where am i" --no-resume --json + cleanup + if [ "$RC" -ne 0 ]; then fail T10 "exit code $RC" + elif ! req_has 0 "Working directory: $WORK"; then fail T10 "system prompt does not name the real cwd $WORK" + else pass T10; fi +} + +# ------------------------------------------------------------ +# T11 — a cloned repo's ./.swarm-code.json is untrusted: its hooks never +# run, its endpoint/api_key never receive the prompt, and its +# permissions can only tighten. Listing the directory under +# trusted_projects in the user's settings restores the full file. +# ------------------------------------------------------------ +t11() { + new_case t11 + mkdir -p "$CASE_HOME/.swarm-code" + cat >"$CASE/scenario.json" <"$CASE_HOME/.swarm-code/settings.json" <"$WORK/.swarm-code.json" <"$WORK/.swarm-code.json" <"$CASE/scenario2.json" <<'EOF' +{"responses": [{"type": "text", "content": "TRUSTED_T11"}]} +EOF + start_mock "$CASE/scenario2.json" || { fail T11 "mock 2 failed to start"; return; } + cat >"$CASE_HOME/.swarm-code/settings.json" < — the "model" field of request #n to the mock. +req_model() { + python3 - "$REQLOG" "$1" <<'PYEOF' +import json, sys +for line in open(sys.argv[1]): + r = json.loads(line) + if r["n"] == int(sys.argv[2]): + print(r["body"].get("model", "")) +PYEOF +} + +# ------------------------------------------------------------ +# T12 — network isolation. 0.0.0.0 dials this machine on Linux/macOS, so +# it stands in for "a non-local host that actually answers": each +# refused case must leave the mock with no request at all. +# ------------------------------------------------------------ +t12() { + new_case t12 + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [{"type": "text", "content": "GATE_T12_A"}, + {"type": "text", "content": "GATE_T12_B"}]} +EOF + start_mock "$CASE/scenario.json" || { fail T12 "mock failed to start"; return; } + local url + for url in "http://127.0.0.1@0.0.0.0:$PORT" "http://localhost:1@0.0.0.0:$PORT" \ + "HTTP://0.0.0.0:$PORT" "http://127.0.0.1.x.invalid:$PORT"; do + RUN_ENDPOINT="$url" run_swarm -p "t12 startup" --no-resume --json + if [ "$RC" -ne 1 ] || [ "$(req_count)" -ne 0 ]; then + cleanup; fail T12 "startup gate let $url through (rc $RC, $(req_count) requests)"; return + fi + done + # An API key is not an opt-in to a remote host. + RUN_ENV=("SWARM_CODE_API_KEY=k") + RUN_ENDPOINT="http://0.0.0.0:$PORT" run_swarm -p "t12 key" --no-resume --json + RUN_ENV=() + if [ "$RC" -ne 1 ] || [ "$(req_count)" -ne 0 ]; then + cleanup; fail T12 "an api_key bypassed the gate (rc $RC)"; return + fi + # .profile_override is read at dial time — past the startup check. An + # override left by an earlier session loses to any env var that is set, + # so the harness's endpoint/model env is dropped and settings.json names + # a local (dead) endpoint that passes startup: the override then wins + # and is what the agent would dial, and the gate must refuse it. + mkdir -p "$CASE_HOME/.swarm-code" + printf '{"endpoint": "http://127.0.0.1:1", "model": "test"}\n' \ + >"$CASE_HOME/.swarm-code/settings.json" + printf '{"endpoint": "http://0.0.0.0:%s", "model": "evil-override"}\n' "$PORT" \ + >"$CASE_HOME/.swarm-code/.profile_override" + RUN_ENDPOINT=- RUN_UNSET="SWARM_CODE_MODEL" run_swarm -p "t12 override" --no-resume --json + rm -f "$CASE_HOME/.swarm-code/.profile_override" "$CASE_HOME/.swarm-code/settings.json" + if [ "$(req_count)" -ne 0 ]; then + cleanup; fail T12 "a non-local .profile_override endpoint was dialed"; return + fi + if ! grep -q "network isolation: refusing" "$CASE/stderr.txt"; then + cleanup; fail T12 "override: no refusal notice (was the override applied at all?)"; return + fi + # providers[]: the non-local first entry is refused, the local one used. + RUN_ENV=("SWARM_CODE_PROVIDERS_JSON=[{\"endpoint\":\"http://0.0.0.0:$PORT\",\"model\":\"evil-provider\"},{\"endpoint\":\"http://127.0.0.1:$PORT\"}]") + run_swarm -p "t12 providers" --no-resume --json + RUN_ENV=() + if [ "$RC" -ne 0 ]; then cleanup; fail T12 "providers: exit code $RC"; return; fi + if [ "$(req_count)" -ne 1 ] || [ "$(req_model 0)" != "test" ]; then + cleanup; fail T12 "providers: non-local provider dialed (model $(req_model 0))"; return + fi + if ! grep -q "network isolation: refusing" "$CASE/stderr.txt"; then + cleanup; fail T12 "providers: no refusal notice on stderr"; return + fi + # Control: with the explicit opt-in the same host IS reachable, so the + # refusals above were the gate, not a dead address. + RUN_ENV=("SWARM_CODE_ALLOW_REMOTE=1") + RUN_ENDPOINT="http://0.0.0.0:$PORT" run_swarm -p "t12 allowed" --no-resume --json + RUN_ENV=() + cleanup + if [ "$RC" -ne 0 ] || [ "$(req_count)" -ne 2 ]; then fail T12 "ALLOW_REMOTE=1 control failed (rc $RC)" + else pass T12; fi +} + +# ------------------------------------------------------------ +# T16 — grep with no `path` in MCP server mode: rg used to be run with no +# path and inherited stdin, so it searched the JSON-RPC stream — +# blocking until the client closed it and swallowing the NEXT request. +# ------------------------------------------------------------ +t16() { + new_case t16 + printf 'needle-t11 here\n' >"$WORK/a.txt" + local req1 req2 + req1='{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"grep","arguments":{"pattern":"needle-t11"}}}' + req2="{\"jsonrpc\":\"2.0\",\"id\":2,\"method\":\"tools/call\",\"params\":{\"name\":\"read\",\"arguments\":{\"path\":\"$WORK/a.txt\"}}}" + ( + cd "$WORK" || exit 97 + { printf '%s\n' "$req1"; sleep 1; printf '%s\n' "$req2"; sleep 2; } | + HOME="$CASE_HOME" perl -e 'alarm 40; exec @ARGV' "$BIN" --mcp-server \ + >"$CASE/stdout.txt" 2>"$CASE/stderr.txt" + ) + RC=$? + if [ "$RC" -ne 0 ]; then fail T16 "MCP server exit code $RC" + elif ! grep -q '"id":1' "$CASE/stdout.txt"; then fail T16 "no response to the grep request" + elif ! grep '"id":1' "$CASE/stdout.txt" | grep -q 'a.txt:1:needle-t11'; then + fail T16 "grep searched the wrong input: $(grep '"id":1' "$CASE/stdout.txt" | head -c 300)" + elif ! grep -q '"id":2' "$CASE/stdout.txt"; then fail T16 "the second JSON-RPC request was swallowed" + else pass T16; fi +} + +# ------------------------------------------------------------ +# T17 — every command-running tool shares the classifier: `background` +# used to run a hardline command ungated. And the classifier is +# token-aware: words inside quoted args / grep patterns don't trip it. +# ------------------------------------------------------------ +t17() { + new_case t17 + local sentinel="$WORK/owned-by-background" + cat >"$CASE/scenario.json" <"$sc/settings.json" + cat >"$CASE/scenario.json" <"$CASE_HOME/.swarm-code/settings.json" + cat >"$CASE/scenario.json" <"$CASE/scenario2.json" <<'EOF' +{"responses": [ + {"type": "tool_calls", "calls": [ + {"id": "call_rm", "name": "bash", "arguments": {"command": "rm -rf ~/victim"}}]}, + {"type": "text", "content": "DANGER_DENIED_T14"} +]} +EOF + start_mock "$CASE/scenario2.json" || { fail T14 "mock 2 failed to start"; return; } + run_swarm -p "t14 danger" --no-resume --json + cleanup + if [ ! -d "$CASE_HOME/victim" ]; then fail T14 "headless auto-approved a dangerous rm -rf ~"; return + elif ! req_has 1 "permission denied"; then fail T14 "dangerous command was not denied"; return + fi + # The opt-in restores auto-approval. + echo '{"permissions": {"bash": "ask"}}' >"$CASE_HOME/.swarm-code/settings.json" + start_mock "$CASE/scenario.json" || { fail T14 "mock 3 failed to start"; return; } + RUN_ENV=("SWARM_CODE_HEADLESS_APPROVE=1") + run_swarm -p "t14 approved" --no-resume --json + RUN_ENV=() + cleanup + if [ ! -e "$sentinel" ]; then fail T14 "SWARM_CODE_HEADLESS_APPROVE=1 did not approve the ask" + elif [ "$RC" -ne 0 ]; then fail T14 "approved: exit code $RC" + else pass T14; fi +} + +# ------------------------------------------------------------ +# T15 — the request body (prompts, tool output, secrets) is never dumped +# to a shared world-readable path. Only SWARM_CODE_DEBUG=1 writes it, +# to ~/.swarm-code/last-body.json with mode 600. /tmp is shared with +# other runs, so the check is "OUR marker is absent", not "no file". +# ------------------------------------------------------------ +file_mode() { python3 -c 'import os,sys; print(oct(os.stat(sys.argv[1]).st_mode & 0o777))' "$1"; } + +t15() { + new_case t15 + local marker="T15_BODY_MARKER_$$_$RANDOM" + local dump="$CASE_HOME/.swarm-code/last-body.json" + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [{"type": "text", "content": "BODY_T15_A"}, + {"type": "text", "content": "BODY_T15_B"}, + {"type": "text", "content": "BODY_T15_C"}]} +EOF + start_mock "$CASE/scenario.json" || { fail T15 "mock failed to start"; return; } + local fmt + for fmt in inband native; do + RUN_ENV=("SWARM_CODE_TOOL_FORMAT=$fmt") + run_swarm -p "t15 $fmt $marker" --no-resume --json + RUN_ENV=() + if [ "$RC" -ne 0 ]; then cleanup; fail T15 "$fmt: exit code $RC"; return; fi + if grep -qs "$marker" /tmp/swarm-code-last-body.json; then + cleanup; fail T15 "$fmt: request body written to /tmp/swarm-code-last-body.json"; return + fi + if [ -e "$dump" ]; then cleanup; fail T15 "$fmt: body dumped without SWARM_CODE_DEBUG"; return; fi + done + RUN_ENV=("SWARM_CODE_TOOL_FORMAT=inband" "SWARM_CODE_DEBUG=1") + run_swarm -p "t15 debug $marker" --no-resume --json + RUN_ENV=() + cleanup + if [ "$RC" -ne 0 ]; then fail T15 "debug: exit code $RC" + elif ! grep -qs "$marker" "$dump"; then fail T15 "SWARM_CODE_DEBUG=1 did not write $dump" + elif [ "$(file_mode "$dump")" != "0o600" ]; then fail T15 "debug dump is mode $(file_mode "$dump"), want 600" + elif grep -qs "$marker" /tmp/swarm-code-last-body.json; then fail T15 "debug: body also written to /tmp" + else pass T15; fi +} + +# ------------------------------------------------------------ + +# ------------------------------------------------------------ +# T18 — a tool call cut off at the output-token limit (finish_reason +# "length", arguments ending mid-string) is NOT run: the file is never +# written, the model is told why and asked to reissue, and history +# keeps "{}" instead of the cut-off blob. +# ------------------------------------------------------------ +t18() { + new_case t18 + python3 - "$CASE/scenario.json" "$WORK" <<'PYEOF' +import json, sys +out, work = sys.argv[1], sys.argv[2] +full = json.dumps({"path": work + "/config.py", + "content": "SETTINGS = {\n 'debug': False,\n 'db_url': 'postgres://prod-db/app'\n}\n"}) +cut = full[:full.index("postgres://prod") + len("postgres://prod")] +json.dump({"responses": [ + {"type": "tool_calls", "finish": "length", + "calls": [{"id": "call_cut", "name": "write", "arguments": cut}]}, + {"type": "text", "content": "REISSUE_ACK_T11"}]}, open(out, "w")) +PYEOF + start_mock "$CASE/scenario.json" || { fail T18 "mock failed to start"; return; } + run_swarm -p "write the config" --no-resume --json + cleanup + local out; out="$(final_json)" + if [ -e "$WORK/config.py" ]; then fail T18 "truncated write was executed: $(cat "$WORK/config.py")" + elif [ "$RC" -ne 0 ]; then fail T18 "exit code $RC" + elif ! req_has 1 "error: not executed"; then fail T18 "model never told the cut-off call did not run" + elif ! req_has 1 "none of them ran"; then fail T18 "truncation nudge missing" + elif req_has 1 "postgres://prod"; then fail T18 "cut-off arguments were sent back verbatim" + elif ! echo "$out" | grep -q "REISSUE_ACK_T11"; then fail T18 "final text missing: $out" + else pass T18; fi +} + +# ------------------------------------------------------------ +# T19 — arguments cut mid-string WITHOUT a length finish_reason (the lenient +# json_decode accepts them) are caught by the strict JSON check. +# ------------------------------------------------------------ +t19() { + new_case t19 + local sentinel="$WORK/SENTINEL_T12" + python3 - "$CASE/scenario.json" "$sentinel" <<'PYEOF' +import json, sys +out, sentinel = sys.argv[1], sys.argv[2] +json.dump({"responses": [ + {"type": "tool_calls", + "calls": [{"id": "call_m", "name": "bash", + "arguments": '{"command": "touch ' + sentinel}]}, + {"type": "text", "content": "MALFORMED_ACK_T12"}]}, open(out, "w")) +PYEOF + start_mock "$CASE/scenario.json" || { fail T19 "mock failed to start"; return; } + run_swarm -p "touch it" --no-resume --json + cleanup + if [ -e "$sentinel" ]; then fail T19 "malformed (cut) command was executed" + elif [ "$RC" -ne 0 ]; then fail T19 "exit code $RC" + elif ! req_has 1 "were not valid JSON"; then fail T19 "model never told the arguments were malformed" + else pass T19; fi +} + +# ------------------------------------------------------------ +# T20 — a stream the user interrupted (the runtime's "[Request interrupted +# by user]" marker) never runs the tool calls it carried, and the turn +# ends without calling the model again. +# ------------------------------------------------------------ +t20() { + new_case t20 + local sentinel="$WORK/SENTINEL_T13" + python3 - "$CASE/scenario.json" "$sentinel" <<'PYEOF' +import json, sys +out, sentinel = sys.argv[1], sys.argv[2] +json.dump({"responses": [ + {"type": "tool_calls", "content": "Creating the file.\n\n[Request interrupted by user]", + "calls": [{"id": "call_i", "name": "bash", "arguments": {"command": "touch " + sentinel}}]}, + {"type": "text", "content": "SHOULD_NOT_BE_CALLED_T13"}]}, open(out, "w")) +PYEOF + start_mock "$CASE/scenario.json" || { fail T20 "mock failed to start"; return; } + run_swarm -p "make the file" --json + cleanup + local journal; journal="$(journal_file)" + if [ -e "$sentinel" ]; then fail T20 "tool call from an interrupted stream was executed" + elif [ "$(req_count)" -ne 1 ]; then fail T20 "expected 1 request (turn ends), got $(req_count)" + elif ! grep -q 'call_i' "$journal" || ! grep -q '\[interrupted\]' "$journal"; then + fail T20 "journal lacks the [interrupted] result for the call" + else pass T20; fi +} + +# ------------------------------------------------------------ +# T21 — headless reports only THIS run's answer. Resume is the default, so +# run 2's history holds run 1's reply; a run 2 whose request fails +# (HTTP 400, or no server at all) must be status error / exit 1, not +# run 1's answer with status ok. +# ------------------------------------------------------------ +t21() { + new_case t21 + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [{"type": "text", "content": "FIRST_RUN_ANSWER_T14"}]} +EOF + start_mock "$CASE/scenario.json" || { fail T21 "mock failed to start"; return; } + run_swarm -p "t21 first" --json + cleanup + if [ "$RC" -ne 0 ] || ! final_json | grep -q FIRST_RUN_ANSWER_T14; then + fail T21 "run 1 did not succeed: rc=$RC $(final_json)"; return + fi + + cat >"$CASE/scenario2.json" <<'EOF' +{"responses": [{"type": "http", "status": 400, "body": "{\"error\":{\"message\":\"bad request\"}}"}]} +EOF + start_mock "$CASE/scenario2.json" || { fail T21 "mock 2 failed to start"; return; } + run_swarm -p "t21 second" --json + cleanup + local out2; out2="$(final_json)" + if [ "$RC" -eq 0 ]; then fail T21 "run 2 (HTTP 400) exited 0: $out2"; return; fi + if echo "$out2" | grep -q FIRST_RUN_ANSWER_T14; then fail T21 "run 2 (HTTP 400) reported run 1's answer: $out2"; return; fi + if ! echo "$out2" | grep -q '"status":"error"'; then fail T21 "run 2 (HTTP 400) status not error: $out2"; return; fi + + # Run 3: nothing listening on the port any more. + run_swarm -p "t21 third" --json + local out3; out3="$(final_json)" + if [ "$RC" -eq 0 ]; then fail T21 "run 3 (server down) exited 0: $out3" + elif echo "$out3" | grep -q FIRST_RUN_ANSWER_T14; then fail T21 "run 3 (server down) reported run 1's answer: $out3" + elif ! echo "$out3" | grep -q '"status":"error"'; then fail T21 "run 3 status not error: $out3" + else pass T21; fi +} + +# ------------------------------------------------------------ +# T22 — a small context window (SWARM_CODE_MAX_TOKENS=32768) gets a positive, +# window-scaled budget (16384): no "compacting" on every step, and the +# context meter reads x/16k (the old budget was -35616). +# ------------------------------------------------------------ +t22() { + new_case t22 + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [ + {"type": "tool_calls", "calls": [ + {"id": "call_t15", "name": "bash", "arguments": {"command": "echo small-window-t22"}}]}, + {"type": "text", "content": "SMALL_WINDOW_OK_T15"} +]} +EOF + start_mock "$CASE/scenario.json" || { fail T22 "mock failed to start"; return; } + RUN_ENV="SWARM_CODE_MAX_TOKENS=32768" run_swarm -p "t22 run it" --no-resume --json + cleanup + if [ "$RC" -ne 0 ]; then fail T22 "exit code $RC" + elif grep -q "compacting" "$CASE/stderr.txt"; then + fail T22 "compacted with a tiny history: $(grep compacting "$CASE/stderr.txt" | head -1)" + elif ! req_has 0 "/16k tok"; then fail T22 "context meter does not show the 16k budget" + elif ! final_json | grep -q SMALL_WINDOW_OK_T15; then fail T22 "final text missing: $(final_json)" + else pass T22; fi +} + +# ------------------------------------------------------------ +# T23 — /compact never loses history: with nothing old enough to summarize +# it is a no-op (no LLM call, no extra summary); an earlier summary is +# merged into the new one, not stacked; a failed summarizer (503) +# leaves the history untouched instead of eliding it. +# ------------------------------------------------------------ +t23() { + new_case t23 + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [], "silent": [{"type": "text", "content": "SHOULD_NOT_BE_CALLED"}]} +EOF + seed_journal 6 + start_mock "$CASE/scenario.json" || { fail T23 "mock failed to start"; return; } + run_swarm -p "/compact" --json + cleanup + if [ "$RC" -ne 0 ]; then fail T23 "noop: exit code $RC"; return; fi + if [ "$(silent_count)" -ne 0 ]; then fail T23 "noop: summarizer called for 12 messages"; return; fi + if [ "$(jcount 'Summary of earlier')" -ne 0 ]; then fail T23 "noop: a summary was prepended"; return; fi + if [ "$(wc -l <"$(journal_file)")" -ne 12 ]; then fail T23 "noop: journal changed ($(wc -l <"$(journal_file)") lines)"; return; fi + + new_case t16b + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [], "silent": [{"type": "text", "content": "NEW_MERGED_SUMMARY_T16"}]} +EOF + seed_journal 15 "OLD_SUMMARY_FACT_T16" + start_mock "$CASE/scenario.json" || { fail T23 "merge: mock failed to start"; return; } + run_swarm -p "/compact" --json + cleanup + if [ "$RC" -ne 0 ]; then fail T23 "merge: exit code $RC"; return; fi + if [ "$(silent_count)" -ne 1 ]; then fail T23 "merge: expected 1 summarizer call, got $(silent_count)"; return; fi + if ! grep -q OLD_SUMMARY_FACT_T16 "$REQLOG"; then fail T23 "merge: earlier summary not passed to the summarizer"; return; fi + if [ "$(jcount 'Summary of earlier')" -ne 1 ]; then fail T23 "merge: $(jcount 'Summary of earlier') summaries in the journal (stacked?)"; return; fi + if [ "$(jcount NEW_MERGED_SUMMARY_T16)" -ne 1 ]; then fail T23 "merge: new summary not journaled"; return; fi + if [ "$(jcount '"q15"')" -ne 1 ]; then fail T23 "merge: most recent user message not kept verbatim"; return; fi + + new_case t16c + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [], "silent": [ + {"type": "http", "status": 503, "body": "{\"error\":{\"message\":\"overloaded\"}}"}, + {"type": "http", "status": 503, "body": "{\"error\":{\"message\":\"overloaded\"}}"}, + {"type": "http", "status": 503, "body": "{\"error\":{\"message\":\"overloaded\"}}"}, + {"type": "http", "status": 503, "body": "{\"error\":{\"message\":\"overloaded\"}}"}]} +EOF + seed_journal 15 + start_mock "$CASE/scenario.json" || { fail T23 "fail: mock failed to start"; return; } + run_swarm -p "/compact" --json + cleanup + # (>= 1: the runtime may retry a 5xx at the curl level.) + if [ "$(silent_count)" -lt 1 ]; then fail T23 "fail: summarizer never called" + elif [ "$(wc -l <"$(journal_file)")" -ne 30 ]; then fail T23 "fail: 503 summarizer lost messages ($(wc -l <"$(journal_file)") of 30 left)" + elif [ "$(jcount 'compaction failed')" -ne 0 ]; then fail T23 "fail: journaled a compaction-failed placeholder" + else pass T23; fi +} + +# ------------------------------------------------------------ +# T24 — compaction mid-turn keeps the live request. 12 tool rounds of 7000 +# chars against a 32K window cross the budget at ~round 9, when the +# user message is more than 16 messages back: every request must +# still carry it (it used to be summarized away, leaving requests with +# no user message); only the pre-turn history is summarized. +# ------------------------------------------------------------ +t24() { + new_case t24 + python3 - "$CASE/scenario.json" <<'PYEOF' +import json, sys +# Distinct commands: the guardrail stops identical repeated calls. +calls = [{"type": "tool_calls", "calls": [{"id": "c%d" % i, "name": "bash", + "arguments": {"command": "head -c 7000 /dev/zero | tr '\\0' A; echo round-%d" % i}}]} + for i in range(12)] +json.dump({"responses": calls + [{"type": "text", "content": "LONG_TURN_DONE_T17"}], + "silent": [{"type": "text", "content": "SUMMARY_OF_OLD_T17"}]}, + open(sys.argv[1], "w")) +PYEOF + seed_journal 5 + start_mock "$CASE/scenario.json" || { fail T24 "mock failed to start"; return; } + RUN_ENV="SWARM_CODE_MAX_TOKENS=32768" run_swarm -p "LIVE_REQUEST_T17 run the dozen commands" --json + cleanup + local verdict + verdict="$(python3 - "$REQLOG" <<'PYEOF' +import json, sys +seen_silent, after, missing = False, 0, [] +for line in open(sys.argv[1]): + r = json.loads(line) + if r["kind"] == "silent": + seen_silent = True + continue + msgs = json.dumps(r["body"]["messages"]) + if "LIVE_REQUEST_T17" not in msgs: + missing.append(r["n"]) + if seen_silent and "SUMMARY_OF_OLD_T17" in msgs: + after += 1 +if not seen_silent: + print("no compaction happened") +elif missing: + print("requests without the live user message: %s" % missing) +elif after == 0: + print("no request carried the summary after compaction") +else: + print("ok") +PYEOF +)" + if [ "$RC" -ne 0 ]; then fail T24 "exit code $RC" + elif [ "$verdict" != "ok" ]; then fail T24 "$verdict" + elif [ "$(silent_count)" -ne 1 ]; then fail T24 "expected 1 summarizer call, got $(silent_count)" + elif ! final_json | grep -q LONG_TURN_DONE_T17; then fail T24 "final text missing: $(final_json)" + else pass T24; fi +} + +# ------------------------------------------------------------ +# T25 — a fatal 4xx keeps the turn's completed work. +# a) "maximum context length" after a big tool result: trimmed and retried +# once — the retry carries the stub, the turn succeeds; +# b) a plain 400 after two writes: the journal keeps both write results +# (it used to keep only the user message), and a resumed run sends them; +# c) the overflow retry happens once, not in a loop. +# ------------------------------------------------------------ +t25() { + new_case t25 + python3 - "$CASE/scenario.json" "$WORK" <<'PYEOF' +import json, sys +out, work = sys.argv[1], sys.argv[2] +over = {"type": "http", "status": 400, + "body": json.dumps({"error": {"message": "This model's maximum context length is 8192 tokens. However, you requested 9000 tokens."}})} +json.dump({"responses": [ + {"type": "tool_calls", "calls": [{"id": "call_big", "name": "bash", + "arguments": {"command": "head -c 12000 /dev/zero | tr '\\0' B"}}]}, + {"type": "tool_calls", "calls": [{"id": "call_w", "name": "write", + "arguments": {"path": work + "/a.txt", "content": "A"}}]}, + over, + {"type": "text", "content": "RECOVERED_T18"}]}, open(out, "w")) +PYEOF + start_mock "$CASE/scenario.json" || { fail T25 "a: mock failed to start"; return; } + run_swarm -p "t25 build it" --no-resume --json + cleanup + if [ "$RC" -ne 0 ]; then fail T25 "a: exit code $RC: $(final_json)"; return; fi + if ! final_json | grep -q RECOVERED_T18; then fail T25 "a: no recovery: $(final_json)"; return; fi + if ! req_has 3 "chars elided"; then fail T25 "a: retry did not carry the trimmed result"; return; fi + if req_has 3 "BBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBB"; then fail T25 "a: retry still carried the 12KB result"; return; fi + if ! req_has 3 "call_w"; then fail T25 "a: retry lost the completed write"; return; fi + + new_case t18b + python3 - "$CASE/scenario.json" "$WORK" <<'PYEOF' +import json, sys +out, work = sys.argv[1], sys.argv[2] +json.dump({"responses": [ + {"type": "tool_calls", "calls": [{"id": "call_w1", "name": "write", + "arguments": {"path": work + "/a.txt", "content": "A"}}]}, + {"type": "tool_calls", "calls": [{"id": "call_w2", "name": "write", + "arguments": {"path": work + "/b.txt", "content": "B"}}]}, + {"type": "http", "status": 400, "body": "{\"error\":{\"message\":\"bad request\"}}"}]}, + open(out, "w")) +PYEOF + start_mock "$CASE/scenario.json" || { fail T25 "b: mock failed to start"; return; } + run_swarm -p "t25 write two files" --json + cleanup + local journal; journal="$(journal_file)" + if [ "$RC" -eq 0 ]; then fail T25 "b: exit 0 after a fatal 400"; return; fi + if [ ! -f "$WORK/a.txt" ] || [ ! -f "$WORK/b.txt" ]; then fail T25 "b: writes did not land"; return; fi + if [ "$(grep -c '"tool_call_id":"call_w[12]"' "$journal")" -ne 2 ]; then + fail T25 "b: journal lost the completed writes: $(cat "$journal")"; return + fi + cat >"$CASE/scenario2.json" <<'EOF' +{"responses": [{"type": "text", "content": "RESUMED_T18"}]} +EOF + start_mock "$CASE/scenario2.json" || { fail T25 "b: mock 2 failed to start"; return; } + run_swarm -p "t25 what did you write" --json + cleanup + if ! req_has 0 "call_w1" || ! req_has 0 "call_w2"; then fail T25 "b: resumed request has no record of the writes"; return; fi + + new_case t18c + python3 - "$CASE/scenario.json" <<'PYEOF' +import json, sys +over = {"type": "http", "status": 400, + "body": json.dumps({"error": {"message": "context_length_exceeded"}})} +json.dump({"responses": [ + {"type": "tool_calls", "calls": [{"id": "call_big", "name": "bash", + "arguments": {"command": "head -c 12000 /dev/zero | tr '\\0' B"}}]}, + over, over, over, {"type": "text", "content": "SHOULD_NOT_GET_HERE"}]}, open(sys.argv[1], "w")) +PYEOF + start_mock "$CASE/scenario.json" || { fail T25 "c: mock failed to start"; return; } + run_swarm -p "t25 loop" --no-resume --json + cleanup + if [ "$RC" -eq 0 ]; then fail T25 "c: exit 0 after repeated overflow" + elif [ "$(req_count)" -ne 3 ]; then fail T25 "c: expected 3 requests (one retry), got $(req_count)" + else pass T25; fi +} + +# ------------------------------------------------------------ +# T26 — text round-trips byte for byte: "
", "<" and a literal < +# (JS source) in tool arguments and prose, native and inband. A +# u003c -> "<" "repair" pass (for a long-fixed runtime bug) turned +# "<" into "\<" — in files the model wrote and in its prose. +# ------------------------------------------------------------ +t26() { + new_case t26 + python3 - "$CASE" "$WORK" <<'PYEOF' +import json, sys +case, work = sys.argv[1], sys.argv[2] +# The file the model means to write: markup, a bare "<", and a JS < +# escape that must land as the six characters backslash-u-0-0-3-c. +want = 's = "
";\nlt = "<";\njs = "\\u003cp\\u003e";\n' +open(case + "/want.txt", "w").write(want) +args = json.dumps({"path": work + "/esc.js", "content": want}) +# Some models JSON-escape "<" as < inside the arguments: decoded once +# by the argument parse, it must become a plain "<". +args_escaped = args.replace('"
"', '"\\u003cdiv\\u003e"') +prose = 'Use
, not \\u003cdiv\\u003e, when the text says "<".' +open(case + "/prose.txt", "w").write(prose) +json.dump({"responses": [ + {"type": "tool_calls", "calls": [{"id": "call_esc", "name": "write", "arguments": args_escaped}]}, + {"type": "text", "content": prose}]}, open(case + "/native.json", "w")) +json.dump({"responses": [ + {"type": "text", "content": "Writing it.\ncall:write" + args_escaped}, + {"type": "text", "content": prose}]}, open(case + "/inband.json", "w")) +PYEOF + local fmt + for fmt in native inband; do + rm -f "$WORK/esc.js" + start_mock "$CASE/$fmt.json" || { fail T26 "$fmt: mock failed to start"; return; } + RUN_ENV="SWARM_CODE_TOOL_FORMAT=$fmt" run_swarm -p "write esc.js" --no-resume --json + cleanup + if [ "$RC" -ne 0 ]; then fail T26 "$fmt: exit code $RC"; return; fi + if ! cmp -s "$WORK/esc.js" "$CASE/want.txt"; then + fail T26 "$fmt: file content changed: $(cat "$WORK/esc.js" 2>/dev/null)"; return + fi + if ! python3 -c 'import json,sys; d=json.loads(open(sys.argv[1]).read().strip().splitlines()[-1]); sys.exit(0 if d["summary"]==open(sys.argv[2]).read() else 1)' \ + "$CASE/stdout.txt" "$CASE/prose.txt"; then + fail T26 "$fmt: prose changed: $(final_json)"; return + fi + done + pass T26 +} + +# ------------------------------------------------------------ +# T27 — the persisted /profile override: an env var that is set beats one +# left by an earlier session (a stale profile model was sent instead +# of SWARM_CODE_MODEL); the profile's chat_template_kwargs reach the +# request (they were dropped); /model changes only the model and +# keeps the profile's settings (it overwrote them). +# ------------------------------------------------------------ +t27() { + new_case t27 + mkdir -p "$CASE_HOME/.swarm-code" + cat >"$CASE_HOME/.swarm-code/settings.json" <<'EOF' +{"profiles": {"qwen2": {"model": "qwen-2-stale-model", "endpoint": "http://127.0.0.1:9", + "chat_template_kwargs": {"enable_thinking": false}}}} +EOF + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [{"type": "text", "content": "PROFILE_OK_T20"}]} +EOF + start_mock "$CASE/scenario.json" || { fail T27 "mock failed to start"; return; } + run_swarm -p "/profile qwen2" --no-resume --json + if [ "$RC" -ne 0 ]; then cleanup; fail T27 "/profile run exit code $RC"; return; fi + run_swarm -p "t27 hello" --no-resume --json + cleanup + local model ct + model="$(req_field 0 model)"; ct="$(req_field 0 chat_template_kwargs)" + if [ "$RC" -ne 0 ]; then fail T27 "run after /profile: exit $RC (endpoint override beat env?)"; return; fi + if [ "$model" != '"test"' ]; then fail T27 "stale override beat SWARM_CODE_MODEL: model=$model"; return; fi + if [ "$ct" != '{"enable_thinking": false}' ]; then fail T27 "profile chat_template_kwargs lost: $ct"; return; fi + + start_mock "$CASE/scenario.json" || { fail T27 "mock 2 failed to start"; return; } + run_swarm -p "/model other-model" --no-resume --json + run_swarm -p "t27 again" --no-resume --json + cleanup + ct="$(req_field 0 chat_template_kwargs)" + if [ "$RC" -ne 0 ]; then fail T27 "run after /model: exit $RC" + elif [ "$ct" != '{"enable_thinking": false}' ]; then fail T27 "/model wiped the profile's kwargs: $ct" + elif [ "$(req_field 0 model)" != '"test"' ]; then fail T27 "after /model: env model not sent: $(req_field 0 model)" + else pass T27; fi +} + +# ------------------------------------------------------------ +# T28 — the -p prompt: README's `-p --json "prompt"` (a flag right after +# -p used to mean "read stdin" — 0 requests, exit 1), either flag +# order, prompts that START with "-" (Scheduler/Flows pass user text +# right after -p; "-- summarize…" was dropped), and stdin via `-p -` +# or no positional argument. +# ------------------------------------------------------------ +t28() { + new_case t28 + python3 - "$CASE/scenario.json" <<'PYEOF' +import json, sys +json.dump({"responses": [{"type": "text", "content": "OK_T21_%d" % i} for i in range(6)]}, + open(sys.argv[1], "w")) +PYEOF + printf 't28 from stdin\n' >"$CASE/stdin1.txt" + printf 't28 stdin fallback\n' >"$CASE/stdin2.txt" + start_mock "$CASE/scenario.json" || { fail T28 "mock failed to start"; return; } + local n=0 why="" + t28_case() { # + local want="$1"; shift + run_swarm "$@" + if [ "$RC" -ne 0 ]; then why="[$*] exit $RC"; return 1; fi + if ! req_has "$n" "$want"; then why="[$*] request $n lacks prompt '$want'"; return 1; fi + if ! grep -q "OK_T21_$n" "$CASE/stdout.txt"; then why="[$*] stdout lacks the answer"; return 1; fi + n=$((n + 1)) + } + t28_case "t28 list the test files" -p --json "t28 list the test files" --no-resume && + t28_case "t28 other order" --no-resume --json -p "t28 other order" && + t28_case "-- summarize the diff t28" --no-resume -p "-- summarize the diff t28" && + t28_case "-x t28 dash prompt" -p "-x t28 dash prompt" --json --no-resume && + RUN_STDIN="$CASE/stdin1.txt" t28_case "t28 from stdin" --no-resume -p - --json && + RUN_STDIN="$CASE/stdin2.txt" t28_case "t28 stdin fallback" -p --json --no-resume + local ok=$? + cleanup + if [ "$ok" -ne 0 ]; then fail T28 "$why"; else pass T28; fi +} + +# ------------------------------------------------------------ + +# ------------------------------------------------------------ +# T29 — `swarm-code trust`: an untrusted repo's hook doesn't run and the +# notice names the command; after `swarm-code trust` in that +# directory the same file applies in full (the hook runs). +# ------------------------------------------------------------ +t29() { + new_case t29 + cat >"$CASE/scenario.json" <<'EOF' +{"responses": [{"type": "text", "content": "TRUST_T29_A"}, + {"type": "text", "content": "TRUST_T29_B"}]} +EOF + start_mock "$CASE/scenario.json" || { fail T29 "mock failed to start"; return; } + cat >"$WORK/.swarm-code.json" <"$CASE/trust.txt" 2>&1 ) + if ! grep -q "^trusted " "$CASE/trust.txt"; then + cleanup; fail T29 "swarm-code trust failed: $(head -c 200 "$CASE/trust.txt")"; return + fi + run_swarm -p "t29 trusted" --no-resume --json + cleanup + if [ ! -e "$WORK/HOOK_RAN" ]; then fail T29 "trusted repo's hook did not run" + elif ! final_json | grep -q "TRUST_T29_B"; then fail T29 "second run failed: $(final_json)" + else pass T29; fi +} + +# Multi-agent / MCP / scheduler / persistence cases (A1..). +. "$ROOT/tests/integration/agents_cases.sh" echo "integration: binary $BIN" echo "integration: scratch $TMP" -t1 -t2 -t3 -t4 -t5 -t6 -t7 -t8 +# `run.sh t4 t11` (or INTEG_ONLY="t4 t11") runs just those cases; no +# arguments runs them all. +ALL_TESTS="t1 t2 t3 t4 t5 t6 t7 t8 t9 t10 t11 t12 t13 t14 t15 t16 t17 t18 t19 t20 t21 t22 t23 t24 t25 t26 t27 t28 t29 agents_cases" +for t in ${*:-${INTEG_ONLY:-$ALL_TESTS}}; do "$t"; done echo "----------------------------------------" echo "integration: $PASS passed, $FAIL failed" diff --git a/tests/smoke.sh b/tests/smoke.sh index f4c8a03..a9267ed 100755 --- a/tests/smoke.sh +++ b/tests/smoke.sh @@ -54,6 +54,14 @@ case "$MCP_OUT" in *"$ESC"*) FAIL "--mcp-server leaked ANSI onto stdout (execution_context feedback gate broken)" ;; esac +# 6. --mcp-server answers MCP ping with an EMPTY result (spec), not -32601. +PING_OUT=$(printf '%s\n' '{"jsonrpc":"2.0","id":9,"method":"ping"}' \ + | $BIN --mcp-server 2>/dev/null) || true +case "$PING_OUT" in + *'"id":9,"result":{}'*) ;; + *) FAIL "--mcp-server ping did not return an empty result: $PING_OUT" ;; +esac + rm -rf "$EMPTY_HOME" trap - EXIT INT TERM echo "smoke ok"