Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
44 commits
Select commit Hold shift + click to select a range
0d1b3d5
fix: clean headless stdout, real cwd, ESC ends the turn, read/edit co…
claude Sep 24, 2026
58f1538
fix(bash): newline-delimited wrapper so comments/heredocs don't break it
claude Sep 24, 2026
7a4f3d6
fix(config): treat a repo's ./.swarm-code.json as untrusted
claude Sep 24, 2026
6152645
fix(grep,glob): never read stdin; surface regex errors to the model
claude Sep 24, 2026
df50ebb
fix(net): parse endpoint URLs properly and gate every LLM dial
claude Sep 24, 2026
c175a5f
fix(agent): never run tool calls from a truncated or interrupted turn
claude Sep 24, 2026
578f446
fix(mcp): structured call status, EOF = connection lost, spec-correct…
claude Sep 24, 2026
9e8a321
fix(pathguard): protect swarm-code's control files, match case-insens…
claude Sep 24, 2026
b8ce347
fix(mcp-server): ping, unknown-tool code, and JSON-RPC envelope valid…
claude Sep 24, 2026
91f945f
fix(agent): headless mode no longer auto-approves every 'ask'
claude Sep 24, 2026
99b5ff8
fix(headless): report only this run's answer, never a resumed one
claude Sep 24, 2026
ac15802
fix(llm): dump the request body only with SWARM_CODE_DEBUG=1, mode 600
claude Sep 24, 2026
bb5cba5
fix(file_watch): no shell injection, portable mtime; clamp wait timeouts
claude Sep 24, 2026
d3a0766
docs: changelog for the security fixes; doctor flags a gated endpoint
claude Sep 24, 2026
ef60570
Merge wip/sec: project-config trust, network gate, control files, hea…
claude Sep 24, 2026
0cfdd81
fix(scheduler): validate schedule.json, never clobber it, deny danger…
claude Sep 24, 2026
bec9027
fix(permissions): token-aware command classifier for every shell tool
claude Sep 24, 2026
49dbf0b
fix(budget): scale the output reserve and compact buffer to the window
claude Sep 24, 2026
cca0f5e
fix(compact): keep the live turn, never lose history on a failed summary
claude Sep 24, 2026
9fb0209
fix(background): per-session private task directory
claude Sep 24, 2026
d70396d
fix(subagent): isolated guardrails, enforced type allow-lists, bounde…
claude Sep 24, 2026
766404c
fix(agent): a fatal 4xx keeps completed work; retry context overflow …
claude Sep 24, 2026
71e2c15
fix(hooks): pass payloads via a 0600 temp file/stdin; fail closed
claude Sep 24, 2026
344f8e5
fix(flows): validate workflow shape before running; cap concurrent tasks
claude Sep 24, 2026
8bdc8f9
fix(hooks): matcher alternatives, tool families, case-insensitive
claude Sep 24, 2026
2b07fbd
fix(llm): stop rewriting "u003c"-style text in model output
claude Sep 24, 2026
a18db41
fix(redact): redact decoded values before encoding; catch env/URL/JSO…
claude Sep 24, 2026
2c5ec4a
fix(read): select the line window from the whole file; explicit markers
claude Sep 24, 2026
18b9ade
fix(edit): refuse NUL-containing and >1MB files instead of corrupting…
claude Sep 24, 2026
d81069b
fix(skills): reject path-traversal slugs in recall_skill / forget_skill
claude Sep 24, 2026
246501a
fix(session-search): reindex journals whose size/mtime changed since …
claude Sep 24, 2026
3ea9066
fix(run_tests): quote repo_path, managed timeout, keep output on failure
claude Sep 24, 2026
bd2805f
fix(profile): env beats a stale override; keep kwargs; /model keeps t…
claude Sep 24, 2026
b352665
test: remove the temp directories the skills / session-search tests c…
claude Sep 24, 2026
f067c80
fix(tools): ~ expansion, edit makes parents, bash failures count, sch…
claude Sep 24, 2026
64925db
docs: note CommandGuard in CONTRIBUTING and the gate's scope in SECURITY
claude Sep 24, 2026
f29bc6c
Merge wip/agents: MCP client/server, scheduler, subagents, flows, red…
claude Sep 24, 2026
f14a72f
fix(cli): -p takes the prompt after known flags and prompts starting …
claude Sep 24, 2026
e094e20
Merge wip/tools: bash wrapper, command classifier, grep/read/edit, ho…
claude Sep 24, 2026
9f56b83
test: close t_headless_default_allowed_still_run (lost in the wip/too…
claude Sep 24, 2026
8e2da5b
Merge wip/loop: cut-off tool calls, headless result, budget, compacti…
claude Sep 24, 2026
9ece596
fix: browser_screenshot goes through the write guard; /schedule shows…
claude Sep 24, 2026
a7a8933
feat: swarm-code trust / untrust / trust --list and /trust in the REPL
claude Sep 24, 2026
d7cccb0
fix(doctor): don't probe a remote endpoint without SWARM_CODE_ALLOW_R…
claude Sep 24, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 26 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,32 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Security

- **A cloned repo's `./.swarm-code.json` is untrusted.** It may set only
`model`, `max_tokens`, `vision`, `chat_template_kwargs`, `llm_timeout_ms`
and *tighten* `permissions`; hooks, `mcpServers`, `endpoint`, `api_key`,
`providers`, `profiles` and `fallback_profile` in it are ignored with a
one-line notice. Opt a repo in with `swarm-code trust` (or `/trust`), which
adds it to `"trusted_projects"` in `~/.swarm-code/settings.json`;
`swarm-code untrust` and `swarm-code trust --list` manage the list.
- **Network gate parses URLs properly and covers every LLM dial.** Userinfo
(`http://127.0.0.1@host`), uppercase schemes, name-prefix and numeric-IP
tricks no longer pass, and fallback / `providers` / `/profile` override
endpoints are checked at the point of dial. **Changed:** an API key no
longer bypasses the gate — remote endpoints need `SWARM_CODE_ALLOW_REMOTE=1`
as documented; scheme-less endpoints (`host:port`) are refused.
- **The model can't write swarm-code's control files** (`~/.swarm-code/`
hooks, `schedule.json`, `.profile_override`, sessions, `settings.json`,
`.swarm-code.json`); `memory/` and `skills/` stay writable. Sensitive-path
checks are case-insensitive (`~/.SSH`).
- **Headless no longer auto-approves an 'ask'.** A dangerous command, an
explicit `"ask"` permission or an MCP tool is denied in `-p` / cron /
`/flows` runs unless `SWARM_CODE_HEADLESS_APPROVE=1` is set.
- **The raw request body is no longer written to
`/tmp/swarm-code-last-body.json`** (or refreshed in `~/.swarm-code/` on every
call); `SWARM_CODE_DEBUG=1` writes `~/.swarm-code/last-body.json`, mode 600.

## [1.1.0] - 2026-07-15

The responsiveness release: the agent never blocks the terminal and the
Expand Down
1 change: 1 addition & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@
| `main.sw` | CLI flags, env config, headless mode, the main loop |
| `agent.sw` | Prompt assembly, tool-call parsing/serialising, turn loop |
| `ToolExecutor.sw` | Shared context, hook, guardrail, and permission boundary |
| `CommandGuard.sw` | Token-aware shell-command risk classifier (hardline / dangerous) used by the permission gate for every command-running tool |
| `tools.sw` | Raw tool handlers (`do_bash`, `do_read`, …) and `exec_raw()` dispatch |
| `ToolRegistry.sw` | Tool identity plus execution-context allow/block policy |
| `ToolSchemas.sw` | OpenAI-compatible function schemas for native tool calling |
Expand Down
25 changes: 21 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,8 +30,8 @@ search: docker compose (3 hits)
> recall_skill deploy-mally-otp # pull full playbook from ~/.swarm-code/skills/
SKILL.md loaded. Building burrito binary …

> /schedule add 1h "review open PRs" # heartbeat-driven cron
job 3 added (every 1h)
> /schedule "1h" "review open PRs" # heartbeat-driven cron (or "hourly", "daily 09:00")
✓ scheduled job 3 (1h): review open PRs
```

## Quickstart
Expand Down Expand Up @@ -78,7 +78,21 @@ Point it at any OpenAI-compatible endpoint via `~/.swarm-code/settings.json`. Pr
}
```

Remote endpoints are opt-in — set `SWARM_CODE_ALLOW_REMOTE=1` (local-network-only by default). Optional semantic memory recall uses `SWARM_CODE_EMBED_ENDPOINT`.
Remote endpoints are opt-in — set `SWARM_CODE_ALLOW_REMOTE=1` (local-network-only by default; an API key alone is not an opt-in). The check applies to every URL the LLM layer dials — primary, fallback, `providers`, and `/profile` switches. Optional semantic memory recall uses `SWARM_CODE_EMBED_ENDPOINT`.

A repo-local `./.swarm-code.json` is **untrusted** (it ships with whatever you cloned): it may set `model`, `max_tokens`, `vision`, `chat_template_kwargs` and `llm_timeout_ms`, and may only *tighten* `permissions`. Its hooks, MCP servers, endpoints, API keys, providers and profiles are ignored, with a one-line notice. To let a repo you trust apply its file in full, run `swarm-code trust` in it (or `/trust` in the REPL; `swarm-code untrust` reverses it, `swarm-code trust --list` shows the list). That adds the directory to `"trusted_projects"` in `~/.swarm-code/settings.json`, which you can also edit by hand.

### Hooks

`settings.json` can run shell commands around tool calls; a `PreToolUse` hook that exits non-zero (or times out after 60s, or cannot be started) blocks the call, and the hook's output is shown to the model:

```json
{ "hooks": {
"PreToolUse": [ { "matcher": "bash", "command": "./scripts/check-cmd.sh" } ],
"PostToolUse": [ { "matcher": "edit|write", "command": "make fmt" } ] } }
```

A matcher is `*` or `|`-separated, case-insensitive alternatives, each matching a tool whose name contains it — plus tool families: `bash` also fires for every other tool that runs a shell command (`background`, `bg_server`, `run_tests`, `file_watch`, `log_wait`), and `edit` also for `multi_edit`. The tool arguments arrive as JSON on the hook's **stdin** and in the private file `$SWARM_CODE_ARGS_FILE`; `$SWARM_CODE_ARGS` carries them inline only when under 100KB (else `$SWARM_CODE_ARGS_OMITTED=1`), so a hook that must see every call should read stdin. `$SWARM_CODE_EVENT` / `$SWARM_CODE_TOOL` name the event and tool. Executable scripts in `~/.swarm-code/hooks/` (`pre_tool.sh`, `post_tool.sh`, `pre_llm.sh`, `post_llm.sh`) get the same treatment via stdin / `$SWARM_HOOK_DATA_FILE` — see `src/Hooks.sw`.

## Features

Expand Down Expand Up @@ -119,8 +133,11 @@ Panel agents run under the fail-closed `council_panel` context: they may inspect
swarm-code runs shell commands, reads and writes files, and can reach the network — so it is built fail-closed:

- **Local-network-only by default**; remote endpoints require an explicit `SWARM_CODE_ALLOW_REMOTE=1`.
- A cloned repo's `./.swarm-code.json` cannot run hooks, start MCP servers, redirect the endpoint/key, or loosen permissions unless you trust the directory (`swarm-code trust`).
- The `write`/`edit` tools refuse swarm-code's own control files (`~/.swarm-code/` hooks, schedule, settings, sessions, profile override; `.swarm-code.json`) — only `~/.swarm-code/memory/` and `skills/` are writable — and credential dirs (`.ssh`, `.aws`, `.gnupg`, …) case-insensitively.
- Every tool runs through one **`ToolExecutor` policy boundary** — context allow-lists, argument-rewriting hooks, guardrails, and permissions — *before* any raw handler executes, and **fails closed** on a missing or unknown execution context.
- A **hardline command blocklist** (`rm -rf /`, `mkfs`, `dd`, fork bombs, …) cannot be bypassed by environment overrides.
- Headless runs (`-p`, cron jobs, `/flows` children) never auto-approve a call that needs permission — a dangerous command, an explicit `"ask"` setting, or an MCP tool — unless you set `SWARM_CODE_HEADLESS_APPROVE=1`.
- A **hardline command blocklist** (`rm -rf /`, `mkfs`, `dd` to a disk, halt/reboot, fork bombs, …) cannot be bypassed by environment overrides. It covers every tool that runs a shell command (`bash`, `background`, `bg_server`, `run_tests`), and commands are parsed like `sh` does — respellings such as `rm -fr /`, `dd of=/dev/sda if=…` or `sh -c '…'` are caught, while words inside quotes, `echo`/`grep` arguments or heredoc bodies are not flagged. Denials name the matched pattern.
- Subagents, MCP, and council contexts run under restricted (often read-only) policies.
- Secrets are redacted from session logs and trajectory exports.

Expand Down
37 changes: 34 additions & 3 deletions SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,22 +23,53 @@ backported.
## Security model

- **Network isolation by default.** Only local-network endpoints are allowed
unless you explicitly set `SWARM_CODE_ALLOW_REMOTE=1`.
unless you explicitly set `SWARM_CODE_ALLOW_REMOTE=1` (an API key is not an
opt-in). The check runs on every URL the LLM layer dials — primary,
fallback, `providers`, and the `/profile` override — and parses it the way
curl does: http(s) only, no `user@` credentials, IP literals only as strict
dotted quads or bracketed IPv6.
- **Untrusted project config.** A repository's `./.swarm-code.json` can only
set harmless keys (`model`, `max_tokens`, …) and *tighten* permissions. Its
hooks, MCP servers, endpoint / API key / providers / profiles, and any
permission loosening are ignored (with a notice) unless the directory is
listed under `trusted_projects` in `~/.swarm-code/settings.json` —
`swarm-code trust` (or `/trust`) adds the current directory for you.
Trust only repos you control: a trusted repo's hooks and MCP servers run
commands on your machine.
- **Single policy boundary.** Every tool call — from the main agent, subagents,
the council, and the MCP server — passes through `ToolExecutor`: context
allow-lists, argument-rewriting hooks, guardrails, and permissions are applied
*before* any raw handler runs. Execution **fails closed** when the execution
context is missing or unknown.
- **No unattended approvals.** In headless mode (`-p`, scheduled jobs, `/flows`
children) there is nobody to answer a permission prompt, so a call that
needs one — a dangerous command, a tool you set to `"ask"`, an MCP tool —
is denied. `SWARM_CODE_HEADLESS_APPROVE=1` opts back into auto-approval
(`/flows` children still hard-deny dangerous commands).
- **Hardline command blocklist.** Destructive commands (`rm -rf /`, `mkfs`, `dd`
to a device, fork bombs, and similar) are blocked and **cannot be bypassed by
environment overrides**.
environment overrides**. The check covers every tool that runs a shell
command (`bash`, `background`, `bg_server`, `run_tests`) and parses the
command like `sh` (quotes, separators, `$(…)`, `sh -c`, wrappers such as
`sudo`/`env`/`xargs`), so respellings are caught; it is a floor against
accidents, not a sandbox — a command assembled at runtime can't be judged.
- **Protected paths.** The `write` / `edit` / `multi_edit` tools refuse
credential locations (`.ssh`, `.gnupg`, `.aws`, `/etc`, …) and swarm-code's
own control files — everything under `~/.swarm-code/` except the `memory/`
and `skills/` data dirs (hooks there run on every tool call), plus
`.swarm-code.json`. Matching is case-insensitive (macOS filesystems are).
`SWARM_CODE_UNSAFE_WRITES=1` lifts these guards.
- **Restricted contexts.** Subagents, MCP-server, and council-panel contexts run
under narrowed, often read-only, tool policies.
- **Secret redaction.** Known secret patterns are redacted from session logs and
trajectory exports.
trajectory exports. The raw LLM request body is written to disk only with
`SWARM_CODE_DEBUG=1`, to `~/.swarm-code/last-body.json` with mode 600.

## Known limitations

- Protected paths are enforced by the file tools (`write`, `edit`,
`multi_edit`, `browser_screenshot`), not by the shell: a `bash` command can
still write anywhere your user can, so review the commands you approve.
- The council panel's read-only isolation is **tool-level**, not yet a filesystem
sandbox — panelists inherit the read tool's filesystem visibility.
- swarm-code runs the commands you (or a model you configured) direct it to. Run
Expand Down
Loading
Loading