Skip to content

fix(core): index files in a new directory the watcher cannot see into - #1664

Open
sammywachtel wants to merge 1 commit into
basicmachines-co:mainfrom
sammywachtel:fix/watch-new-directory-contents
Open

sammywachtel wants to merge 1 commit into
basicmachines-co:mainfrom
sammywachtel:fix/watch-new-directory-contents

Conversation

@sammywachtel

@sammywachtel sammywachtel commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Why

On Linux, files copied into a running server in bulk can stay on disk and out of the index, whole folders at a time. They are not in search, the graph or any tool. Nothing is logged, and the watcher looks idle, as if it had finished. A restart fixes it, because the startup scan indexes whatever is on disk.

The cause is how a new directory gets watched. watchfiles (built on notify's inotify backend) learns about a new directory from the directory above it, and then adds a watch on the new directory itself. If the new directory cannot be read at that instant, adding the watch fails, and notify discards the error (self.add_watch(path, true, false).ok() in notify 8's inotify.rs). From then on, no file written into that directory produces an event.

install -d -o <user> -g <group> run as root creates directories this way: the directory exists for a moment, owned by root and closed, before it is handed over. A restore or an SSH copy made as root and chowned afterwards has the same window. The watcher runs as the unprivileged user, so it loses that race.

Three things in the watch service then make the loss permanent:

  1. The directory's own creation event does arrive, because the parent's watch is fine. But local_storage_event_input_from_watchfiles_change drops every directory event (if metadata is None or metadata.is_dir: return None).
  2. The periodic restart (watch_project_reload_interval) builds a new watcher, which does watch the directory. So later edits are seen. But nothing scans for the files that landed before. A restart also throws away the changes the watcher has collected but not yet handed over: the Rust watch() clears them when it sees the stop event. That loss is deterministic whenever a tick lands during a slow batch.
  3. Nothing compares the disk with the index while the server runs.

Evidence

  • watchfiles alone, no server: a watcher running as uid 1000 in a Linux container, with 49 Markdown files in 8 folders copied in as root with install -d -o 1000 -g 1000 -m 0755 followed by install -o 1000 -g 1000 per file. Three runs reported 7, 7 and 11 of the 49 files. CPU load made no difference. Controls: the same copy without -o/-g (directories root-owned but readable) reported 49 of 49. The same copy with the watcher running as root reported 49 of 49.
  • The server (Postgres, --cpus 1), with the same copy into an empty project directory while it ran: 12 to 36 of 49 files indexed, then idle for up to 15 minutes. Whole folders were missing, and which folders varied from run to run. This happened with the default 300 s reload interval and with it raised to an hour. In one run, a second, independent watchfiles process started in the same container reported only 9 of 49, so the events never reached the server at all.
  • With this change, the same server scenario indexed 49 of 49 (in 195 s with the container otherwise idle, and in 185 s with two busy loops competing for its CPU).

What changed

  • index/watch_service.py, new index/watch_reconcile.py:
    • New directories are walked. A batch that reports a new directory gets an added change for every eligible file inside it, using the same hidden, ignore and symlink rules as the startup scan. The walk runs in a worker thread.
    • A renamed directory moves its rows. When a batch reports a new directory and a gone one together, the indexed files under the gone one get deleted changes, so the existing move processor pairs them by checksum. Without this, the walk above would index a renamed directory a second time under its new path.
    • Restarts need a reason. A maintenance tick restarts the watcher only when the project set changed, or a new directory appeared (so that the new directory gets watched). Each restart is followed, once indexing is quiet, by a reconcile.
    • A periodic reconcile. On every tick that does not restart, each project is walked in a worker thread and compared with its entities' file_path (one query). Every settled file the index does not know is indexed through the normal watcher path, with a warning that names the files: Index reconcile: N file(s) in project X were on disk but not in the index, so the watcher missed them; indexing now. "Settled" means unchanged for 60 s, judged on the later of mtime and ctime, so a file still in the watcher's debounce window is left alone. A copy that preserves mtime still has a fresh ctime. The reconcile skips a tick while a batch is in flight, just after one, or while the startup scan is still running. A file it already tried that still will not index is not retried until it changes.
  • services/initialization.py: tells the watch service whether the startup scan is still running.
  • config_models.py: the watch_project_reload_interval description says what the tick now does.

Testing

New tests/index/test_watch_bulk_copy.py:

  • test_a_copy_into_a_directory_the_watcher_cannot_watch_is_indexed (Linux, non-root only): runs the real watch loop with real indexing. people/ is created with mode 000, opened to 755, and then 8 notes are written into it. All 8 must be indexed within a 20 s ceiling.
  • test_negative_control_without_the_fix_the_copy_stays_unindexed: the same scenario with the directory walk and the reconcile patched out. None of the 8 notes is indexed. This shows the scenario reproduces the defect.
  • test_bulk_copy_into_a_busy_watcher_is_fully_delivered: a tick with nothing to do no longer restarts the watcher, so a copy that lands during a slow batch is delivered (ceiling 20 s). On main it fails with never delivered: [...].
  • test_negative_control_restart_on_every_tick_loses_the_copy: the same scenario with the old restart-on-every-tick behavior, which loses the copy.
  • test_a_reported_new_directory_indexes_the_files_inside_it and test_a_directory_moved_inside_the_project_moves_its_rows (entity ids are kept).
  • Reconcile tests:
    • it indexes a file nothing reported and logs that it had to;
    • a second pass is silent;
    • a file still settling is left alone;
    • a file that will not index is tried once, and again after it changes;
    • it waits while a batch is in flight or the startup scan runs.

Runs:

  • tests/index, tests/services, tests/test_config.py on macOS, SQLite: 968 passed, 5 skipped. The two Linux-only tests skip on macOS, where FSEvents watches the whole tree and has no per-directory watch.
  • tests/index on Linux (Debian, Python 3.12, non-root user): 197 passed, including both end-to-end tests.
  • ruff check, ruff format, ty check src tests test-int: clean.
  • Not run: the Postgres suites.

Risks / follow-ups

  • A file edited inside an unwatched directory before the watcher restarts keeps its old indexed content until it is edited again or the server restarts. The reconcile looks for missing files, not changed ones. The window is at most one reload interval.
  • A directory moved out of the project still leaves its entities behind, as it does today. Only a move within one batch (gone plus new) is expanded.
  • The reconcile costs one walk and one file_path query per project per tick, 300 s by default. It reads no file contents unless a file is missing from the index.
  • The root cause is notify discarding the failed add_watch. Reporting that error, or rescanning a directory whose watch failed, belongs in notify or watchfiles, and I have not filed it there. This change makes the watch service correct whatever those libraries do.

On Linux, files copied into a running server in bulk could stay on disk and
out of the index, whole folders at a time, with nothing logged.

watchfiles (over notify's inotify backend) learns of a new directory from its
parent and then adds a watch of its own. If the directory cannot be read at
that instant the add fails and notify discards the error, so no file written
into it ever produces an event. `install -d -o <user>` run as root creates
directories this way, and so does any copy or restore made as root and
chowned afterwards. The directory's own creation event still arrives, but the
watcher dropped directory events. The periodic restart re-added the watch
without looking for files that had landed meanwhile, and it discarded changes
collected but not yet handed over.

- A reported new directory is walked and its files indexed in that batch. A
  directory renamed inside the project moves its rows instead of doubling them.
- A maintenance tick restarts the watcher only when the project set changed or
  a new directory appeared, and each restart is followed by a reconcile.
- A periodic reconcile indexes settled files on disk that the index does not
  know, and logs a warning naming them. Its walk runs in a worker thread.

Signed-off-by: sammywachtel <subp@wachtel.us>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant