Repository navigation
fix(core): index files in a new directory the watcher cannot see into - #1664
Open
sammywachtel wants to merge 1 commit into
Open
sammywachtel wants to merge 1 commit into
sammywachtel wants to merge 1 commit into
Conversation
On Linux, files copied into a running server in bulk could stay on disk and out of the index, whole folders at a time, with nothing logged. watchfiles (over notify's inotify backend) learns of a new directory from its parent and then adds a watch of its own. If the directory cannot be read at that instant the add fails and notify discards the error, so no file written into it ever produces an event. `install -d -o <user>` run as root creates directories this way, and so does any copy or restore made as root and chowned afterwards. The directory's own creation event still arrives, but the watcher dropped directory events. The periodic restart re-added the watch without looking for files that had landed meanwhile, and it discarded changes collected but not yet handed over. - A reported new directory is walked and its files indexed in that batch. A directory renamed inside the project moves its rows instead of doubling them. - A maintenance tick restarts the watcher only when the project set changed or a new directory appeared, and each restart is followed by a reconcile. - A periodic reconcile indexes settled files on disk that the index does not know, and logs a warning naming them. Its walk runs in a worker thread. Signed-off-by: sammywachtel <subp@wachtel.us>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
On Linux, files copied into a running server in bulk can stay on disk and out of the index, whole folders at a time. They are not in search, the graph or any tool. Nothing is logged, and the watcher looks idle, as if it had finished. A restart fixes it, because the startup scan indexes whatever is on disk.
The cause is how a new directory gets watched. watchfiles (built on notify's inotify backend) learns about a new directory from the directory above it, and then adds a watch on the new directory itself. If the new directory cannot be read at that instant, adding the watch fails, and notify discards the error (
self.add_watch(path, true, false).ok()in notify 8'sinotify.rs). From then on, no file written into that directory produces an event.install -d -o <user> -g <group>run as root creates directories this way: the directory exists for a moment, owned by root and closed, before it is handed over. A restore or an SSH copy made as root and chowned afterwards has the same window. The watcher runs as the unprivileged user, so it loses that race.Three things in the watch service then make the loss permanent:
local_storage_event_input_from_watchfiles_changedrops every directory event (if metadata is None or metadata.is_dir: return None).watch_project_reload_interval) builds a new watcher, which does watch the directory. So later edits are seen. But nothing scans for the files that landed before. A restart also throws away the changes the watcher has collected but not yet handed over: the Rustwatch()clears them when it sees the stop event. That loss is deterministic whenever a tick lands during a slow batch.Evidence
install -d -o 1000 -g 1000 -m 0755followed byinstall -o 1000 -g 1000per file. Three runs reported 7, 7 and 11 of the 49 files. CPU load made no difference. Controls: the same copy without-o/-g(directories root-owned but readable) reported 49 of 49. The same copy with the watcher running as root reported 49 of 49.--cpus 1), with the same copy into an empty project directory while it ran: 12 to 36 of 49 files indexed, then idle for up to 15 minutes. Whole folders were missing, and which folders varied from run to run. This happened with the default 300 s reload interval and with it raised to an hour. In one run, a second, independent watchfiles process started in the same container reported only 9 of 49, so the events never reached the server at all.What changed
index/watch_service.py, newindex/watch_reconcile.py:addedchange for every eligible file inside it, using the same hidden, ignore and symlink rules as the startup scan. The walk runs in a worker thread.deletedchanges, so the existing move processor pairs them by checksum. Without this, the walk above would index a renamed directory a second time under its new path.file_path(one query). Every settled file the index does not know is indexed through the normal watcher path, with a warning that names the files:Index reconcile: N file(s) in project X were on disk but not in the index, so the watcher missed them; indexing now. "Settled" means unchanged for 60 s, judged on the later of mtime and ctime, so a file still in the watcher's debounce window is left alone. A copy that preserves mtime still has a fresh ctime. The reconcile skips a tick while a batch is in flight, just after one, or while the startup scan is still running. A file it already tried that still will not index is not retried until it changes.services/initialization.py: tells the watch service whether the startup scan is still running.config_models.py: thewatch_project_reload_intervaldescription says what the tick now does.Testing
New
tests/index/test_watch_bulk_copy.py:test_a_copy_into_a_directory_the_watcher_cannot_watch_is_indexed(Linux, non-root only): runs the real watch loop with real indexing.people/is created with mode 000, opened to 755, and then 8 notes are written into it. All 8 must be indexed within a 20 s ceiling.test_negative_control_without_the_fix_the_copy_stays_unindexed: the same scenario with the directory walk and the reconcile patched out. None of the 8 notes is indexed. This shows the scenario reproduces the defect.test_bulk_copy_into_a_busy_watcher_is_fully_delivered: a tick with nothing to do no longer restarts the watcher, so a copy that lands during a slow batch is delivered (ceiling 20 s). Onmainit fails withnever delivered: [...].test_negative_control_restart_on_every_tick_loses_the_copy: the same scenario with the old restart-on-every-tick behavior, which loses the copy.test_a_reported_new_directory_indexes_the_files_inside_itandtest_a_directory_moved_inside_the_project_moves_its_rows(entity ids are kept).Runs:
tests/index,tests/services,tests/test_config.pyon macOS, SQLite: 968 passed, 5 skipped. The two Linux-only tests skip on macOS, where FSEvents watches the whole tree and has no per-directory watch.tests/indexon Linux (Debian, Python 3.12, non-root user): 197 passed, including both end-to-end tests.ruff check,ruff format,ty check src tests test-int: clean.Risks / follow-ups
file_pathquery per project per tick, 300 s by default. It reads no file contents unless a file is missing from the index.add_watch. Reporting that error, or rescanning a directory whose watch failed, belongs in notify or watchfiles, and I have not filed it there. This change makes the watch service correct whatever those libraries do.