Skip to content

fix: skip binary and non-UTF-8 files opened as source text - #429

Open
vibgrate-team wants to merge 1 commit into
mainfrom
cursor/binary-non-utf8-source-cb2f
Open

vibgrate-team wants to merge 1 commit into
mainfrom
cursor/binary-non-utf8-source-cb2f

Conversation

@vibgrate-team

Copy link
Copy Markdown
Contributor

Summary

vg build and vg scan no longer treat a binary, media, or non-UTF-8 file as source text just because it sits under the walk or has a source-like name (payload.js, README.md).

Those bytes are not decoded. The file is left out of the code map and out of scan text. One stderr notice names the count and the first few root-relative paths, and shows how to ignore the file. The notice never includes file contents. A file already covered by .gitignore or --exclude stays quiet. Valid UTF-8 source, including non-ASCII text, is still read. This is separate from the filesystem-root guard and from symlink handling.

notice: skipped 2 files that are not UTF-8 text (README.md, payload.js). vg does not read binary or non-UTF-8 files as source. Ignore them with --exclude or a .gitignore rule. Example: vg build --exclude 'README.md'
vg build --exclude 'payload.js'
vg build --exclude '*.bin' --exclude 'vendor-blobs/**'
vg scan --exclude 'legacy/blob.js'

When vg scan also builds the code map, that build can print its own notice for the same paths. Each line is still one summary.

How to verify

pnpm test test/non-utf8-source.test.ts
pnpm test
pnpm lint
pnpm typecheck

The fixture writes a small binary blob (PNG-like header, a NUL, invalid UTF-8, and an ASCII marker) to payload.js and README.md, plus a real src/keep.ts and a logo.png. It asserts:

  • discovery and buildGraph do not throw
  • the map stays byte-stable across two builds and contains keep, not the blob
  • stderr is the single sorted notice and does not contain the marker or an absolute path
  • .gitignore and --exclude 'payload.js' omit the file from the notice
  • runCoreScan finishes, and the artifact does not contain the marker

Related issues

Closes #294

Checklist

  • pnpm test passes
  • pnpm lint is clean
  • pnpm typecheck is clean
  • Docs updated (README / DOCS / ARCHITECTURE) where behavior changed
  • Determinism preserved — identical input still produces identical graph.json / report output (content-hashed IDs, stable sorts; no time, randomness, or filesystem-order dependence)
  • No proprietary or internal references — public, Apache-2.0 content only
  • Commits use Conventional Commits and are signed off (git commit -s, DCO)

Notes for reviewers

Known media extensions were already dropped from the scan walk. This change is the case where the file would have been opened as text: a misleading extension, or a context file such as README.md. A NUL, a UTF-16 BOM, or any non-UTF-8 sequence is enough to skip. Files above the per-file cap are judged on an 8 KiB prefix so a large blob is not loaded just to be refused.

Open in Web Open in Cursor 

Walks for vg build and vg scan no longer decode a blob just because
its name looks like source. Those files are skipped, one stderr notice
names them and shows a vg --exclude example, and the bytes stay out of
the map and the scan artifact.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: vibgrate-team <vibgrate-team@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bug: actionable error when binary or non-UTF8 files are treated as source text

2 participants