Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 31 additions & 2 deletions DOCS.md
Original file line number Diff line number Diff line change
Expand Up @@ -113,6 +113,7 @@ For a quick overview, see the [README](./README.md). This document covers everyt
- [Thresholds](#thresholds)
- [Scanner Toggles](#scanner-toggles)
- [Symlinks](#symlinks)
- [Binary and non-UTF-8 files](#binary-and-non-utf-8-files)
- [Extended Scanners](#extended-scanners)
- [Platform Matrix](#platform-matrix)
- [Dependency Risk](#dependency-risk)
Expand Down Expand Up @@ -2136,7 +2137,7 @@ Switches (flags that take no value, such as `--vulns`, `--offline` or `--no-grap

By default, the scan writes `.vibgrate/scan_result.json`. Use `--no-local-artifacts` or `--max-privacy` to suppress local JSON artifact files.

`vg scan` does not follow symlinks while it indexes the tree. A skipped link is named once on stderr. See [Symlinks](#symlinks).
`vg scan` does not follow symlinks while it indexes the tree. A skipped link is named once on stderr. See [Symlinks](#symlinks). A binary or non-UTF-8 file that would otherwise be read as source is skipped. The scan prints one notice on stderr; when the scan also builds the code map, that build prints its own. See [Binary and non-UTF-8 files](#binary-and-non-utf-8-files).

For offline drift scoring, pass `--package-manifest <file>` with a downloaded manifest bundle such as `https://github.com/vibgrate/manifests/latest-packages.zip`. The manifest shape, the fail-closed errors, and what offline mode skips are in [Offline scan with a package-version manifest](#offline-scan-with-a-package-version-manifest).

Expand Down Expand Up @@ -2728,7 +2729,7 @@ Maps source code into a graph artifact, enabling all downstream queries (`vg sho
| `--attestation <file>` | `.vibgrate/attestation.intoto.jsonl` | Where `--attest` writes, and where `--verify` reads |
| `--pub <path>` | — | Public key PEM that pins the signer for `--verify` |

`vg build` does not follow symlinks while it discovers files. A skipped link is named once on stderr. See [Symlinks](#symlinks). `.gitignore` and `--exclude` still apply.
`vg build` does not follow symlinks while it discovers files. A skipped link is named once on stderr. See [Symlinks](#symlinks). `.gitignore` and `--exclude` still apply. A binary or non-UTF-8 file that would otherwise be read as source is skipped and named once on stderr. See [Binary and non-UTF-8 files](#binary-and-non-utf-8-files).

**Local by default — no git churn.** The first time vg writes into `.vibgrate/` it also creates `.vibgrate/.gitignore`, keeping the graph artifacts (`graph.json`, `graph.html`, `GRAPH_REPORT.md`, `facts.jsonl`, `mcp-navigation.json`) and the cache out of git — so builds, auto-refreshes, and MCP use never leave your branch dirty. Run `vg share` when you want the map committed for your team (it rewrites that ignore file). vg never touches an existing `.vibgrate/.gitignore`, so edit it (or leave it empty) to manage the ignores yourself.

Expand Down Expand Up @@ -5132,6 +5133,34 @@ text only; it does not hide this notice.
notice: skipped 3 symlinks (alias.ts, nested/cycle, via). vg does not follow symlinks. Point the root at the link target, or pass --exclude or a narrower root.
```

### Binary and non-UTF-8 files

`vg build` and `vg scan` read source as UTF-8 text. A file under the walk can still be a binary, a media blob, or another encoding, including one whose extension looks like source (`payload.js`, `README.md`). vg does not decode those bytes. It leaves the file out of the map and out of scan text, prints one notice for that walk on stderr, and continues. `vg scan` builds a code map after the scan unless you pass `--no-graph`, so the map build can print a second notice for the same paths. The command does not crash, and neither notice includes file bytes.

Known media extensions (images, fonts, audio, video, archives) are already left out of the scan walk. This notice is for a file vg would otherwise have opened as text. Output files do not change because of the notice itself, and the exit code stays the same when the rest of the command succeeds.

The notice is the count, then the first few root-relative paths in sorted order (at most five; further files are a `+N more` count). Paths are relative to the root you passed. Absolute paths are not printed. The line goes to stderr, so `--format json` and `vg build --json` keep a JSON document on stdout. `--quiet` hides promotional text only; it does not hide this notice.

```text
notice: skipped 2 files that are not UTF-8 text (README.md, payload.js). vg does not read binary or non-UTF-8 files as source. Ignore them with --exclude or a .gitignore rule. Example: vg build --exclude 'README.md'
```

To leave a file out without a notice, list it in `.gitignore` or pass `--exclude`. Both commands accept the flag. A config `exclude` list is merged in as well.

```bash
vg build --exclude 'payload.js'
vg build --exclude '*.bin' --exclude 'vendor-blobs/**'
vg scan --exclude 'legacy/blob.js'
```

```gitignore
# binary checked in under a source-like name
payload.js
*.bin
```

A source file that is valid UTF-8, including non-ASCII text, is still read. Save a file that uses another encoding as UTF-8 if it should be part of the map.

---

## Extended Scanners
Expand Down
2 changes: 2 additions & 0 deletions src/core-open/run-core-scan.ts
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,7 @@ import { loadConfig, appendExcludePatterns } from './config.js';
import { pathExists, readJsonFile, writeJsonFile, writeTextFile, ensureDir, FileCache, quickTreeCount } from './utils/fs.js';
import { portableValue } from './utils/portable-path.js';
import { assertSafeWalkRoot } from './utils/root-safety.js';
import { emitSkippedNonUtf8Notice } from './utils/source-text.js';
import { detectVcs } from './utils/vcs.js';
import { isCiEnvironment, hasVibgrateWorkflow } from './utils/ci-env.js';
import { resolveRepositoryName } from './utils/repository-name.js';
Expand Down Expand Up @@ -805,6 +806,7 @@ export async function runCoreScan(
}
}

emitSkippedNonUtf8Notice(fileCache.skippedNonUtf8);
fileCache.clear();

if (allProjects.length === 0) {
Expand Down
39 changes: 34 additions & 5 deletions src/core-open/utils/fs.ts
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,8 @@ import {
UnsafeRootError,
type WalkBudgetState,
} from './root-safety.js';
import { emitSkippedSymlinkNotice, rememberSkippedSymlink } from './skipped-symlinks.js';
import { emitSkippedSymlinkNotice, rememberSkippedSymlink, walkRelativePath } from './skipped-symlinks.js';
import { isUtf8SourceText } from './source-text.js';


const execFileAsync = promisify(execFile);
Expand Down Expand Up @@ -399,6 +400,13 @@ export class FileCache {
return this._skippedLargeFiles;
}

/** Binary or non-UTF-8 files a text read refused. Paths are root-relative. */
private _skippedNonUtf8: string[] = [];

get skippedNonUtf8(): readonly string[] {
return this._skippedNonUtf8;
}

// ── Directory walking ──

/**
Expand Down Expand Up @@ -770,7 +778,13 @@ export class FileCache {
}
}

const content = await fs.readFile(abs, 'utf8');
const buf = await fs.readFile(abs);
if (!isUtf8SourceText(buf)) {
this.noteNonUtf8(abs);
// Cache the refusal so later scanners do not read the blob again.
return '';
}
const content = buf.toString('utf8');
if (content.length > TEXT_CACHE_MAX_BYTES) {
// Too large for cache — evict so we don't hold it
this.textCache.delete(abs);
Expand Down Expand Up @@ -824,6 +838,15 @@ export class FileCache {
this.jsonCache.clear();
this.existsCache.clear();
this.sizeCache.clear();
this._skippedNonUtf8 = [];
}

/** Record a text read that refused binary or non-UTF-8 bytes. Path only. */
private noteNonUtf8(abs: string): void {
const rel = this._rootDir ? walkRelativePath(this._rootDir, abs) : '';
const label = rel || path.basename(abs);
if (!label || this._skippedNonUtf8.includes(label)) return;
this._skippedNonUtf8.push(label);
}

/** Number of file content entries currently held */
Expand Down Expand Up @@ -1135,12 +1158,18 @@ export function stripBom(text: string): string {
}

export async function readJsonFile<T>(filePath: string): Promise<T> {
const txt = await fs.readFile(filePath, 'utf8');
return JSON.parse(stripBom(txt)) as T;
const buf = await fs.readFile(filePath);
if (!isUtf8SourceText(buf)) {
const base = path.basename(filePath);
throw new SyntaxError(`${base || 'file'} is not UTF-8 text`);
}
return JSON.parse(stripBom(buf.toString('utf8'))) as T;
}

export async function readTextFile(filePath: string): Promise<string> {
return fs.readFile(filePath, 'utf8');
const buf = await fs.readFile(filePath);
if (!isUtf8SourceText(buf)) return '';
return buf.toString('utf8');
}

export async function pathExists(p: string): Promise<boolean> {
Expand Down
182 changes: 182 additions & 0 deletions src/core-open/utils/source-text.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,182 @@
import * as fs from 'node:fs';

/**
* Decide whether bytes are safe to open as source text, and format the one
* notice a walk prints when they are not.
*
* A NUL, a UTF-16 BOM, or any byte sequence that is not UTF-8 means the file
* is binary or another encoding. Callers skip it. The notice names paths
* only — never file bytes — and tells the operator how to ignore the file.
*/

/** How many leading bytes a large file is judged on before a full read. */
export const NON_UTF8_PREFIX_BYTES = 8192;

/**
* Marker on a parse row that was refused because the file is not UTF-8 text.
* The build drops the row. It is not a graph warning and not file contents.
*/
export const NON_UTF8_SKIP_MARK = 'vg-skip-non-utf8';

/** How many root-relative paths one notice lists. Further files are a count. */
export const SKIPPED_NON_UTF8_NOTICE_CAP = 5;

const HARD_FULL_READ_CAP = 32 * 1024 * 1024;

function comparePath(a: string, b: string): number {
return a < b ? -1 : a > b ? 1 : 0;
}

/**
* True when `bytes` must not be decoded as source text.
* `complete` means `bytes` is the whole file. A prefix may end on a cut
* multibyte character; that tail is not, by itself, a reason to skip.
*/
export function isNonUtf8Source(bytes: Uint8Array, complete: boolean): boolean {
if (bytes.length === 0) return false;
if (hasUtf16Bom(bytes)) return true;
for (let i = 0; i < bytes.length; i++) {
if (bytes[i] === 0) return true;
}
const sample = complete ? bytes : trimIncompleteUtf8Tail(bytes);
try {
new TextDecoder('utf-8', { fatal: true }).decode(sample);
return false;
} catch {
return true;
}
}

/** True when the whole buffer is UTF-8 text with no NUL and no UTF-16 BOM. */
export function isUtf8SourceText(bytes: Uint8Array): boolean {
return !isNonUtf8Source(bytes, true);
}

/**
* Read `abs` and return its text, or `null` when the bytes are not UTF-8
* source text. Throws when the file cannot be read.
*/
export function readUtf8SourceSync(abs: string): string | null {
const buf = fs.readFileSync(abs);
if (!isUtf8SourceText(buf)) return null;
return buf.toString('utf8');
}

/**
* True when a walked file should be left out of source-text handling.
* Files up to `fullReadCap` are checked in full. Larger files are judged
* on a prefix so a blob is never slurped just to be refused. `0` means
* prefix-only. Unreadable files return false; the caller already has a
* path for those.
*/
export function shouldSkipNonUtf8File(abs: string, fullReadCap: number): boolean {
let size: number;
try {
size = fs.statSync(abs).size;
} catch {
return false;
}
if (size === 0) return false;
const cap = fullReadCap > 0 ? Math.min(fullReadCap, HARD_FULL_READ_CAP) : 0;
try {
if (cap > 0 && size <= cap) {
return !isUtf8SourceText(fs.readFileSync(abs));
}
const n = Math.min(size, NON_UTF8_PREFIX_BYTES);
const buf = Buffer.alloc(n);
const fd = fs.openSync(abs, 'r');
try {
const got = fs.readSync(fd, buf, 0, n, 0);
return isNonUtf8Source(buf.subarray(0, got), got === size);
} finally {
fs.closeSync(fd);
}
} catch {
return false;
}
}

/**
* One stderr line when a walk skips binary or non-UTF-8 files. `null` when
* `relPaths` is empty. Paths are de-duplicated, sorted, and capped. Absolute
* paths and control characters are omitted. File bytes are never included.
*/
export function formatSkippedNonUtf8Notice(relPaths: readonly string[]): string | null {
const seen = new Set<string>();
const unique: string[] = [];
let hidden = 0;
for (const raw of relPaths) {
const key = raw.split('\\').join('/');
if (!key || key === '.' || seen.has(key)) continue;
seen.add(key);
const rel = displayRel(key);
if (!rel) hidden++;
else unique.push(rel);
}
unique.sort(comparePath);
const total = unique.length + hidden;
if (total === 0) return null;
const shown = unique.slice(0, SKIPPED_NON_UTF8_NOTICE_CAP);
const extra = unique.length - shown.length;
const list = extra > 0 ? `${shown.join(', ')}, +${extra} more` : shown.join(', ');
const noun = total === 1 ? 'file that is not UTF-8 text' : 'files that are not UTF-8 text';
const names = list.length > 0 ? ` (${list})` : '';
const example = shown[0] ? shellQuote(shown[0]) : "'*.bin'";
return (
`notice: skipped ${total} ${noun}${names}. ` +
'vg does not read binary or non-UTF-8 files as source. ' +
`Ignore them with --exclude or a .gitignore rule. Example: vg build --exclude ${example}`
);
}

/** Write {@link formatSkippedNonUtf8Notice} to stderr. No-op when there is nothing to say. */
export function emitSkippedNonUtf8Notice(relPaths: readonly string[]): void {
const line = formatSkippedNonUtf8Notice(relPaths);
if (!line) return;
process.stderr.write(`${line}\n`);
}

function hasUtf16Bom(bytes: Uint8Array): boolean {
if (bytes.length < 2) return false;
const b0 = bytes[0]!;
const b1 = bytes[1]!;
return (b0 === 0xff && b1 === 0xfe) || (b0 === 0xfe && b1 === 0xff);
}

/** Drop a trailing partial UTF-8 sequence so a prefix is not refused for being cut. */
function trimIncompleteUtf8Tail(bytes: Uint8Array): Uint8Array {
const n = bytes.length;
if (n === 0) return bytes;
let i = n - 1;
let cont = 0;
while (i >= 0 && cont < 3 && (bytes[i]! & 0xc0) === 0x80) {
cont++;
i--;
}
if (i < 0) return bytes;
const lead = bytes[i]!;
let need = 0;
if ((lead & 0x80) === 0) need = 1;
else if ((lead & 0xe0) === 0xc0) need = 2;
else if ((lead & 0xf0) === 0xe0) need = 3;
else if ((lead & 0xf8) === 0xf0) need = 4;
else return bytes;
const have = n - i;
if (have < need) return bytes.subarray(0, i);
return bytes;
}

/** A path safe to print. `null` when it must not appear in the notice. */
function displayRel(rel: string): string | null {
if (!rel || rel === '.' || rel.startsWith('../') || rel === '..' || rel.startsWith('/')) return null;
if (/^[A-Za-z]:/.test(rel)) return null;
for (let i = 0; i < rel.length; i++) {
const c = rel.charCodeAt(i);
if (c < 0x20 || c === 0x7f) return null;
}
return rel;
}

function shellQuote(value: string): string {
return `'${value.replace(/'/g, `'\\''`)}'`;
}
Loading
Loading