Context
While tracing a betweenness=0.047 bridge in a mid-sized TypeScript monorepo (17k+ nodes, 1,400 communities), the top God Node showed up as a cross-community hub between three communities — Controllers (Finance/Admin), Acquisition Survey Domain, and a doc-cluster.
Running graphify path/graphify explain revealed the bridge was fake: a function assertClubAccess() in permissions.ts had .getById() as a neighbor, but the .getById() node was labeled as living in "Acquisition Survey Domain" community — in reality the edge was to team.controller.ts:L25 .getById(), which was a same-name method-ID collision across controllers. The old (pre-#1504) node-ID scheme glued all .getById() into one node, and clustering + labeling fell apart.
After a fresh rebuild (delete manifest.json → full re-extract) with path-qualified IDs:
- Node count: 17,580 → 14,011
- Communities: 1,399 → 1,087 (tighter clusters)
assertClubAccess() degree: 5 → 4, now single community (Controllers Core)
The bug + collision warning fired correctly:
[graphify] note: this graph uses the pre-#1504 node-ID scheme;
rebuild with `graphify extract --force` to get path-qualified IDs
(fixes same-name-file collisions).
Problem: this notice only appears on graphify query/path/explain output, not on GRAPH_REPORT.md, and graphify extract --force isn't actually a CLI command (graphify update --force is the real path, and even that doesn't trigger a re-extract — just overrides shrink-guard). The user has to delete manifest.json manually, which is an unobvious operation.
Proposals
1. Expose a dedicated rebuild subcommand
Something like graphify rebuild [--path .] that:
- Backs up
manifest.json to .manifest-backup-<ts>.json
- Deletes
manifest.json + .graphify_old.json
- Triggers the full detect → extract → merge pipeline with
root= relative paths (same as current --update)
- Prints "Semantic cache preserved: X/Y files" at the end
Rationale: semantic re-extraction stayed 100% cache-hit in my case (288/288 docs), so the rebuild cost was only AST (seconds, no tokens). The command looks expensive but almost never is.
2. Surface the pre-#1504 banner inside GRAPH_REPORT.md
Right now, users who read the report but never run query miss the warning. Add a top-of-report block:
> ⚠️ **Graph uses pre-#1504 node-ID scheme.** Same-name files/methods may collide, inflating betweenness
> and distorting community membership. Run `graphify rebuild` to regenerate with path-qualified IDs.
> (Likely cost: seconds + 0 tokens if your semantic cache is intact.)
Check for the collision by looking for duplicate label with different source_file in graph.json nodes — if the duplicate rate > 0.5%, print the banner.
3. (Optional) Suggested Questions / Surprising Connections should flag collision-driven bridges
High-betweenness nodes whose degree is dominated by same-label-different-source neighbors are statistical artifacts. If (unique source_files among neighbors) / degree < 0.5, demote the question to "likely a node-ID collision — run graphify rebuild."
Why this matters
Users new to graphify see the bridge, trust the clustering, and run queries against noise. The remediation path exists but is buried in a single-line note; the warning + fix flow should be loud enough that users don't trace a false bridge for 20 minutes before finding the escape hatch.
Happy to send a PR for #1 (the subcommand) if the direction looks right.
Context
While tracing a
betweenness=0.047bridge in a mid-sized TypeScript monorepo (17k+ nodes, 1,400 communities), the top God Node showed up as a cross-community hub between three communities — Controllers (Finance/Admin), Acquisition Survey Domain, and a doc-cluster.Running
graphify path/graphify explainrevealed the bridge was fake: a functionassertClubAccess()inpermissions.tshad.getById()as a neighbor, but the.getById()node was labeled as living in "Acquisition Survey Domain" community — in reality the edge was toteam.controller.ts:L25.getById(), which was a same-name method-ID collision across controllers. The old (pre-#1504) node-ID scheme glued all.getById()into one node, and clustering + labeling fell apart.After a fresh rebuild (delete
manifest.json→ full re-extract) with path-qualified IDs:assertClubAccess()degree: 5 → 4, now single community (Controllers Core)The bug + collision warning fired correctly:
Problem: this notice only appears on
graphify query/path/explainoutput, not onGRAPH_REPORT.md, andgraphify extract --forceisn't actually a CLI command (graphify update --forceis the real path, and even that doesn't trigger a re-extract — just overrides shrink-guard). The user has to deletemanifest.jsonmanually, which is an unobvious operation.Proposals
1. Expose a dedicated rebuild subcommand
Something like
graphify rebuild [--path .]that:manifest.jsonto.manifest-backup-<ts>.jsonmanifest.json+.graphify_old.jsonroot=relative paths (same as current--update)Rationale: semantic re-extraction stayed 100% cache-hit in my case (288/288 docs), so the rebuild cost was only AST (seconds, no tokens). The command looks expensive but almost never is.
2. Surface the pre-#1504 banner inside
GRAPH_REPORT.mdRight now, users who read the report but never run
querymiss the warning. Add a top-of-report block:Check for the collision by looking for duplicate
labelwith differentsource_fileingraph.jsonnodes — if the duplicate rate > 0.5%, print the banner.3. (Optional) Suggested Questions / Surprising Connections should flag collision-driven bridges
High-betweenness nodes whose degree is dominated by same-label-different-source neighbors are statistical artifacts. If
(unique source_files among neighbors) / degree < 0.5, demote the question to "likely a node-ID collision — rungraphify rebuild."Why this matters
Users new to graphify see the bridge, trust the clustering, and run queries against noise. The remediation path exists but is buried in a single-line note; the warning + fix flow should be loud enough that users don't trace a false bridge for 20 minutes before finding the escape hatch.
Happy to send a PR for #1 (the subcommand) if the direction looks right.