suggest_questions() and the GRAPH_REPORT.md "Knowledge Gaps" section both treat degree(n) <= 1 as "possible documentation gap," filtered only by _is_file_node, _is_concept_node, and file_type != "rationale" (analyze.py:549-563, report.py:327, shared via _real_node).
_is_file_node (analyze.py:65-90) only recognizes three shapes: a label matching its own filename, a .method() stub, or a bare function_name() stub. _is_concept_node (analyze.py:183-193) only excludes nodes with an empty or extension-less source_file.
Neither catches a plain AST declaration whose label has no trailing () — e.g. a TypeScript type/interface alias, an enum member, a local const, or a JSON config key — even when it has a real source_file and _origin == "ast", and even when its only edge is the contains edge from its own file.
Concrete repro: in a Next.js/Medusa monorepo, three file-local row-shape types declared back-to-back:
type TabRow = { id: string; ... };
type ContributionRow = { id: string; ... };
type TabRecognition = { vendor_id: string; ... };
(used only as manager.execute<ContributionRow[]>(...) generic parameters, never exported) — each got exactly one contains edge and were then surfaced by suggest_questions() as "What connects ContributionRow, TabRecognition, TabRow to the rest of the system? — 1,556 weakly-connected nodes found — possible documentation gaps." They are not a gap; they're working as designed.
This isn't a one-off: in a ~5,000-node graph of that repo, 1,718 of 2,033 weakly-connected nodes have _origin == "ast", vs. 260 _origin == "semantic" (docs/PRD-derived) and 55 untagged. Manually sampling 26 of the semantic ones found zero confirmed real gaps too (mostly risk-register table rows and section headings) — but the AST-origin majority is the larger, more mechanically-fixable false-positive source.
Suggested fix: have _is_file_node (or a sibling check) also exclude any node with _origin == "ast" whose only edge is a single contains edge from its declaring file — regardless of whether the label ends in (). That one change would remove the great majority of the 1,718 false positives in our graph without touching the semantic/doc side of gap detection.
(graphify version: graphifyy 0.9.79, installed via uv tool install)
suggest_questions()and theGRAPH_REPORT.md"Knowledge Gaps" section both treatdegree(n) <= 1as "possible documentation gap," filtered only by_is_file_node,_is_concept_node, andfile_type != "rationale"(analyze.py:549-563,report.py:327, shared via_real_node)._is_file_node(analyze.py:65-90) only recognizes three shapes: a label matching its own filename, a.method()stub, or a barefunction_name()stub._is_concept_node(analyze.py:183-193) only excludes nodes with an empty or extension-lesssource_file.Neither catches a plain AST declaration whose label has no trailing
()— e.g. a TypeScripttype/interfacealias, an enum member, a localconst, or a JSON config key — even when it has a realsource_fileand_origin == "ast", and even when its only edge is thecontainsedge from its own file.Concrete repro: in a Next.js/Medusa monorepo, three file-local row-shape types declared back-to-back:
(used only as
manager.execute<ContributionRow[]>(...)generic parameters, never exported) — each got exactly onecontainsedge and were then surfaced bysuggest_questions()as "What connectsContributionRow,TabRecognition,TabRowto the rest of the system? — 1,556 weakly-connected nodes found — possible documentation gaps." They are not a gap; they're working as designed.This isn't a one-off: in a ~5,000-node graph of that repo, 1,718 of 2,033 weakly-connected nodes have
_origin == "ast", vs. 260_origin == "semantic"(docs/PRD-derived) and 55 untagged. Manually sampling 26 of the semantic ones found zero confirmed real gaps too (mostly risk-register table rows and section headings) — but the AST-origin majority is the larger, more mechanically-fixable false-positive source.Suggested fix: have
_is_file_node(or a sibling check) also exclude any node with_origin == "ast"whose only edge is a singlecontainsedge from its declaring file — regardless of whether the label ends in(). That one change would remove the great majority of the 1,718 false positives in our graph without touching the semantic/doc side of gap detection.(graphify version: graphifyy 0.9.79, installed via
uv tool install)