Context
Working through a trace of a cross-community bridge surfaced by GRAPH_REPORT.md → Suggested Questions, I hit a case where graphify correctly tagged an edge as AMBIGUOUS (confidence 0.35, conceptually_related_to, source_location: null) but the downstream pipeline still treated it as a signal:
- Suggested Questions asked the user to explain the AMBIGUOUS relationship as if it were real.
- Louvain clustering merged the two end-doc clusters into one community because the AMBIGUOUS edge contributed to modularity.
End result: a confirmed false positive consumed exploration time, and a 27-node community was shaped around noise.
Example (anonymized to the pattern, not the specific repo)
Two unrelated planning docs — one about a BE PATCH endpoint for a video-evaluation feature, one about a department-hierarchy schema refactor — have zero cross-references in prose but share surface structure (Goal / Architecture / File Map / Phase templates, same tech stack, same superpowers:... sub-skill hint). The LLM semantic extractor inferred conceptually_related_to at confidence 0.35 / source_location: null.
The AMBIGUOUS flag worked perfectly — the problem is only that consumers of the graph (suggest_questions, cluster) weight it the same as EXTRACTED/INFERRED edges.
Proposals
1. suggest_questions should skip or demote AMBIGUOUS edges
Currently (graphify.analyze.suggest_questions) a bridge with confidence == AMBIGUOUS is phrased identically to a high-confidence one:
"What is the exact relationship between A and B? Edge tagged AMBIGUOUS (relation: conceptually_related_to) - confidence is low."
The disclaimer is good, but the question still spends user attention. Two options:
- (a) Skip AMBIGUOUS-only bridges unless the user passes
--include-ambiguous.
- (b) Still emit the question but suppress it from the top-N and append it in a separate "Low-confidence hints" section at the bottom of the report, so the user sees it after the real signal.
My preference: (b), keeps the audit trail without burying real bridges.
2. Clustering should weight AMBIGUOUS edges below EXTRACTED/INFERRED
Louvain modularity in graphify.cluster currently reads weight from the raw confidence_score (0.35 for AMBIGUOUS in my case). Two cheap wins:
- Multiply AMBIGUOUS weight by a configurable factor (default 0.5) before passing to
nx.community.louvain_communities(..., weight='weight').
- Expose it as
cluster(G, ambiguous_scale=0.5) so power users can bump it to 0.0 (drop entirely) or 1.0 (current behavior).
This keeps the community graph more faithful to EXTRACTED structure without changing the edge data itself. Users exploring via /graphify path still see the AMBIGUOUS edge; only the auto-cluster stops being warped by it.
Backward compatibility
Both changes are additive (new flags, defaults preserve existing behavior if you choose ambiguous_scale=1.0). Nothing in the schema or output file format changes.
Related
- Honesty Rules in
SKILL.md already say "Never invent an edge. If unsure, use AMBIGUOUS." This proposal just completes the loop by making downstream consumers honor the confidence tier.
confidence_score is already written to graph.json — this is a pure consumer-side change.
Happy to open a PR for either or both if the maintainers agree with the direction.
Context
Working through a trace of a cross-community bridge surfaced by
GRAPH_REPORT.md→ Suggested Questions, I hit a case wheregraphifycorrectly tagged an edge asAMBIGUOUS(confidence 0.35,conceptually_related_to,source_location: null) but the downstream pipeline still treated it as a signal:End result: a confirmed false positive consumed exploration time, and a 27-node community was shaped around noise.
Example (anonymized to the pattern, not the specific repo)
Two unrelated planning docs — one about a BE PATCH endpoint for a video-evaluation feature, one about a department-hierarchy schema refactor — have zero cross-references in prose but share surface structure (Goal / Architecture / File Map / Phase templates, same tech stack, same
superpowers:...sub-skill hint). The LLM semantic extractor inferredconceptually_related_toat confidence 0.35 /source_location: null.The AMBIGUOUS flag worked perfectly — the problem is only that consumers of the graph (
suggest_questions,cluster) weight it the same as EXTRACTED/INFERRED edges.Proposals
1.
suggest_questionsshould skip or demote AMBIGUOUS edgesCurrently (
graphify.analyze.suggest_questions) a bridge withconfidence == AMBIGUOUSis phrased identically to a high-confidence one:The disclaimer is good, but the question still spends user attention. Two options:
--include-ambiguous.My preference: (b), keeps the audit trail without burying real bridges.
2. Clustering should weight AMBIGUOUS edges below EXTRACTED/INFERRED
Louvain modularity in
graphify.clustercurrently readsweightfrom the rawconfidence_score(0.35 for AMBIGUOUS in my case). Two cheap wins:nx.community.louvain_communities(..., weight='weight').cluster(G, ambiguous_scale=0.5)so power users can bump it to 0.0 (drop entirely) or 1.0 (current behavior).This keeps the community graph more faithful to EXTRACTED structure without changing the edge data itself. Users exploring via
/graphify pathstill see the AMBIGUOUS edge; only the auto-cluster stops being warped by it.Backward compatibility
Both changes are additive (new flags, defaults preserve existing behavior if you choose
ambiguous_scale=1.0). Nothing in the schema or output file format changes.Related
SKILL.mdalready say "Never invent an edge. If unsure, use AMBIGUOUS." This proposal just completes the loop by making downstream consumers honor the confidence tier.confidence_scoreis already written tograph.json— this is a pure consumer-side change.Happy to open a PR for either or both if the maintainers agree with the direction.