Skip to content

reshape: parser/ (subtree) + graph/ + bin/axiom-graph. One repository for the whole pipeline - #469

Merged
swapnilpaliwal-sd merged 328 commits into
mainfrom
reshape
Sep 14, 2026
Merged

swapnilpaliwal-sd merged 328 commits into
mainfrom
reshape

Conversation

@swapnilpaliwal-sd

@swapnilpaliwal-sd swapnilpaliwal-sd commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #468. Also fixes #575 (the parser now lives here). Do not merge until every acceptance item below is checked.

Layout

bin/axiomcode        ONE command for any repository: `bin/axiomcode <src> <out> [--library <src-or-ir>,…]` — parses every language present, solves each, leaves <out>/<lang>/graph.sqlite. Libraries may be source trees (parsed for you) or IR roots. `parser` / `engine` / `test` are the pieces.
parser/              AxiomCodeAI/parser merged in with its full history (312 commits, rewritten under the parser/ prefix); npm workspace; its internals untouched
graph/               the engine (was src/): <lang>/, pipeline/, bundle/, and test/ — the engine's suites, defaulting to parser/dist in this repository
binaries/ .github/   arrive with #455, rebased onto this

Two package.jsons remain: the parser keeps its own; the root declares workspaces: ["parser"], so npm install && npm run build builds both.

What changed, mechanically

  • git mv src graph; every repository-path reference rewritten (tsconfig, package.json, constants/paths.ts, the executor, suites, tools). Fixture directories named src/ inside test cases and the JDK-layout fixture in jdk-ir-stamp-test.sh are untouched (one was wrongly rewritten at first; the suite caught it).
  • The --debug CSV output directory is renamed graph/ → csv/: .gitignore ignored graph/ as an output name, which would have hidden the new source folder.
  • Every suite's parser default is parser/dist/index.js here instead of a sibling ../Parser checkout; AXIOM_PARSER still overrides.
  • The engine's suites move to graph/test/ (each package owns its tests; the parser's stay where its repository keeps them). test/java/README.md removed (its run instructions are now the root README's Tests section); graph/typescript/RESULTS.md removed (a measurement snapshot naming corpus projects, referenced by nothing).

Parser changes (the parser is in this repository now)

  • Fix for parser#182: two TypeScript or Python projects in one tree were analysed concurrently into one folder with truncating writes, the last to finish won, the other's IR vanished silently. A package plus a loose tests/ folder is two projects, so any tree with tests beside the package lost one. Each project now analyses into its own scratch folder and extract.ts merges per relation. The multi-program fixture, which had documented the defect, now asserts the fix; the evaluation harness's stage-siblings-separately workaround is shortfall-gated and stays dormant.
  • --per-language (opt-in): <ir>/java/, <ir>/typescript/, <ir>/python/, one folder per language present, holding only that language's tables. Default layout unchanged; the parser's own suites and the engine suites (flat) are green.

Tested on a mixed Java + TypeScript + Python project with libraries in all three: one command, three graphs, every by-construction edge present (dispatch fans to the overrides, library targets named, tests reach the client, main reaches run); with the JDK IR added as a library root the getClass().getSimpleName() sites resolve to named java.lang edges.

Acceptance (#468)

  • npm install && npm run build from the root builds parser and engine
  • Java 39/39, TypeScript 53/53, Python 15/15 from graph/test/ with no AXIOM_PARSER, the in-repo parser is the one exercised; goldens unchanged
  • bin/axiomcode all on a mixed Java+TypeScript+Python tree → three graphs from one command; on the parser repository itself (TypeScript, ~1,200 files: 7,716 methods, 24,875 sites, solved in 10 s) and on the two-service FastAPI project, schema_queries.callers_of answers on them
  • every test script finds the repository root by its marker (no ../../..); the move to graph/test/ is what exposed five level-counted paths that had silently gone one level wrong
  • parser suites in-tree: npx tsx parser/src/test/{java-extractor,typescript,python}-tests.ts, all pass. (npm test -w parser is vitest, which has no files to find; the tsx scripts are the parser's suite, as its README says.)
  • git log --follow graph/java/engine/resolution/virtual-dispatch.dl crosses the rename; the parser's history is native: git log -- parser/ lists its commits back to "Initial commit: AxiomCode Parser", and git log --follow / git blame on any parser/src/* file work (the history was rewritten with git filter-repo --to-subdirectory-filter parser before merging, so no subtree boundary exists, commit ids differ from the parser repository, messages/authors/dates do not)
  • no reference to ../Parser or to src/ as the engine root remains outside fixtures and the parser's own tree
  • build: prebuilt engine binaries, CI builds Linux/macOS/Windows on every merge and commits them under engine/binaries/ #455 rebased onto this (engine/binaries/ → binaries/, workflow paths), next

During the transition

While the parser repository still receives pushes, its new commits are brought over by the same rewrite (git filter-repo --to-subdirectory-filter parser on a fresh clone, then git merge); the 8 open parser issues can be transferred here with gh issue transfer at cutover (comments and labels follow; old URLs redirect), the closed ones stay readable in the archived repository. Archive that repository once contributors have moved.

swapnilpaliwal-sd and others added 30 commits August 22, 2026 16:04
…illed

The question was whether node.add(node).name() is a parser or an engine problem,
by analogy with Java where E -> concrete return type was engine work. Checking
rather than reasoning separated two answers.

The engine could already do it. Every hop is emitted: outer CALL -> its RECEIVER
child -> the inner CALL -> that call site, resolved METHOD -> core.base.Node.add
-> returnTypeName "Node" -> py_type core.base.Node. Six hops, all foreign keys,
no text matching. On the Java comparison that is genuinely the engine's to
compose.

But the parser also has a one-hop rule for exactly this shape, and it was
silently broken: buildInnerCallReturnIndex called returnedTypeOf with
`typesByName: new Map()` — a stub I never filled — so an annotated return could
never resolve, and `-> "Node"` looked up Node in nothing.

  sealed corpus: 25/28 linked, 3 correctly BUILTIN, zero missed
                 100% of the truly resolvable

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Resolution rate was the wrong measure. It bills the parser for layer-4 work
that §0.5 assigns to the engine, so a parser fully in spec scores badly.

The right question is whether the hops are present. For every unresolved
call site this walks what an engine would have to walk — receiver to
binding to assignment target to parent to value; self.x to a py_field row;
super() to ordered py_type_base rows — and returns FEASIBLE, BLOCKED with
the missing fact named, or UNDECIDABLE where no static answer exists.

On vendored asyncio: 9,789 sites, 2,848 resolved, 2,512 undecidable, 4,394
feasible, 35 blocked. Of the 7,277 statically answerable, an engine could
reach 99.5% from the facts as they stand. The whole parser bill is 28
self.x with no py_field row anywhere, 3 SELF calls with no bases to walk,
2 super() with no bases, and 2 py_type_base rows carrying no position.

The number swung 97.7 → 85.2 → 99.5 across three revisions of the checker
on identical parser output. First it assumed three verdicts instead of
checking them; then it over-corrected into being stricter than an engine,
looking for a field on the exact type rather than walking the MRO and
demanding a same-scope assignment for a global written one scope out. Both
mistakes produced a confident wrong answer.

So it is mutation-tested rather than believed: stripping ASSIGNMENT_TARGET
parents moves it to 2,223 blocked, deleting py_field to 760, deleting
py_type_base to 130, blanking bindingLinkHash to 2,223.

FEASIBLE means the facts are there, not that the engine will be right. It
still has to implement C3, the scope walk and argument flow correctly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Six reference kinds, each scored only where the referenced entity is
declared in the analysis root — a rate over all references measures the
corpus, not the parser.

On vendored asyncio: field refs, method refs and type refs in expressions
all link at 100%, variable refs at 99.9%, type references at 99.6%, base
classes at 95.8%, object creation at 92.9%, and method calls at 41.1%.
Overall 34,289 of 37,813, and zero mislinked anywhere: not one
referencedEntityHash points into the wrong table, so every edge that exists
is a true edge.

Two of these rows first read 0, for two different reasons and both mine.
The expression kind enum is NAME_REFERENCE and ATTRIBUTE_ACCESS, not NAME
and ATTRIBUTE, so the filter matched nothing — a zero that is
indistinguishable from "no such references exist". Then method references
were identified by name-matching the last dotted segment against declared
method names, which swept in 5,733 field accesses sharing a name with some
method and reported 232 mislinks. The parser already labels every reference
with referencedEntityKind; re-deriving that classification badly was the
whole defect.

Mutation-tested: repointing METHOD refs at a bogus hash surfaces 1,093
mislinks, and blanking bindingLinkHash drops the rate to 22.0%.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PEP 695 needed no 3.12 runtime: tree-sitter-python 0.21 already parses the
syntax, so `class Box[T]` reads on the pinned 3.10 toolchain.

Three grammar surprises, each found by reading the tree rather than the
reference. `type_parameter` is not unique to PEP 695 — Dict[str, int] uses the
same node under generic_type — so only a direct child of a class, function or
type alias counts, otherwise every subscript generic in the file would be
reported as a declared parameter. `*Ts` and `**P` both parse to splat_type with
no distinction, though a TypeVarTuple is a sequence of types and a ParamSpec a
whole parameter list, so they are told apart by source text. And PEP 696
defaults do not parse at all in 0.21 — that failure surfaces as a py_parse_gap
ERROR_NODE rather than being hidden, which is what the relation exists for.

Two bugs found while testing: a type alias nests its parameters under
type > generic_type rather than as a direct child, so the direct-child rule
silently found none; and once fixed it found them twice, because
`type Alias[T] = list[T]` has two generic_type children and only the left one
declares.

Verified on all eight forms: plain, multiple, bounded, tuple-constrained,
TypeVarTuple, ParamSpec, generic method, type alias.

Also confirmed with evidence that py_type_base.position preserves C3 order:
class CAB(A, B) gives A at 0 and B at 1, class CBA(B, A) gives B at 0 and A at
1, and CPython's __mro__ agrees with both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The claim was that the remaining method-call gap is engine work, with facts
already emitted for all but 35 sites. This checks it the way a consumer would:
read only the CSVs and ask, for each unlinked call, whether a path to a concrete
type exists using nothing but foreign keys.

On vendored-asyncio's 7,147 unlinked sites:

  DERIVABLE    219   engine work — 158 field carries a type or constructor,
                     58 inner call linked so its return is readable,
                     3 parameter annotation
  NEEDS_FACT    87   parser work
  NOT_STATIC ~4,300  2,118 builtin target with no py_method possible,
                     1,618 local whose type the source never states,
                     818 outside the analysis root,
                     525 attribute with no stated type

So the remaining gap is mostly not closable by anyone, rather than waiting on the
engine. Roughly 219 engine, 87 parser, and the rest needs the source to say
something it does not say.

My own instrument over-reported twice before I trusted it: it missed inherited
attributes by looking only at the declaring class, and counted sys.stderr.write()
as a self-attribute by not excluding module-rooted receivers. 211 -> 87.

The 87 are dominated by one real pattern: a classmethod factory that constructs
into a local and writes through it — `self = cls.__new__(cls); self._sslobj = x`.
The field extractor requires the write target to be the method's receiver, and in
a classmethod that is `cls`, so `self` there is an ordinary local. Recorded
rather than fixed, because the rule needs to say which locals count as the
instance, and getting that wrong invents fields on the wrong class.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ermost def

A3 found my two oracles contradicting each other on self inside a closure:

    def readline(self):
        def nreadahead():
            return self.peek(1)

emit_oracle looked only at self.fn[-1], so `self` matched no parameter of
nreadahead and fell through to NAME. They matched the other oracle and it
broke 47 files against this one.

SELF is right, and not only because it is the useful answer. symtable marks
this `self` is_free=True rather than a parameter, which is exactly why NAME
cannot work: §3's bare-name recipe terminates at PARAMETER → argument flow,
and self is never an explicit argument at any call site, so that walk
provably dead-ends. The invariant that makes SELF safe — "exactly the
enclosing class or a subclass" — survives closure capture, because a
closed-over name cannot be rebound to another type without becoming local.

The rule now walks outward and stops at the first scope that binds the
name. SELF only when that scope is a function declared directly in a class
body and the name is its first positional parameter; CLS when that
parameter is named cls or the method is a classmethod; NAME when an
intervening scope takes it as its own parameter or writes it locally.

That last clause is where the two oracles still differ, and there the other
one is unsound. `def inner(self)` called as inner(None), and `self =
object()` inside a nested def, both bind something that is not the
instance. Answering SELF there manufactures a false edge, which costs more
than the missing edge it replaces.

253 call sites flip NAME → SELF on the torture corpus, 7,710 → 7,963.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A3 built an independent gate that follows each unlinked call to a concrete
type using only FKs, and got 219 derivable where mine said 4,394. Checking
rather than conceding, every measurement moved toward theirs.

The bias had one shape in three places. Of 840 NAME receivers with a fully
intact hop chain, only 392 reach a value from which a type is derivable —
the other 448 have every fact present and terminate in nothing, 316 of them
assigned from a call whose return type nobody wrote down. The bare-name
branch counted a site feasible when a binding of that name existed anywhere
in the corpus, which inflated 991 of 1,157, because a binding is not a
callee. And FOR_TARGET, WITH_TARGET and TUPLE_UNPACK were credited because
bindingOrigin names the mechanism, when knowing a name is bound from an
iterable is not knowing the element type.

Fixed, with a transitive value walk that is depth-bounded and cycle-guarded:
4,394 → 3,946 → 3,221 → 1,473, and NOT_STATIC is now its own verdict at
1,930 rather than being folded into engine work. That distinction is the
point — engine work implies recoverable and this mostly is not.

1,473 still exceeds their 219. The residue is 386 parameter bindings and
192 attribute chains this does not yet follow to a type, plus bare-name
calls that link to a function where theirs requires a type. Tuning further
would be fitting to their answer rather than to the truth, so it stops here
with the disagreement stated.

On the parser bill their 87 beats this gate's 35 and is adopted: this one
cannot see the cls.__new__(cls) factory pattern at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Other agents need the build to tell them WHICH FACT MOVED, not that a
percentage drifted. "services/inheritance.py py_call_site.resolvedCalleeKind
METHOD -> UNRESOLVED" is a bug report; 91.2 falling to 90.8 is not.

src/test-data/python/verified holds 8 files admitted by measurement rather
than by eye: every call site in each links to a declared entity or is a
builtin, so the ceiling is exactly 100% and any later unresolved call is a
regression. 1,358 facts across 16 relations are frozen in _golden/.

Two rules keep it honest. Nobody hand-writes expected facts — the parser
proposes and the golden records, since a hand-written expectation only
tests whether its author and the implementer read the spec the same way.
And --bless re-runs the CPython oracle first and refuses to freeze output
it disagrees with, so a bug cannot be pinned in place where every later fix
would read as a regression.

Package structure is preserved. The first cut flattened core/base.py to a
single filename, which breaks `from .base import ...` — the corpus would
have kept passing its own goldens while testing different code.

The diff drops hash columns from the EXPLANATION but not the comparison.
Renaming one method cascaded four scope hashes and truncated the actual
change off the end of every line.

PEP 695 gets its own gate because the regimes are not comparable: the same
source yields a different scope tree under PY3_0_11 and PY3_12_PLUS, which
is why emissionRegime is in py_module's PK. CPython 3.12's ast is the
oracle, and the parser agrees on all 12 type parameters — name, position,
bound and ownerKind, including TypeVarTuple and ParamSpec.

TYPE_PARAM, TYPE_ALIAS and TYPE_PARAM_BOUND exist in PythonScopeKind but no
extractor emits them, so §2.20's pyScopeLinkHash points at the enclosing
scope instead of a wrapper. The gate asserts that gap rather than passing
over it: an enum value is not an emitted fact.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
py_block was the largest zero-coverage relation: 17 columns, the youngest code,
and the one data flow depends on. Now checked against ast for the set plus kind,
condition text and caught exception types.

It found two systematic defects immediately, both of which I had fixed before in
other relations and not here. The block's endColumn came from the tree-sitter
block node, which swallows a trailing comment — `return -1  # incomplete` pushed
the span past the code on 32 of 4,220 blocks, identical to the py_type.endLine
defect and fixed the same way. And conditionText kept the parentheses of
`if (a and b):`, so the same test read differently by spelling, and disagreed
with the FK I had already unwrapped for that reason.

  4,215 / 4,220 matched, 0 content disagreements over 6,335 comparisons

The service version was never hashed. Java takes an identifier and hashes it;
the Python analyzer took `serviceVersionLinkHash` from options and wrote it raw,
so a column named ...LinkHash carried an unhashed string and a rule ported from
Java would compare a hash against a literal. Now derives with the same helper and
prefix — "v1.2.3" gives SERVICE_VERSION_9479983c… both ways. py_module's PK
includes this value, so every module hash changes; that is correct rather than
unfortunate, but goldens need regenerating.

Also documents PythonTypeParameterKind as deliberately unemitted: §2.20 declares
fourteen columns and none carries a parameter kind, so *Ts and **P are
indistinguishable in the facts today. The extractor recovers the distinction from
source text and it needs somewhere to put it — raised with A0 rather than
deleted, since the distinction is real.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Python had grown seven entry points where Java has one, and the cost was
not tidiness: nobody could say what "the tests pass" meant, because no
single command ran them all. python-tests.ts is that command.

Suites are ordered by what they prove, strongest evidence first, because a
failure early makes everything after it uninterpretable — if the oracle
disagrees with CPython there is no point asking whether the parser agrees
with the oracle. Oracle self-test, then schema guard, then goldens,
closed-world, PEP 695, and the torture sweep. 6/6 pass.

src/test-data/python/categories mirrors src/test-data/java: one directory
per entity kind, twelve categories covering every methodKind, every
fieldOrigin, every import shape, every block form, all four comprehension
scopes, two lambdas on one line, and a layered integration sample.

Every file is admitted by measurement: each call site links or is a
builtin, enforced by golden-gate --select. Getting there took four rounds —
the first cut passed 6 of 12, and the failures were all real gaps
(return-type flow, two-hop attribute chains, unflagged builtin exceptions).
Rather than weaken the rule, each file now avoids the unresolvable CALL
while keeping the construct under study, with a comment saying which and
why. The corpus is a tripwire, and a tripwire that depends on known-missing
inference is just a second copy of the backlog.

Goldens grow from 1,358 facts over 8 files to 3,696 over 20.

Also syncs the four ratified enums into the schema: scopeKind gains
TYPE_PARAM_BOUND with the three-level nesting written out, py_type_parameter
is no longer marked deferred and documents its kind and variance enums, and
py_comment leaves the not-yet-emitted list. 54 of 56 enums now compared.

Deletes src/test/python-scratch: 47 files, referenced by nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two findings from A3. The first does not hold: LINE_COMMENT is in §2.17 and
is the first value of eleven, matching PythonCommentKind exactly, which is
why the enum guard passes it. The list wraps across three markdown lines,
so a grep that catches only the continuation lines misses it — worth noting
because §2.15's edgeRole spans nine.

The second is right, and it is my bug from the enum sync. I wrote a `kind`
enum block into §2.20 for a column the relation does not have. The
consequence is real: `class Variadic[T, *Ts, **P]` emits three rows
identical but for paramName and position, where CPython's ast returns
TypeVar, TypeVarTuple and ParamSpec. A TypeVarTuple is a sequence of types
and a ParamSpec a whole parameter list, so Callable[P, R] and tuple[*Ts]
cannot be told apart.

The worse half is that my guard validated the phantom. It compared the
documented enum against a real TypeScript enum and passed, having never
asked whether the column exists. gen_decls now cross-checks every enum
block against the relation's declared column list and rejects one written
for a column that is not there — it catches this case immediately.

The doc now states the gap instead of describing a column that isn't there,
and carries the proposal: kind at position 12, shifting serviceVersionLinkHash
to 13 and the key to 14. That keeps the convention that the version hash sits
immediately before the key, and only the trailing pair moves. It is safe now
and never again — nothing consumes this relation positionally and no golden
holds a row of it, because verified/ is a 3.10 corpus with no PEP 695 in it.

Left red pending ratification.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Ratifies A0's proposal. Without this column `*Ts` and `**P` were
indistinguishable in emitted facts: tree-sitter gives both `splat_type`
with a bare identifier under it, so the TypeVarTuple/ParamSpec
distinction lives only in the source text. PythonTypeParameterKind
existed as an enum with no column to live in.

Inserted rather than appended, on A0's reasoning that the insert is free
only until a golden freezes a row of this relation -- after which the
column would have to go on the end and break the convention that the
version hash sits immediately before the key. Nothing consumes
py_type_parameter positionally today and verified/ is a 3.10 corpus with
no PEP 695 in it, so the window is open now and never again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The strongest oracle available for py_scope and py_binding -- symtable IS
what the compiler consults, so a disagreement there is a defect and not a
modelling preference -- existed only as a throwaway script. It scored
10,769 files at 99.9% clean and then the script was gone, leaving a
number nobody can reproduce or re-run against a change. An instrument
that cannot be re-run is not an instrument.

Two ordering traps are handled on the oracle side, both of which would
otherwise be charged to the parser:

  - A SymbolTable on 3.10 carries lineno and no column, so columns must
    come from a separate ast walk that gets zipped with the symtable
    children. But CPython does not create comprehension blocks in source
    order: symtable_handle_comprehension visits the outermost iterable in
    the ENCLOSING scope before entering the block, so in
    `[a for a in [b for b in q]]` the inner listcomp's block is created
    first and the two are siblings. A source-ordered zip pairs each scope
    with its neighbour's position and manufactures mirror-image
    BIND_MISSING/BIND_SPURIOUS pairs.

  - symtable.c visits Try as body, ORELSE, HANDLERS, finalbody -- the
    else clause before the except clauses, which is neither source order
    nor ast field order.

The zip asserts kind and name agree at every position and reports
GATE_DESYNC rather than blaming the parser when they do not; that is how
the Try case surfaced.

is_annotated is named as unchecked rather than quietly skipped, because
ANNOTATED_ONLY is a different predicate and wiring the two together
would have scored a disagreement as a pass on every annotated
assignment.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…lias

tree-sitter-python applies PEP 695's soft `type` keyword greedily, so the
ordinary idiom for setting an attribute on an object's class parses
cleanly as a type_alias_statement: the leading `type` becomes the
keyword, `(obj).attr` is accepted as the alias name, and the assigned
value is presented as a type. Nothing is marked as an error.

The damage was larger than it first looked. The `call` node is not
mis-shaped, it is ABSENT, so `type` has no identifier node anywhere and
symtable reported a global reference we did not. Worse, the whole
statement fell through to a walk that visits statement children, and a
type_alias_statement has none -- so a nested `compute(val)` on the right
hand side produced no expression rows and no call site either. The
statement was silently invisible end to end.

Found by the symtable gate, as six BIND_MISSING rows all named `type`
across unittest, test_urllib and test_contextlib. `type` was the only
builtin ever missing while len, print, isinstance and int were all
present, which is what pointed at a name-specific grammar path rather
than a filter.

Recovered: the assignment target, the annotation, the value, and every
nested call. Not recoverable: the outer `type(...)` call itself, since
there is no node to mint a call site from. That is emitted as a
py_parse_gap with the new SOFT_KEYWORD_MISPARSE kind and the existing
MISPARSED_SILENTLY disposition, which is the case that relation was
built for -- absence here is otherwise indistinguishable from the idiom
not appearing.

The three-way decomposition lives in one helper rather than in each
extractor. An annotated target nests a `constrained_type` in the middle,
because `x: y` is also PEP 695 bound syntax, and digging to the wrong
layer marks the ANNOTATION as an assignment target -- which it did until
the helper existed.

unittest goes from 4 dirty files to 43/43 clean, 3273/3273 scopes,
14984/14984 bindings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The language reference is explicit that all identifiers are converted to
NFKC while parsing, so two spellings can be one name: the mathematical
script Unicode is `Unicode`, and MICRO SIGN U+00B5 is GREEK SMALL LETTER
MU U+03BC. Keeping raw source text did not just miss the link between a
definition and a use spelled differently -- it INVENTED one. `Unicode =
1` followed by a script-spelled reference produced a local binding plus a
spurious GLOBAL_IMPLICIT for a global that does not exist, where symtable
has a single local. A false edge is worse than a missing one: a data-flow
query that follows it gets a confident wrong answer.

Normalisation goes before private-name mangling because CPython folds in
the tokeniser, so mangling already sees a normalised name; the other
order would leave a fullwidth `__x` unmangled.

Found by the symtable gate on test_unicode_identifiers.py.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PEP 572 binds a comprehension's walrus target in the nearest enclosing
function, and we modelled that as free/nonlocal unconditionally. When the
enclosing function declares the name `global`, symtable propagates the
declaration into the comprehension instead:

    def f():
        global G
        [G := 1 for _ in range(1)]   # listcomp: G global and declared

We reported free and nonlocal there, which points a consumer at a
function-local cell that does not exist while the name really resolves to
a module global.

DEF_LOCAL stays set alongside DEF_GLOBAL, and the owner still records the
target. The walrus does ASSIGN, so symtable has is_assigned true while
is_local is false; dropping DEF_LOCAL fixed the four scope predicates and
broke the assignment one, which the gate caught immediately.

Also fixes the gate itself: it classified scope kind from symtable's
name, but get_type() only distinguishes module/class/function, so
comprehensions are identified by the synthetic names listcomp/setcomp/
dictcomp/genexpr -- and test_peepholer.py defines `def listcomp():`,
whereupon the oracle called a function a comprehension. Kind now comes
from the ast node type, which cannot be spoofed.

CPython's own test suite: 740 of 740 comparable files clean, 44559/44559
scopes, 222396/222396 bindings. The 8 excluded are CPython's
deliberately-malformed encoding and __future__ fixtures, which symtable
refuses too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Expression positions inside a MULTI-LINE f-string are miscomputed by the
interpreter. Reduced to two lines that differ only in quoting:

    a = f'''
    void impl({', '.join(x.d() for x in ar)});
    '''
    b = f"void impl({', '.join(x.d() for x in ar)});"

On the single-line form ast puts the genexpr at the `(` of `join(`, which
is right and is exactly what tree-sitter reports. On the multi-line form
it points into the middle of `x.d()` at a closing paren. CPython fixes up
f-string sub-expression locations only for the single-line case.

This produced a SCOPE_MISSING and a SCOPE_SPURIOUS for the same scope
across torchgen and fontTools, and reading them as parser defects would
have led to "correcting" positions that were already right. Those scopes
are now excluded from position matching and the excluded count is PRINTED
on every run, so the gate reports what it cannot check instead of quietly
dropping it or blaming the wrong side. Single-line f-strings stay in;
they are correct, and they are the common case.

torchgen: 41/41 clean, 12 positions unverifiable, 7994/7994 bindings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Inside a comprehension block CPython's order is generators[0].target,
generators[0].ifs, then (iter, target, ifs) for each remaining generator,
and only then the element. We walked the element first, in source order.

Source order is right often enough to hide this: it diverges only when a
clause and the element BOTH open a scope, and then it swaps their
scopeOrdinals. fontTools' subset_glyphs is exactly that shape -- a dict
comprehension with a listcomp in its value and a genexpr in its `if` --
so the two came out in the wrong order and each was reported at the
other's position.

Also: a dict comprehension visits its VALUE BEFORE ITS KEY.
symtable_handle_comprehension is called with elt=key and value=value and
emits `if (value) VISIT(value); VISIT(elt);`.

Both verified on 3.10.4 rather than read off the C, by giving each
position its own lambda and reading back the resulting block order:

    { (lambda: KEY)() : (lambda: VAL)() for x in y }   -> VAL, KEY
    [ (lambda: ELT)() for x in y if (lambda: IFF)() ]  -> IFF, ELT
    [ (lambda: ELT)() for x in y for z in (lambda: IT2)() ] -> IT2, ELT
    [ 1 for x in y if (lambda: IFF)() for z in (lambda: IT2)() ] -> IFF, IT2

fontTools: 276/276 clean, 6981/6981 scopes, 46334/46334 bindings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Twenty-two columns across the two relations had nothing checking them, on
the relation the schema treats as Python's answer to Java's annotations.
ast is strong ground truth here: a decorator list is a plain list of
expressions on a definition, so presence, order, the dotted name, the
call/bare distinction and every argument are stated outright.

Three parser defects, all of the same family -- a column that is empty or
wrong where it could be right:

  - argumentCount was set only for the call form, so `@property` reported
    "" and "zero arguments" was indistinguishable from "not known". A
    decorator that is not a call takes zero arguments; that is a fact.

  - dottedPath was empty for the BARE form, so `@contextlib.contextmanager`
    was joinable by dottedPath and `@classmethod` was not, and every
    consumer had to special-case the single segment.

  - dottedPath was the NORMALISED TEXT of the expression, so PEP 614
    decorators put `null(null)` and `[null][0].__call__.__call__` into a
    column whose only purpose is to be joined against a resolvable
    entity. Nothing can ever match those, so they were not links but the
    appearance of links, and a consumer counting resolvable decorators
    would have counted them. It is now a name chain or empty, and empty
    is a true and checkable statement. kind stays SYNTACTIC and
    dottedPath carries resolvability -- two different facts, and
    collapsing them loses one.

Three findings that were the GATE's fault and are worth recording,
because each nearly led to "fixing" correct code:

  - The context enum spells the class case TYPE_DECLARATION, not CLASS.
  - Keying decorators by (owner name, position) pooled every `test#0` in
    testpatch.py across a dozen classes. A decorator's own start line is
    unique within a file, since two cannot share a line.
  - A list argument is deliberately SPLIT into one row per element
    sharing the argument's position and distinguished by arrayIndex, as
    Java does for array-valued annotations, so `@jump_test(2, 1, [1, 1,
    2])` is three arguments in five rows.

isKeyword settled as "written as name=value", so `**kwds` is not one: it
has no keyword name, and calling it one lets a consumer look up the value
for keyword X and find a row with no X. The star form stays recoverable
from argumentValue.

CPython's test suite: 744 of 744 comparable files clean, 4250/4250
decorators, 41870 column comparisons.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
On 600 stdlib modules the contexts ISINSTANCE_TYPE, ISSUBCLASS_TYPE and
RAISE_TYPE were emitted ZERO times, and so was the owner kind
EXPRESSION, while `isinstance(x, Foo)` and `raise ValueError(...)` appear
in nearly every file. All four are declared in the enums, so nothing was
waiting on a schema decision -- the facts were simply never produced. A
dead enum value is invisible: it looks like the construct does not occur.

These matter more than their column count suggests, and for the reason
that decides which bucket an unresolved call site falls into. A receiver
whose type cannot be read off its declaration is often pinned exactly
once, by an `isinstance` guard, and that guard is the only static
evidence there will ever be. Emitting it moves sites out of "not
knowable statically" and into "joinable" without inventing anything: the
new references are resolved by the same pass that resolves an
annotation, so each either names a type in the corpus or stays empty.

Measured over the same 600 modules: py_type_reference goes from 2242
rows to 5652, and 531 of the 3410 new rows already resolve to a concrete
in-corpus type in a SINGLE-FILE pass, before cross-module resolution
runs at all.

The owner is the CALL expression, not the narrowed variable. Both are
defensible, but the call carries the position, and a consumer that wants
the variable reads the call's first argument -- whereas owning by the
variable would throw away WHERE the narrowing holds, which is exactly
what makes it sound to use.

`isinstance(x, (A, B))` is flattened to one reference per alternative,
since each is separately resolvable and a composite would be neither.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`except ValueError as e:` is the one narrowing construct that types a
VARIABLE outright -- no inference, no guard to reason about, no
flow-sensitivity to get right. Inside that handler `e` IS a ValueError.
EXCEPT_TYPE was dead in the emitted facts, so `e.detail` had nothing
pointing at a type and no join could recover one.

The reference is therefore owned by the BINDING rather than by an
expression wherever the handler binds a name, which is what lets a
consumer resolve a use of `e` without having to notice that an except
clause was involved at all. Where there is no `as`, the type is still
recorded and owned by the exception expression.

`except (A, B) as e:` flattens to one reference per alternative, as
isinstance does: `e` is one of them and each resolves separately, while a
composite would resolve to nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`class Foo: x: Bar` produced a type reference owned by the class-body
BINDING with context VARIABLE_ANNOTATION, and nothing owned by the field.
FIELD_TYPE and owner kind FIELD were both dead. So "what type does field
x of Foo declare?" could not be answered by joining from py_field: a
consumer had to know that a class attribute is ALSO a binding, find the
class body scope, and match on name -- a join nobody should have to
discover, and one that returns nothing rather than failing when got
wrong.

The tree matches Java's convention exactly, which was the thing to get
right rather than invent. Java documents
`Map<String, List<? extends Number>>` as Map d0, String d1, List d1,
wildcard d2, Number d3 with parentReferenceHash chaining, and uses
TypeRefContext.FIELD_TYPE with ReferenceOwnerKind.FIELD. Python now gives
`Dict[str, List[Optional[Bar]]]` the same shape:

    d0 p0 SUBSCRIPT Dict
      d1 p0 NAME      str
      d1 p1 SUBSCRIPT List
        d2 p0 OPTIONAL  Optional
          d3 p0 NAME      Bar      -> resolved to PY_TYPE

Verified on Dict/Callable/Tuple/Union fields: 0 orphan parents and 0 rows
whose depth is not the parent's plus one.

`self.w: Baz` in a method body is covered too, since the annotation is
recorded per FIELD rather than per class-body statement.

The binding-owned duplicate is suppressed inside a class body. Emitting
both put two rows on one annotation with different owners, which
double-counts every annotated attribute in any tally of how many type
references resolve; the field is the canonical owner and the binding
stays reachable through py_field.

Also adds link-coverage.ts, which reports REAL coverage by splitting
unresolved links into RESOLVABLE (the name matches an entity this run
emitted -- the parser's gap) and EXTERNAL (nothing by that name exists --
not knowable). A raw resolved/total ratio measures how many imports a
project has, not the parser.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Most of the "87 parser-owned sites" was a measurement artifact. The
feasibility gate treated ANY two-segment attribute receiver as an
attribute of the enclosing class, so `mock.call_args.get()` and
`result.errors.append()` were looked up in the enclosing class's MRO,
found nothing, and were charged to the parser. On unittest that was 262
receivers against 134 genuinely self-rooted ones. It also double-booked
"py_field exists but states no type" as BOTH NOT_STATIC and NEEDS_FACT,
when the source simply never states a type and there is no fact to emit.

Correcting the meter took 223 -> 6, and the 6 were two real causes:

  - `self.__dict__['x'] = v`, usually through a local alias:

        __dict__ = self.__dict__
        __dict__['_mock_children'] = {}

    This names attribute `x` exactly as `self.x = v` does, and classes
    that must bypass a custom __setattr__ write this way. None of those
    attributes had a py_field row, so no call on them could link to
    anything. The key must be a string LITERAL, so this is decided
    statically -- a computed key is skipped rather than guessed. Now
    emitted, closing 4 of the 6.

  - The MIXIN pattern, closing the other 2. `class CallableMixin(Base)`
    uses `self.call_args_list`, which exists only once `class
    Mock(CallableMixin, NonCallableMock)` combines it with a SIBLING
    base. It is not on CallableMixin's MRO and cannot be: there may be
    several combinations supplying different types, so linking it means
    picking one, which is a guess wearing the shape of a fact. Given its
    own verdict rather than folded into NOT_STATIC, because "unknowable"
    and "knowable only from the other side of a combination" are
    different and a reader should be able to tell them apart.

unittest: NEEDS_FACT 223 -> 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Resolving a virtual call to one target is the engine's job. What the
parser owes is the substrate CHA, RTA and interprocedural flow consume,
and asserting that the substrate is sufficient is worth nothing. This
builds all five layers using only the emitted CSVs -- no parser objects,
no in-process state -- and reports what is constructible and what is
blocked for want of a fact.

On unittest, 548 types and 3227 methods:

  1. C3 linearisation      548/548 classes, deepest MRO 8, and ZERO bases
                           missing a position. Order is what makes
                           class C(A,B) different from class C(B,A), and
                           C3 is undefined without it.
  2. Override edges        562 built across the hierarchy. So overrides
                           are DERIVABLE and no override column is
                           needed -- worth stating, because adding one
                           was the obvious reflex and would have been
                           redundant emission.
  3. CHA, self/cls         3788 sites, 98.2% get a candidate set, 97.5%
                           monomorphic.
  4. RTA                   274 types constructed somewhere, narrowing the
                           candidate set.
  5. Argument -> parameter 8236 edges bound, including 140 that need the
                           constructor hop type -> __init__, since a
                           constructor call resolves to the TYPE and a
                           type has no parameters.

The split that matters is receiver kind. For self/cls dispatch the
receiver's static type IS the enclosing class, so CHA runs with no
inference at all -- hence 98.2%. For any other receiver the type must
come from flow first, and using the enclosing class there is simply
wrong; it is what made the earlier feasibility numbers overstate parser
gaps, so the gate refuses to do it and counts those separately: of 3371
non-self sites, 884 are already linked, 341 have type evidence the parser
emitted for the receiver, and 2146 have no static evidence anywhere.

That last number is the honest ceiling, not a defect. It is duck typing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ectly

python-work held two unlike things: the contract and the process. The
contract has moved to src/schema/python — the frozen schema, the generated
.dl, the generator, and the two decision-procedure specs. Everything else
is gone: 480K of coordination logs, 17M of staging corpora that
torture-corpus.ts and vendor-corpus.ts regenerate from the pinned
interpreter, and the process notes.

gen_decls resolves its own paths now. Run from the new directory it compared
ZERO enums and still exited 0 on the arity check, which is the worst kind of
pass — a guard reporting success because it looked at nothing.

Also deletes nine transient .*-out directories and gitignores the pattern.
Every gate writes one; none is read across runs.

The move exposed a change the goldens caught immediately. Decorator
dottedPath went from HANDLERS[0] to HANDLERS, which is right — fullText
still carries @handlers[0], so nothing is lost. Type references gained
eleven EXCEPT_TYPE rows, which is new coverage. But every BARE decorator now
carries argumentCount=0 where §2.12 says "", and 0 asserts "called with
nothing" rather than "never called".

That last one could not be caught by re-blessing, because a golden only sees
UNINTENDED change and would have laundered the violation into the baseline.
So stated rules are now checked directly, every run, apart from the goldens.
The gate reports "goldens hold, but the spec does not" and fails.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two things, the first a correction A0 filed and was right about.

argumentCount goes back to "" for a decorator that is not a call, per
§2.12. I had changed it to '0' reasoning that "zero arguments" is a fact
and an empty column is invisible to a query. That was wrong: 0 asserts
CALLED WITH NOTHING, and `@property` was never called at all -- there is
no argument list to have a length. `@f()` and `@f` differ in exactly
this, so collapsing them makes the column unable to express the
difference, which is the reverse of the problem I thought I was fixing.

Then the measurement that matters. A closed-world corpus makes "every
call site must link" mechanically true but says nothing about whether
each links to the RIGHT target, and a corpus written by the same author
as the parser cannot settle that. So the corpus is GENERATED -- 302
files, 9480 lines, 60 packages, no import outside the tree -- and the
verdict comes from running it under sys.settrace, which reports for every
function entry the caller's frame and the callee's code object. That is
precisely the claim resolvedCalleeHash makes, so CPython adjudicates it.

2321 runtime call edges:

    MATCH      1731   74.6%
    PROPERTY    120    5.2%   attribute access, no call node -- correct
    WRONG         0    0.0%
    UNLINKED    420   18.1%
    NO_SITE      50    2.2%

    precision of emitted links: 100.0%

Zero wrong targets is the number worth having. Every link the parser
emits is the one CPython takes, so nothing downstream is poisoned; the
gap is recall.

Three corrections to the gate itself, each of which had it blaming the
parser: py_call_site carries no filePath and joins through
pyModuleLinkHash; py_module.filePath is ALREADY relative, so running
path.relative on it again walked out of the tree and matched nothing
(100% NO_SITE); and a @Property getter runs on ATTRIBUTE ACCESS, so it
correctly has no call site while CPython still reports entering it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both were found by the runtime gate rather than by reading code, and both
are ordinary Python that any real project contains.

1. `import pkg.mod` then `pkg.mod.Klass()` did not resolve.

   The dotted-receiver path only tried `importedModules.get(segments[0])`,
   which for `import pkg.mod` binds the PACKAGE `pkg` -- and `mod` is a
   submodule, not a member of the package's exports, so the lookup failed
   and the next line declared the call external. It is not external: the
   module is in the corpus. Every call through one of the two ordinary
   import forms was being written off as third-party.

   Now the whole receiver is tried as a module path, longest prefix
   first, so `a.b.c.D()` prefers module `a.b.c` over module `a.b` plus an
   attribute walk -- which is what Python itself does.

2. A return annotation was resolved in the CALLER's namespace.

   An annotation is written in the scope of the method that carries it,
   so `-> "B"` on a method of p/b.py means p.b.B. A caller doing
   `from p.b import B as Alias` has no name `B` at all, so the lookup
   failed and the fluent chain died: `Alias.of(1).describe()` was
   unresolved while `B.of(1).describe()` -- the same call reached by a
   different name -- resolved. Now the declaring module's namespace is
   consulted first, with the caller's as fallback.

Against CPython at runtime on the 302-file corpus, 2321 traced edges:

    MATCH     1731 -> 1851   74.6% -> 79.8%
    WRONG        0 ->    0    precision stays 100%
    UNLINKED   420 ->  300

What remains unlinked is three shapes, and none is a missing fact: chains
of two or more hops, where every individual hop IS linked and walking
them is the engine's job by the boundary already documented here; mixin
dispatch, where the attribute is supplied by a sibling base and no single
answer exists; and loop variables drawn from a heterogeneous list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
"UNLINKED: 300" is not a number anyone can act on. It mixed calls a
consumer reaches by joining facts we already emit with calls no static
analysis will ever reach, and only the second kind is a ceiling. The
split is computed by PERFORMING the join a consumer would -- reaching
assignments to the receiver, then the type each right-hand side produces
-- rather than by asserting the join is possible.

  UNLINKED split (300):
    DERIVABLE from emitted facts     180   60.0%
          120  chain: inner call resolved, read its return type
           60  reaching-assignment: single receiver type
    NOT DERIVABLE                    120   40.0%
           60  local: no type reaches the receiver
           60  self: type known but does not declare it (mixin/sibling base)

Three defects in the check itself, each of which made the parser look
worse than it is:

  - It was not TRANSITIVE. `p.chain().chain().report()` has an unresolved
    inner call, but that call's own receiver is derivable and its return
    type follows. Stopping at the first unlinked hop reported the whole
    chain as unreachable when only its first hop needed a join -- the
    measurement saying "impossible" about work that is merely not done.
    Recursion is depth-bounded at six.

  - A bare `self` receiver was routed through reaching assignments,
    which looked for an assignment to `self`, found none, and called the
    type unknown -- when it is the one type always known. It is the
    enclosing class.

  - `cls(value)` in a classmethod factory is a constructor call whose
    callee NAME is `cls`, not a class name, so 50 sites were reported as
    having no row when the row existed and resolved. NO_SITE is now 0.

Final picture on 2321 runtime edges: 1901 linked by the parser (81.9%),
120 property accesses correctly carrying no call node, 180 derivable by
an engine join, 120 genuinely unresolvable -- duck-typed containers and
mixin dispatch where no single answer exists. Precision stays 100%.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
swapnilpaliwal-sd and others added 4 commits September 13, 2026 15:36
@type; recall absences (#179)

Fixes #176. Squashed from js; corpus 41/41, TypeScript 44/44, Python 21/21; scrub 1 hit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
 #163 (#180)

A REGRESSION I INTRODUCED in #149. The member group key answers "which member
of this owner is this", and I applied it to any callable emitted inside a type
context. An arrow assigned to a `const` inside a method is not a member — it is
an expression that merely occurs inside the class, and its `tsTypeLinkHash`
names the class only because that is the enclosing emit context.

Every arrow shares the sentinel name `<arrow>`, so
`md5(ownerGroupKey || "<arrow>" || false)` was ONE value for every arrow in a
class body:

  L5:22  L6:30  L9:22  L10:30  L13:25   -> one declarationGroupKey, overloadIndex 0..4

Five distinct callables across three different methods, emitted as though they
were overloads of each other, so a consumer commits a call through one `const`
to a different method's arrow. The module-scope arrows beside them were always
correct, because there is no owner there at all — which is the shape the
class-scoped ones now match.

`isClassElement` / `isTypeElement` are tsc's own predicates for membership, so
the fix is that test rather than a list of kinds to exclude. Class methods,
accessors, constructors, and the call and construct signatures of a reopened
interface all keep the key #149 gave them.

A static block is the one ClassElement excluded: tsc gives it no symbol, and
two in one class collide on `<static-block>` for exactly the reason arrows did.
Found while writing the test, not reported.

The test asserts BOTH directions, because a predicate can be wrong either way:
non-members carry no key, and a real two-signature class overload still shares
one. It also asserts the invariant the collision broke — no group key is shared
by differently-named callables. Reverting the fix reports four arrows, two
static blocks, and 6 groups where the fix leaves 4.

45/45 checks pass; the #149 and #93 checks still pass unchanged.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
… (#181)

`export default class NamedClass {}` was recorded as `default`, so a NAMED
default export and an ANONYMOUS one came out as two rows identical apart from
the module — while `ts_export` had kept both names all along
(`exportedName=default`, `localName=NamedClass`).

The binding is not wrong, and that is the point. `default` is
`InternalSymbolName.Default`, and tsc agrees outright: for
`export default class NamedClass {}` the symbol's `escapedName` is literally
`"default"`, while `declaration.name.text` is `NamedClass`. Two names, and the
parser was recording one of them twice.

So they are separated at the one place they differ. `escapedName` and
`declarationGroupKey` keep the binding's `default`; `name`, and the
`qualifiedName` built from it, take the declaration's own identifier when it
has one. An anonymous `export default class {}` still reads `default`, because
nothing else is available.

  named-class.ts  name=NamedClass            qn=…#NamedClass  escapedName=default
  anon-class.ts   name=default               qn=…#default     escapedName=default
  named-fn.ts     name=namedDefaultFunction                   escapedName=default
  anon-fn.ts      name=default                                escapedName=default
  control.ts      name=ControlClass                           escapedName=ControlClass

The merge partition is the regression that mattered and it is unmoved: 1,351
declaration sites in 1,304 groups, 27 merged, set equality with tsc in both
directions.

The test asserts both halves, because the fix separates two things that had
been one: the own name appears, AND `escapedName` is still `default`. Reverting
reports the name, the qualifiedName, the indistinguishability, and the
function.

46/46 checks pass.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The engine — rules per language, pipeline, bundle — moves from src/ to graph/, so the
repository's top level names its parts (graph, and next: parser, binaries, skills, bin).
Every repository-path reference is rewritten (tsconfig, package.json, constants/paths.ts,
the executor, the suites and tools); fixture directories named src/ inside test cases are
untouched. The --debug CSV output directory is renamed from graph/ to csv/: .gitignore
ignored graph/ as an output directory, which would have hidden the new source folder.
@swapnilpaliwal-sd swapnilpaliwal-sd added build Build, packaging and developer setup enhancement New feature or request platform OS / toolchain portability labels Sep 14, 2026
@swapnilpaliwal-sd swapnilpaliwal-sd self-assigned this Sep 14, 2026
The parser's 312 commits are rewritten under the parser/ prefix (git filter-repo
--to-subdirectory-filter) before merging, so every historical commit already lives under
parser/: `git log -- parser/`, `git blame` and `git log --follow` on any parser file work
natively, with no subtree boundary. Only the commit ids differ from the parser repository;
messages, authors and dates are unchanged.
…ph, README for the new layout

The parser is now built from parser/ by the root `npm run build` (workspaces), and every
suite's default parser path is parser/dist/index.js in this repository instead of a sibling
checkout. bin/axiom-graph is the one command — parse, solve, bundle — leaving
<out>/graph.sqlite (IR and scratch under <out>/.intermediate). skills/code-graph/SKILL.md
tells an agent how to build the graph and how to query it: read schema_guide, bind the
tested schema_queries, filter on tier, treat a negative answer as a lower bound.
…ayout fixture path wrongly rewritten; drop test/java/README.md and the TypeScript results snapshot

check_staging.py, check_arity.py, literal_gate.py and empty_relation_lint.py built the
engine path from ('src', lang) and were missed by the rewrite; they point at graph/ now.
jdk-ir-stamp-test.sh's fixture mimics the JDK source layout (src/java.base/share/classes)
and had been rewritten as if src/java were the engine — reverted. The Java suite README
is removed (its run instructions move to the root README's Tests section) and so is
graph/typescript/RESULTS.md, a measurement snapshot naming the corpus projects that
nothing referenced.
test/ asserted the engine's edges (each case parses a fixture and solves it), and parser/
already carries its own tests, so each package now owns its tests: graph/test/<lang>/ and
graph/test/tools/. Every root derivation in the suites and tools gains one level; the
parser's tests are untouched. skills/ is removed — it will be authored separately.
@swapnilpaliwal-sd swapnilpaliwal-sd changed the title reshape: parser/ (subtree) + graph/ + bin/axiom-graph + skills/ — one repository for the whole pipeline reshape: parser/ (subtree) + graph/ + bin/axiom-graph. One repository for the whole pipeline Sep 14, 2026
… every test script finds the repository root by its marker instead of counting ../ levels

bin/axiomcode replaces the single-purpose command: `parser <src> <ir>`, `engine --language L
--client-ir <ir> --out <dir>`, `all --language L --src <dir> --out <dir>`, and `test
[java|typescript|python|parser|all]` — no npx, no positional order for the parser to remember.

The suites and tools derived the repository root with `cd $HERE/../../..`, a count that
moving a script silently invalidates — the move to graph/test/ broke seven of them, five of
which failed quietly (a node_modules/typescript lookup, portable-stat.sh sourcing, the corpus
decl path, the oracle sibling path). Every one now walks up to the marker (package.json next
to graph/), so depth no longer matters; the Python tools do the same with pathlib.
…he graph records what it was built from

`parser` and `all` take --version (default: the source tree's git commit, or v1.0.0 when it
is not a checkout) and --exclude-tests (default: included). The version is the parser's
serviceVersionLink, so it stamps every IR row; `all` also passes it to the engine as
run.source_version, with run.source_dir, through a --meta pass-through added to
run-souffle.sh — a graph.sqlite always says which commit it describes.
… detects them from the IR and solves each into <out>/<lang>/graph.sqlite; --language restricts to one
The parser emits every language's tables into one flat directory; the command sorts them
into <ir>/java/, <ir>/typescript/, <ir>/python/ (javascript and csharp as they land) — the
unprefixed Java and config tables to java/, a language with no rows gets no folder — and
`all` runs each engine on its own folder. Movable into the parser once its repository is
archived.
…y takes source trees (parsed for you) or IR roots

No subcommand and no language: the parser emits every language it finds and each is
solved. --library entries that are source trees are parsed once under
<out>/.intermediate/lib/; IR roots (a JDK IR, a per-language IR, a flat IR) are recognised
by their marker tables, never by folder names — a source tree may have a java/ folder.
swapnilpaliwal-sd and others added 3 commits September 13, 2026 20:09
Rename detection carried every edited file from src/ to graph/ on its own. One conflict,
in the bundle test: reshape renamed the bundle's CSV output directory graph/ -> csv/
(the repository now HAS a top-level graph/), and #474 added dispatch_candidates to the
same core-table loop. Both, not either.

The new files — 37 .envelope goldens and test/tools/envelope_report.py — needed placing
by hand: git can infer where an EDITED file went when its directory was renamed, but not
where a file created on the other branch belongs.

Verified on the merged tree: bundle-test.sh passes (3 languages, one schema across them,
canonical queries run, schema doc current), the Python literal gate passes, all 37 goldens
are under graph/test/*/expected, and the three rule files carry method_dispatch_candidate.
…ch other's IR; --per-language output layout

Two TypeScript or two Python projects under one root were analysed concurrently into the
same output directory, and each relation was opened with a truncating write — whichever
project finished last kept its rows and the other's vanished, silently, in an order that
varied between runs. A package plus a loose tests/ folder is two projects, so any tree with
tests beside the package lost one of them. The JavaScript analyzer already unioned its roots
up front; TypeScript and Python now analyse each project into its own scratch folder and
extract.ts merges the relations per language — header once, rows in project order —
without touching the analyzers. The multi-program fixture, which documented the defect so
the harness's stage-siblings-separately workaround stayed honest, now asserts the fix;
that workaround is gated on a detected shortfall and stays dormant.

`--per-language` (index.ts flag; ExtractOptions.layout) writes outputDir/java/,
typescript/, python/, javascript/ — one folder per language that had a project, holding
only that language's tables, the config tables with java/. The default layout is unchanged.
bin/axiomcode uses the flag instead of sorting files itself.
…nse (FSL-1.1-Apache-2.0), AxiomCode Inc.; package.json license fields; README pointer
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

build Build, packaging and developer setup enhancement New feature or request platform OS / toolchain portability

Projects

None yet

1 participant