From b12a6cf7b64b334015f391351af28813b1c2fc35 Mon Sep 17 00:00:00 2001 From: swapnil <78632212+swapnilpaliwal-sd@users.noreply.github.com> Date: Sun, 13 Sep 2026 22:54:32 -0700 Subject: [PATCH 1/2] =?UTF-8?q?README:=20the=20front=20page=20a=20reader?= =?UTF-8?q?=20expects=20=E2=80=94=20tagline,=20quick=20start=20first,=20wh?= =?UTF-8?q?at=20you=20get,=20how=20it=20works,=20accuracy,=20layout,=20dev?= =?UTF-8?q?elopment?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The substance is unchanged (every measured number is kept); the order and shape follow what established projects do: what it is in one line, how to run it in three, then depth. Two open-source project names and two library names are dropped from the public text. --- README.md | 337 ++++++++++++++++++++++-------------------------------- 1 file changed, 138 insertions(+), 199 deletions(-) diff --git a/README.md b/README.md index a9c1b476..66bdf801 100644 --- a/README.md +++ b/README.md @@ -1,250 +1,189 @@ -# AxiomCode code graph - -**A knowledge graph of what code actually does, derived formally rather than guessed.** - -`axiom-code-graph` builds a *type-directed call graph*: for every call site in a codebase it resolves -which function (or set of functions) can actually run, by reasoning over the type system — receiver -types, type hierarchy, overload applicability, generics, closure and function-reference targets — -with the platform library and third-party dependencies linked in as typed signatures. - -The engine is language-independent: it solves over a relational IR, so support for a language is a -matter of emitting that IR. The first front end is JVM-based, and the validation below uses it -because the platform ships its own class-file parser — which lets ground truth be read from compiled -artifacts with no third-party analyzer in the loop. Additional front ends follow the same contract. - -The output is a queryable graph of program structure with provenance on every edge, designed to answer -questions like *"if I change this method, what breaks?"*, *"who can reach this sink?"*, *"what is the -minimum code an agent needs to read to reason about this change?"* — and to be **auditable** when it -answers them. +

AxiomCode Code Graph

+ +

+ A type-resolved call graph of your codebase — derived formally, validated against ground truth, queryable from one SQLite file. +

+ +

+ License: FSL-1.1-Apache-2.0 + Node ≥ 22.5 + Languages +

+ +

+ Quick start · + What you get · + How it works · + Accuracy · + Layout · + Development +

--- -## Why it exists - -An AI agent working on a real codebase has to decide what to read. Today that decision is made by -text: grep, fuzzy search, embeddings. Text-similarity retrieval has two failure modes and both are -expensive: +For every call site in a codebase, the engine resolves **which function — or set of functions — can actually run**, by reasoning over the type system: receiver types, hierarchy, overloads, generics, closures and function references, with libraries linked in as typed signatures. Every edge carries a **confidence tier** and every blind spot is **declared, never dropped**. The result is one `graph.sqlite` per language whose schema is documented inside it, so a person or an AI agent can answer *"if I change this, what breaks?"*, *"who can reach this?"*, *"what do I need to read?"* — and see how sure the answer is. -* **It returns too much.** On a 3,200-file project, a grep-style expansion from a changed method - returns tens of thousands of methods. That is a context window filled with code that has no causal - relationship to the change — tokens paid for noise, and a model whose attention is diluted. -* **It misses the one that matters.** The method that breaks is often the one that never mentions your - method's name — it dispatches through an interface, a lambda stored in a field, an inherited - override. Name matching cannot see those edges. Neither can embeddings. +## Quick start -A call graph fixes both, *if* you can trust it. An untrustworthy call graph is worse than none: a -missing edge is a silent wrong answer, and an over-fanned edge floods the context you were trying to -shrink. So the design goal here is not "produce a graph" — it is **produce a graph whose errors are -known, bounded, and labelled.** +```bash +git clone https://github.com/AxiomCodeAI/axiom-code-graph.git && cd axiom-code-graph +npm install # builds the parser and the engine -## What it is designed to do +bin/axiomcode ./out # source tree in → ./out//graph.sqlite per language found +``` -| capability | what it means in practice | -|---|---| -| **Change impact** | From the methods a commit touched, return the transitive callers — the blast radius. Depth-bounded, so you can ask for "definitely affected" or "possibly affected" separately. | -| **Context selection for agents** | Turn "here is the repo" into "here are the 95 methods causally connected to this change". Measured: 95 of 43,793 methods at depth 1 — **99.78% of the codebase eliminated** while still containing **100%** of the true direct callers. | -| **Reachability & security** | Traverse from entry points or toward sinks (path traversal, deserialization, SSRF, …) to decide whether a vulnerable API is actually reachable from untrusted input, instead of flagging every import of a library. | -| **Confidence-tiered answers** | Every edge is labelled: `known_edge` (one resolved target), `multi_inferred` (a sound dispatch set), `boundary_lib` (client → library), `ambiguous_unknown` (an honestly declared blind spot). A consumer picks its own risk tolerance instead of trusting a flat list. | +Then ask questions: -## What "formal" means here — and what it does not +```sql +-- who calls this method, from where, and how sure are we? +sqlite3 out/java/graph.sqlite ".parameter set :qualified_name 'app.Widget.render'" \ + "$(sqlite3 out/java/graph.sqlite "SELECT sql FROM schema_queries WHERE name='callers_of'")" +``` -Being precise about this matters more than the marketing value of the word. +``` +caller file_path start_line tier kind +InheritanceOverride.main src/InheritanceOverride.java 40 known_edge method +``` -**It is formal in these senses:** +Requirements: **Node ≥ 22.5** and a POSIX shell (Git Bash on Windows). Until the prebuilt engine packages are published ([#478](https://github.com/AxiomCodeAI/axiom-code-graph/pull/478)), the first solve per language also needs [Soufflé](https://souffle-lang.github.io) 2.5 and a C++ compiler to compile the engine once; after that, `npm install` fetches it prebuilt and neither is needed. -* The graph is the **least fixpoint of a declarative rule set** — ~44 Datalog (Soufflé) rule files over - a relational IR of the source. There is no model, no heuristic scoring, no sampling. The same input - yields the same graph, and every edge is traceable to the rules and facts that derived it. -* **Over-approximation is explicit, not accidental.** Where dispatch is genuinely ambiguous the engine - emits the *sound set* of possible targets (`multi_inferred`) rather than a guess, so downstream - reasoning can be sound too. -* **Unknowns are declared.** A call site the engine cannot resolve is emitted as `ambiguous_unknown` — - never dropped. A CI guard asserts that the count of *silently missing* call sites is zero, so - coverage gaps cannot hide. +
+All options -**It is not** a proof of program correctness, a termination/safety verifier, or a model checker. It is -formal *reasoning about program structure*, whose conclusions are then **empirically validated against -an independent implementation** — which is the part most call-graph tools skip. +``` +bin/axiomcode [options] + --library [,…] the platform library and real dependencies — source trees (parsed for you) or IR roots. + Libraries are the type oracle: without them, calls into dependencies are declared unknown. + --language L restrict to one language (java | typescript | python) + --version V stamp the IR with a version (default: the source's git commit, else v1.0.0) + --exclude-tests leave test code out (default: included) + --debug also write csv/*.csv and keep raw/ next to each graph.sqlite + --dispatch-cap N|off fan-width cap on virtual dispatch (default 20; off for unbounded reachability) + +bin/axiomcode parser # the stages separately +bin/axiomcode engine --language L --client-ir / --out +bin/axiomcode test [java|typescript|python|parser|all] +``` +
-## How it is validated +## What you get -Claims about a call graph mean nothing without ground truth built by a *different* toolchain. Every -number below comes from one: +One SQLite database per language, **same schema for every language**, with names, files and lines already joined — no IR, no source, no rule files needed to read it: -| oracle | role | +| table | holds | |---|---| -| the language platform's **own** compiled-artifact parser (for the JVM front end: `java.lang.classfile`, JEP 484) | reference edges read from **compiled artifacts**, not from source. No third-party analyzer. | -| the same parser over application classes, dependency archives and the platform image | independent type hierarchy — used to re-point each edge to its declaring type and to build the dispatch envelope | -| runtime instrumentation | **actually-executed** edges; an executed edge the graph lacks is the strongest possible bug signal | - -Scored against two reference sets, because a call graph has two kinds of truth: **must-have** (the -compiled artifact's declared targets) and the **sound envelope** (plus class-hierarchy dispatch). An edge inside the -envelope is a real possibility; only an edge outside it is a defect. Normalization is symmetric and -documented — anonymous types keyed by supertype, closure bodies folded to their enclosing function, -and bridges / `access$N` / enum `values` / string-concat lowering / autoboxing / invokedynamic / -enhanced-for desugaring / synthesized constructors excluded on **both** sides. +| `call_edges` | the graph: one row per (call site, possible target) with `tier`, provenance and call kind | +| `methods` · `types` · `call_sites` | every callable, type and call site with qualified name, file and line | +| `type_ancestors` · `overrides` | hierarchy and virtual-dispatch pairs | +| `entry_points` · `entry_reachable` | what the runtime invokes, and what it reaches | +| `unresolved_sites` | the declared blind spots, attributed to the method that contains them | +| `schema_guide` · `schema_queries` · `schema_vocab` · `schema_notes` | **the documentation, inside the database**: how to use it, tested canonical queries, every enum value per language with its meaning, per-language caveats | -### Results +Every edge has a tier, so a consumer picks its own risk tolerance: -Hand-crafted constructs with known edges — inheritance and virtual dispatch, anonymous types and -single-method interfaces, overload disambiguation with implicit conversion, receiverless calls: - -> **precision 1.000 · recall 1.000** on application-internal edges, 0 silently dropped call sites - -Change impact on a large open-source database engine (JVM front end — the first one implemented; -other front ends are scored with the same harness): **five consecutive real commits**, 69 changed -methods, seeded at each commit's parent revision; universe 13,398 files / 181,355 methods. Two -bounds, because a call graph has two kinds of truth — **certain** callers (statically resolved -declared targets; a miss is undeniable) and **possible** callers (plus hierarchy dispatch; a result -outside it is a demonstrable false positive). - -| depth | approach | precision | recall vs certain | recall vs possible | F1 | MCC | -|---|---|---|---|---|---|---| -| d1 | **axiom-code-graph** | **0.980** | **0.912** | **0.729** | **0.836** | **0.845** | -| d1 | tree-sitter name matching | 0.397 | 0.922 | 0.707 | 0.508 | 0.529 | -| d1 | conservative resolver | 0.772 | 0.863 | 0.662 | 0.713 | 0.714 | -| d2 | **axiom-code-graph** | **0.989** | **0.911** | **0.438** | **0.607** | **0.658** | -| d2 | conservative resolver | 0.646 | 0.693 | 0.308 | 0.418 | 0.446 | -| d2 | tree-sitter name matching | 0.390 | 0.799 | 0.356 | 0.372 | 0.371 | +| tier | meaning | +|---|---| +| `known_edge` | exactly one resolved target | +| `multi_inferred` | a *sound set* of possible targets (virtual dispatch over instantiated subtypes) | +| `boundary_lib` | the target is in a library — named, not expanded | +| `ambiguous_unknown` | the engine could not resolve the site; kept as a row with a NULL target | -Two properties matter more than the averages: +Full schema: [`graph/bundle/SCHEMA.md`](graph/bundle/SCHEMA.md). -- **It never over-estimates.** Seeded one changed method at a time (59 seeds with callers), at any - depth, on any seed. Exact affected-count on 45/59 seeds at d1, 30/59 at d3 — against 33 and 5 for - name matching. When it is wrong it is conservative. -- **Its errors are few enough to inspect.** At d2: **2 false positives** against 68 and 224. +## Supported languages -Context per change at d1: **9.8 files / 56 K tokens** — 99.3% less than reading the repository, and -1/26 the token spend of an unindexed agent answering the same question. +| language | status | ground truth used for validation | +|---|---|---| +| Java | stable | the JDK's own class-file parser over compiled artifacts; runtime tracing | +| TypeScript | stable | the TypeScript compiler's own resolution | +| Python | stable | CPython bytecode and `sys.settrace` | +| JavaScript | in progress | — | +| C# | planned | — | -Raw *accuracy* is meaningless at this class imbalance — returning **nothing** scores 0.998, since -~99.8% of a codebase is unaffected by any given change. Read recall against each bound instead. +A multi-language repository is one command: the parser emits every language it finds, and each is solved into its own `graph.sqlite`. -**The gap:** at d1, 36 of 133 possible callers are missed — small delegating types that dispatch -through a dependency-declared interface, statics qualified by a type name, and callbacks whose -receiver is a lambda parameter. Restricted to files of 1,000 lines or more, d1 recall is 1.000. +## How it works -## Run +``` + source tree ──▶ parser ──▶ relational IR ──▶ engine (Datalog, per language) ──▶ graph.sqlite + parser/ csv tables graph//engine/*.dl + schema inside + compiled once per platform, shipped prebuilt +``` -```bash -npm install && npm run build # builds the parser and the engine (Node ≥ 22.5) +1. **Parse.** The parser extracts a relational IR — types, methods, expressions, call sites, imports — for every language present, in one pass. +2. **Solve.** Each language's rule set (~40 Soufflé Datalog files) computes the **least fixpoint** over that IR: type resolution, hierarchy, generics, overload applicability, virtual dispatch, closure and function-value flow. No model, no scoring, no sampling: the same input yields the same graph, and every edge is traceable to the rules and facts that derived it. +3. **Bundle.** The raw relations are joined to the IR and written as `graph.sqlite`, with the schema, vocabularies and canonical queries as tables. -bin/axiomcode all --src --out # source → //graph.sqlite, per language found -# --language java | typescript | python restrict to one (the parser emits every language it finds) -# --version V stamp the IR (default: the source's git commit, else v1.0.0) -# --exclude-tests leave test code out (default: included) -# --library [,...] the platform library and real dependencies, when you have their IR -# --debug also write csv/*.csv and keep raw/ +The rules compile to one self-contained executable per language and platform. CI builds them (Linux x64/arm64, macOS arm64, Windows x64) and publishes them on npm as `@axiomcode/engine--`, which `npm install` selects by platform ([#478](https://github.com/AxiomCodeAI/axiom-code-graph/pull/478)). With Soufflé installed, the engine compiles locally instead; a checkout whose rules differ from the published engine never runs a stale binary. -bin/axiomcode parser # the two stages separately: -bin/axiomcode engine --language java --client-ir /java --out # IR is written per language, // -bin/axiomcode test [java|typescript|python|parser|all] -``` +### What "formal" means here -`bin/axiomcode` is the whole pipeline as subcommands: the parser (`parser/`) extracts a relational IR -from the source — every language it finds, in one pass — the engine (`graph/`) solves each language -separately, and `//graph.sqlite` is the result (graphs are per language; a Java→TypeScript -call is not an edge in either). **No -Soufflé and no C++ compiler**: the rules compile to one self-contained executable, CI builds it for -Linux (x86_64, arm64), macOS (arm64) and Windows on every merge to `main` and commits it under -`binaries///`, so a checkout carries the engine for every platform. With `souffle` -installed the engine compiles locally instead. +- The graph is the least fixpoint of a declarative rule set over a relational IR — deterministic and auditable. +- **Over-approximation is explicit.** Genuinely ambiguous dispatch yields the sound set (`multi_inferred`), not a guess. +- **Unknowns are declared.** An unresolvable call site is emitted as `ambiguous_unknown`; a CI guard asserts that the count of silently missing call sites is zero. -**Outputs — the same in every language** ([`graph/bundle/SCHEMA.md`](graph/bundle/SCHEMA.md)) +It is *not* a proof of program correctness or a model checker. It is formal reasoning about program structure whose conclusions are then **validated against an independent implementation** — the part most call-graph tools skip. -``` -/ - graph.sqlite the contract: core tables, the language's ext_* relations, and the schema as tables -and only with --debug: - csv/.csv the core tables as headered, tab-delimited text - raw/ the per-language Soufflé relations, verbatim — engine-internal, not a contract -``` +## Accuracy -An agent that opens `graph.sqlite` needs nothing else: `schema_guide` says how to use it in -reading order, `schema_queries` holds tested SQL for the common questions, `schema_vocab` and -`schema_notes` carry the per-language meaning of every value and every caveat. +Every number below is scored against ground truth built by a **different toolchain** — the language platform's own compiled-artifact parser and runtime instrumentation, never a third-party analyzer. -The core tables are `methods`, `types`, `call_sites`, `call_edges`, `type_ancestors`, `overrides`, -`entry_points`, `entry_reachable`, `unresolved_sites`, `type_instantiated` and `run` — with names, -files and lines already joined in, so a consumer needs nothing but the one file: +**Hand-crafted constructs** (inheritance and virtual dispatch, anonymous types and single-method interfaces, overload disambiguation with implicit conversion, receiverless calls): **precision 1.000 · recall 1.000** on application-internal edges, 0 silently dropped call sites. -```sql --- who calls Widget.render, and where? -SELECT caller.qualified_name, s.file_path, s.start_line, e.tier -FROM call_edges e JOIN methods callee ON callee.id = e.callee_method_id - JOIN methods caller ON caller.id = e.caller_id - JOIN call_sites s ON s.id = e.call_site_id -WHERE callee.qualified_name = 'app.Widget.render'; - --- what can `tier` be in THIS bundle's language, and what does each value mean? -SELECT value, meaning FROM schema_vocab -WHERE table_name = 'call_edges' AND column_name = 'tier' - AND language = (SELECT value FROM run WHERE key = 'language'); -``` +**Change impact on a large open-source database engine** (JVM front end): five consecutive real commits, 69 changed methods, universe of 13,398 files / 181,355 methods. Two bounds, because a call graph has two kinds of truth — *certain* callers (statically resolved declared targets) and *possible* callers (plus hierarchy dispatch): -Where the front ends differ — which values a column can hold, what an id may point at, which -tables a language leaves empty — the difference is recorded in `schema_vocab` and `schema_notes` -inside the database, and in `SCHEMA.md`, generated from the same source (`npm run schema-doc`). -`graph.sqlite` needs Node ≥ 22.5 (`node:sqlite`); on an older Node the CSVs are still written and -the omission is reported. +| depth | approach | precision | recall vs certain | recall vs possible | F1 | MCC | +|---|---|---|---|---|---|---| +| d1 | **AxiomCode** | **0.980** | **0.912** | **0.729** | **0.836** | **0.845** | +| d1 | tree-sitter name matching | 0.397 | 0.922 | 0.707 | 0.508 | 0.529 | +| d1 | conservative resolver | 0.772 | 0.863 | 0.662 | 0.713 | 0.714 | +| d2 | **AxiomCode** | **0.989** | **0.911** | **0.438** | **0.607** | **0.658** | +| d2 | conservative resolver | 0.646 | 0.693 | 0.308 | 0.418 | 0.446 | +| d2 | tree-sitter name matching | 0.390 | 0.799 | 0.356 | 0.372 | 0.371 | -**Knobs** +- **It never over-estimates.** Seeded one changed method at a time (59 seeds), exact affected-count on 45/59 seeds at d1 and 30/59 at d3 — against 33 and 5 for name matching. When it is wrong it is conservative. +- **Its errors are few enough to inspect.** At d2: **2 false positives**, against 68 and 224. +- **Context per change at d1: 9.8 files / 56 K tokens** — 99.3 % less than reading the repository, and 1/26 the token spend of an unindexed agent answering the same question. At d1, 95 of 43,793 methods: **99.78 % of the codebase eliminated while retaining 100 % of true direct callers.** -* `--library a,b,c` — the platform library plus every real dependency. Libraries are a *type oracle* even - when you only want first-party edges; omitting them is the single largest source of unresolved - receivers. -* `--dispatch-cap N|off` (default 20) — how many implementations one virtual call site may fan to. - **Turn it off for unbounded reachability** (taint / sink traversal), where a target behind a wide - dispatch would otherwise be dropped. -* `--jdk-depth` (platform-hop cap), `--lib-depth`, `--taint`. +Raw accuracy is meaningless at this class imbalance (returning nothing scores 0.998); read recall against each bound. The residual d1 gap — 36 of 133 possible callers — is small delegating types dispatching through a dependency-declared interface, statics qualified by a type name, and callbacks whose receiver is a lambda parameter; restricted to files of 1,000+ lines, d1 recall is 1.000. -## Layout +## Repository layout ``` -bin/axiomcode parser | engine | all | test — the pipeline as subcommands -parser/ the IR extractor (its own package; merged in with history) -graph/ the engine - /engine/projections/ IR → typed relations - /engine/containment/ ownership, type nesting - /engine/resolution/ type resolution, hierarchy, generics, virtual dispatch - /engine/expression-resolution/ call sites, callee resolution, overloads, lambdas - /engine/call-edge-generation/ call classes, chain edges, lambda dispatch - /souffle/ relation declarations + export manifest - /templates/ staging maps - pipeline/run-souffle.sh fact staging, engine resolution (committed / compiled / fetched), stage↔solve loop, then the bundle stage - bundle/ the output contract: schema as data (SCHEMA.md), per-language adapters, writers - test// the engine's regression suites, torture harnesses, oracles (graph/test/tools: platform preflights) -binaries/// CI-built engines, committed on merge (ENGINE_ID = the rules they were built from) +bin/axiomcode the command: parser | engine | all | test +parser/ the IR extractor (Java, TypeScript, Python, JavaScript) +graph/ + /engine/ the rules: projections → containment → resolution → expression resolution → call edges + /souffle/ relation declarations and the export manifest + pipeline/run-souffle.sh fact staging, engine resolution (npm package / local compile), stage↔solve loop, bundle stage + bundle/ the output contract: schema as data, SCHEMA.md, per-language adapters, writers + test// regression suites, torture harnesses, oracles +packaging/ the @axiomcode/engine-- package template CI publishes +.github/workflows/ engine builds and the npm publish ``` -## Tests +## Development ```bash -bin/axiomcode test java # the engine's Java suite; --oracle also scores against javac/javap ground truth -bin/axiomcode test typescript # --oracle scores against the TypeScript compiler -bin/axiomcode test python # --oracle scores against CPython bytecode and tracing -bin/axiomcode test parser # the parser's own suites (bin/axiomcode test = everything) +npm install && npm run build # parser + engine +bin/axiomcode test java # regression suite; --oracle scores against javac/javap ground truth +bin/axiomcode test typescript # --oracle scores against the TypeScript compiler +bin/axiomcode test python # --oracle scores against CPython bytecode and tracing +bin/axiomcode test parser # the parser's own suites +bin/axiomcode test # everything ``` -Each suite parses its fixture cases with the parser in this repository (`AXIOM_PARSER` overrides), -solves them, guards that no call site was dropped, and diffs the normalised edges against a -golden. `--keep` retains the per-case work directories (`graph/test//.work//out/graph.sqlite` -is a real bundle to poke at); `--bless` regenerates goldens — review the diff. The torture -harnesses (`graph/test//torture/`) and the TypeScript corpus runner score real projects. +Each suite parses its fixture cases with the parser in this repository, solves them, guards that no call site was dropped, and diffs the normalised edges against a golden. `--keep` retains per-case work directories (`graph/test//.work//out/graph.sqlite` is a real bundle to poke at); `--bless` regenerates goldens — review the diff. Torture harnesses under `graph/test//torture/` score real projects against their oracles. + +Editing rules requires [Soufflé](https://souffle-lang.github.io) 2.5 locally (`brew install souffle`; the pinned version is in `graph/pipeline/engine.conf`); the engine recompiles on the first solve after a rule change. Publishing engines: *Actions → publish-npm → Run workflow* (dry run by default) or push a `v*` tag. ## Known limits -Stated because a graph you can't trust the boundaries of isn't useful: - -* **Unstaged dependencies dominate the residual recall gap** — 79.8% of unresolved receivers point at - libraries that weren't linked in (Guava 44%, slf4j 12%, …). Pass them via `--library`. -* **Generic substitution through library-written type arguments** is incomplete, so a lambda parameter - typed only through a library generic chain (`map.values().forEach(x -> x.m())`) stays unresolved. -* **Nested types are flattened** by the IR (`pkg.Outer.Inner` → `pkg.Inner`); on cassandra 623 types - collide, which costs precision on `multi_inferred`. -* **Function values in parameters or collections** are not tracked (fields and locals are). -* **Reflection** is out of scope by construction, and is reported as `ambiguous_unknown` rather than - silently omitted. +- **Unlinked dependencies dominate the residual recall gap** — most unresolved receivers point at libraries that were not passed via `--library`. +- Generic substitution through library-written type arguments is incomplete, so a lambda parameter typed only through a library generic chain stays unresolved. +- Nested types are flattened by the IR (`pkg.Outer.Inner` → `pkg.Inner`), which costs precision on `multi_inferred` where names collide. +- Function values in parameters or collections are not tracked (fields and locals are). +- Reflection is out of scope by construction and is reported as `ambiguous_unknown`, never silently omitted. ## License From ca71c8431556caae0d975604d70326950ecf79c1 Mon Sep 17 00:00:00 2001 From: swapnil <78632212+swapnilpaliwal-sd@users.noreply.github.com> Date: Sun, 13 Sep 2026 23:15:46 -0700 Subject: [PATCH 2/2] README: what it does and why coding agents need it, in plain language; the three trust properties; languages with their maturity --- README.md | 40 ++++++++++++++++++++++++++-------------- 1 file changed, 26 insertions(+), 14 deletions(-) diff --git a/README.md b/README.md index 66bdf801..92f6e051 100644 --- a/README.md +++ b/README.md @@ -11,6 +11,8 @@

+ What it does · + Languages · Quick start · What you get · How it works · @@ -21,7 +23,29 @@ --- -For every call site in a codebase, the engine resolves **which function — or set of functions — can actually run**, by reasoning over the type system: receiver types, hierarchy, overloads, generics, closures and function references, with libraries linked in as typed signatures. Every edge carries a **confidence tier** and every blind spot is **declared, never dropped**. The result is one `graph.sqlite` per language whose schema is documented inside it, so a person or an AI agent can answer *"if I change this, what breaks?"*, *"who can reach this?"*, *"what do I need to read?"* — and see how sure the answer is. +## What it does + +Point it at a repository and it builds a **call graph**: a map of which function calls which, across the whole codebase. Unlike a text search, it knows about types — so when `shape.area()` could run `Circle.area` or `Square.area`, the graph says *both*, and when a call goes through an interface, an inherited override, or a function stored in a variable, the graph still finds the target. The result is one `graph.sqlite` file per language, with its own documentation inside, that you query with plain SQL. + +**Why coding agents need it.** An AI agent working on a real codebase has to decide what to read. Text search returns too much (thousands of unrelated methods that happen to share a name) and misses what matters (the caller that never mentions the name because it dispatches through an interface). Both cost tokens and both cause wrong answers. With the graph, an agent asks *"what does this change affect?"* and gets the causally connected methods — measured on a large project: **95 of 43,793 methods, with 100 % of the true direct callers included** — instead of reading the repository. + +**Why you can trust it.** Three properties, each checked rather than promised: + +- **Exact and repeatable.** The graph is computed by a fixed set of logical rules, not a model or a heuristic score. The same code produces the *identical* graph every time, and every edge can be traced back to the rule and the facts that produced it. +- **Validated against the compiler.** Results are scored against ground truth from the language's own toolchain — the JDK's class-file parser over compiled bytecode, the TypeScript compiler, CPython's bytecode and tracing — never against another third-party analyzer. On hand-crafted constructs: precision 1.000, recall 1.000. +- **Honest about what it doesn't know.** Every edge carries a confidence tier, and a call the engine cannot resolve is kept as a row that says so — never silently dropped. A test asserts the number of silently missing call sites is zero. + +## Supported languages + +| language | maturity | what "maturity" means here | +|---|---|---| +| **Java** | stable | first front end; validated on hand-crafted constructs (P/R 1.000) and on five real commits of a large open-source project against bytecode ground truth; Spring/DI configuration wiring resolved | +| **TypeScript** | stable | 53 regression cases plus real projects scored against the TypeScript compiler's own resolution; structural typing, overload sets, module graph, `.d.ts` libraries | +| **Python** | stable | MRO, decorators, protocols, dynamic-attribute detection; scored against CPython bytecode and `sys.settrace` on a 600-site torture suite | +| **JavaScript** | in progress | parser complete; engine under review | +| **C#** | planned | — | + +A repository with several languages is one command: the parser emits every language it finds, and each is solved into its own `graph.sqlite`. Graphs are per language — a Java→TypeScript call is not an edge in either. ## Quick start @@ -90,18 +114,6 @@ Every edge has a tier, so a consumer picks its own risk tolerance: Full schema: [`graph/bundle/SCHEMA.md`](graph/bundle/SCHEMA.md). -## Supported languages - -| language | status | ground truth used for validation | -|---|---|---| -| Java | stable | the JDK's own class-file parser over compiled artifacts; runtime tracing | -| TypeScript | stable | the TypeScript compiler's own resolution | -| Python | stable | CPython bytecode and `sys.settrace` | -| JavaScript | in progress | — | -| C# | planned | — | - -A multi-language repository is one command: the parser emits every language it finds, and each is solved into its own `graph.sqlite`. - ## How it works ``` @@ -111,7 +123,7 @@ A multi-language repository is one command: the parser emits every language it f ``` 1. **Parse.** The parser extracts a relational IR — types, methods, expressions, call sites, imports — for every language present, in one pass. -2. **Solve.** Each language's rule set (~40 Soufflé Datalog files) computes the **least fixpoint** over that IR: type resolution, hierarchy, generics, overload applicability, virtual dispatch, closure and function-value flow. No model, no scoring, no sampling: the same input yields the same graph, and every edge is traceable to the rules and facts that derived it. +2. **Solve.** Each language's rule set (~40 Soufflé Datalog files) applies the rules until nothing new can be derived: type resolution, hierarchy, generics, overload applicability, virtual dispatch, closure and function-value flow. No model, no scoring, no sampling — the same input yields the same graph. 3. **Bundle.** The raw relations are joined to the IR and written as `graph.sqlite`, with the schema, vocabularies and canonical queries as tables. The rules compile to one self-contained executable per language and platform. CI builds them (Linux x64/arm64, macOS arm64, Windows x64) and publishes them on npm as `@axiomcode/engine--`, which `npm install` selects by platform ([#478](https://github.com/AxiomCodeAI/axiom-code-graph/pull/478)). With Soufflé installed, the engine compiles locally instead; a checkout whose rules differ from the published engine never runs a stale binary.