Skip to content

feat(doc-links): C# doc-comment references become MENTIONS edges - #2550

Open
DeusData wants to merge 19 commits into
fix/c-typedefs-macros-enumeratorsfrom
feat/doc-links-core
Open

DeusData wants to merge 19 commits into
fix/c-typedefs-macros-enumeratorsfrom
feat/doc-links-core

Conversation

@DeusData

@DeusData DeusData commented Oct 5, 2026

Copy link
Copy Markdown
Owner

Summary

References written in doc comments become graph edges. A C# <see cref="..."/>, <seealso>, <exception> or <inheritdoc cref> that names a definition of the repository is stored as a MENTIONS edge from the documented definition to that target. A reference that cannot be bound with certainty is never guessed: it becomes a row with a reason.

This PR brings the layer itself (extraction of the references, the resolving step, storage, the status report, incremental upkeep) and its first language, C#. The other languages follow as one PR each on this base.

Stacked on #2543 (C typedefs, macros, enumerators; index version 4). Upgrade: the semantic index version goes to 5, so an existing index is rebuilt in full once.

What you get

Edges. MENTIONS, source = the definition whose doc comment holds the reference (the file's File node for a file-level comment), target = the referenced definition. Properties:

{"via":"doc_comment","syntax":"see","tier":"exact","line":42,"count":2}
  • syntax: the markup family (see, seealso, exception, inheritdoc).
  • tier: exact when the reference is a qualified name, or reaches its target through a using, an alias or a path; unique when it is the first scope level with a hit.
  • line: the first line the reference is written on; count: how often this source mentions this target.

Unresolved references. Table doc_link_unresolved(project, rel_path, line, syntax, raw, reason). Reasons:

Reason Meaning
missing nothing of that name is in scope in the repository
ambiguous more than one declaration fits and the markup does not say which
external the name belongs to something outside the repository (for C#: the System namespaces, or a member a base type outside the repository would have to supply)
test_only_target the only declaration is test code that product code cannot see
graph_gap the declaration exists in the source but has no node (for example a member inside #if)
unparseable the reference text is not a valid reference
below_bar_tier the markup family did not pass the precision audit and is switched off
not_indexed reserved; C# never emits it

Status. index_status gets a doc_links block: {"mentions": N, "unresolved": {"<reason>": N, ...}, "status": "ok" | "error"}, a hint when the index predates the layer or the layer failed, and with diagnostics=full up to 50 sample rows.

No new exit codes, no new tool. MENTIONS is queried like any edge (query_graph, trace_path).

How a C# reference is bound

The rule is the compiler's own rule for cref, read from Roslyn's binder (Binder_Lookup.cs, Binder_Crefs.cs) and implemented from the rule, not from its code:

  • The first scope with the name decides; a qualified name is tried as written.
  • A cref never binds an inherited member. Such a reference is external when a base type lies outside the repository, missing otherwise.
  • The name is found first, a parameter list is matched afterwards; generic methods are ignored beside a non-generic one; an alias with a parameter list binds nothing.
  • System is the one namespace root treated as outside the repository by definition.

Several projects in one repository. A type name may be declared in more than one assembly. The layer does not choose between them:

  • An assembly is a project file's name; a source file belongs to the nearest project file at or above it. A project file whose SDK compiles nothing (Microsoft.Build.NoTargets, Microsoft.Build.Traversal) owns no source file.
  • A file with no project above it belongs to a shared tree (the largest directory with no project file in or below it).
  • Two declarations of one name in different assemblies make the reference ambiguous. There is no preference for a nearer directory or for a ref directory.
  • One exception with a proof behind it: an assembly that holds only a stub of the type, plus exactly one implementation in a shared tree and no other implementation anywhere, is one type.
  • MSBuild global usings (<Using>, ImplicitUsings, Directory.Build.props / .targets, imports, conditions) are evaluated from the project files' content. The layer opens no file: project files reach it as data the indexer has already read, and a path written inside a project file is only ever a key into that data.

Measured effect

Parent commit ("before") against this change, C# only.

Corpus References Edges Unresolved rows of them ambiguous
microsoft/semantic-kernel ca40aa72 7,005 3,420 2,456 2 (0.03 %)
dotnet/runtime, consumer slice (6,612 files, 61 % of its references) 30,868 14,522 5,090 3,670 (11.9 %)
dotnet/runtime, complete 50,185 22,458 9,672 7,722 (15.4 %)
  • The complete run's rows: ambiguous 7,722, graph_gap 1,158, external 626, missing 157, unparseable 7, test_only_target 2. References that are neither an edge nor a row are the ones the language resolves inside the comment itself (parameters, type parameters) and URLs.
  • dotnet/runtime is the hard case and it shows. It declares the core types more than once (the shared sources of the core library beside a second, minimal core library), so 3,373 of the slice's 3,670 ambiguous references are names like ArgumentNullException that two shared trees declare. The previous state bound all of these, to whichever declaration came first. Deciding them needs the project reference graph and type accessibility across assemblies, which is a follow-up.
  • An ordinary multi-project repository (semantic-kernel): 2 ambiguous references in 7,005.
  • Determinism: one worker and several workers give the same edges and rows (small corpus, consumer slice). An incremental run equals a full run in 11 of 11 edit steps.
  • Cost of the layer: 46 ms CPU on semantic-kernel, 185 ms on the consumer slice; on the complete dotnet/runtime (1.30 M nodes, 149 s wall for the whole index) the layer's index build takes 1.7 s and resolution 0.34 s CPU.

Incremental

MENTIONS edges and unresolved rows belong to their source file, like CALLS: a re-extracted file rewrites its own. What the edge set cannot show (a file whose references resolve differently because ANOTHER file changed) is decided per language:

  • A removed name: the files whose rows mention it are re-resolved with the changed file.
  • A change that can re-route references anywhere (for C#: a project file, a global using, a using in a file whose type has base types, a deleted file) is not repaired file by file: the run falls back to a full rebuild, as it already does for an added definition name.
  • How often that happens on real commit history was not measured; a planner that knows which files mention a name (follow-up) removes the fallback for added names.

Precision audit

Every markup family ships as edges only if it passes a held-out audit that was registered before any output existed on the audit repositories: at least 60 sampled links per family (or all of them), judged against the source by readers who do not know the resolver's rule or tier; a family passes with a Wilson 95 % lower bound of at least 0.90 and a point estimate of at least 0.95. A family below the bar is switched off (below_bar_tier rows).

Family Links judged Correct Point Lower bound Ships
see 150 (of 18,099) 150 1.000 0.975 yes
seealso 40 (census) 40 1.000 0.912 yes
exception 35 (census) 35 1.000 0.901 yes
inheritdoc 72 (census) 72 1.000 0.949 yes

Held-out repositories: dotnet/aspnetcore, dotnet/efcore, dotnet/orleans at pinned commits. Round 1: 297/297 judged correct (eight readers, one per sheet); round 2: 31/31 agreed (a fresh reader on a seeded tenth); round 3: nothing to settle. The audit ran on the layer as of its last change before the fixes listed under Robustness. Re-run with the final code (this branch, on #2543): aspnetcore and orleans give byte-identical links (10,825 and 1,864); efcore gives the same 5,560 links with four documented methods now carrying their class segment in the source name (same reference, same target: a base-graph change from main) plus one new link from a property that is now a node, judged correct by a fresh blind reader. The unresolved rows of all three are identical.

Robustness

A repository under index is untrusted input. The scanner and the resolver are bounded against it, each bound with a test that fails without it:

  • nesting depth of the C# scanner (64 levels, deeper is not placed), no recursion driven by input;
  • no cost that grows faster than the input: tests hold the work counted by test seams against the input size (twice the input, about twice the work), without a clock;
  • no silent limit: where a bound stops a lookup (more than 256 declarations of one name in one project) the reference becomes an ambiguous row and is counted;
  • a damaged or foreign stored scope fails closed.

Two independent security reviews and a re-check of the second one's fixes ran on this code (read-only reviewers). Of the second review's 18 findings, 13 hold as fixed and 4 hold in part; the re-check's new findings of the same classes are fixed here: a unit's using directives were asked again by every file (now remembered per unit and query; edges and rows identical with and without the memory); a refused scope's type names now count as graph gaps everywhere (the stored marker carries them, so an incremental run equals a full one); a doc-line map that cannot grow and a .props/.targets file malformed before its root no longer leave the status ok or close a scope; an empty stored region name is refused. A third suspected silent failure path (a file holding only directives) cannot occur, since such a file always carries its module definition; a test pins that. Two remain open by design choice and are follow-ups: a shared MSBuild file that reads a per-project value is walked again for every project, and distinct properties each keep an expanded copy of a long value; both cost time or memory on unusual project trees, neither changes edges.

Tests

119 tests in three suites: doc_mentions (72), doc_mentions_cost (9 work-counter tests: twice the input, at most about twice the work, counted by test seams, never timed) and doc_mentions_msbuild (38). Each finding of the reviews and of the author's own audit has a test that fails when its fix is reverted (the reverts were run). The first Linux arm64 run showed three things macOS did not, fixed here: the node pass now binds files in path order (its scratch-table cost had depended on the order the file system lists files in), a test thread now frees its parser caches (LeakSanitizer), and the suite is split so each part stays inside the harness's 900 s per-suite wall clock on gcc aarch64 ASan. The scratch-table test's large file is sized so the extraction budget cannot cut it on that leg, under load either (1,000 fields: 3,008 counted steps against a bound of 4,160, 40,512 with the fix reverted).

Verification

Local three-system ladder on the final, rebased trees:

  • macOS arm64, this PR's tip alone (scripts/test.sh: clean build, contract steps, every suite, binary guards): 8,579 passed, 0 failed, 10 skipped (149 suites).
  • With the documents PR on top (the stacked tree): macOS 8,609 passed, 0 failed, 10 skipped. Linux arm64 (gcc, ASan + LeakSanitizer, test-infrastructure/run.sh test): 8,446 passed, 9 skipped, and one failure, the scratch-table test's large file cut by the extraction budget under six-job load. The test was resized for that (above); the resized suite passes alone and beside five heavy suites, at 3,008 counted steps on Linux and macOS alike. Windows arm64 VM (win.sh test-par): 8,366 passed, 83 skipped. Two route tests of edge_types_probe (routes_laravel_get_survives_xlang_guard_parallel, routes_express_get_survives_xlang_guard) fail there identically on this stack's base (fix(extract): C typedefs become nodes; macros and enumerators get their own QNs #2543's head), so they do not come from this PR. Two suites stopped at the 900 s suite clock while the other legs loaded the host; both passed on an isolated re-run.
  • Lint: cppcheck, clang-format (Homebrew LLVM), no-suppress and the memory-core ratchet.

Known limits

  • #if members, operators, indexers, delegates and events have no nodes in the C# graph, so references to them are graph_gap rows.
  • Projects with the same file name in different directories count as one assembly.
  • A shared tree's files do not get any project's global usings.
  • .proj and .projitems files are not read.
  • A misspelt name under System reads external, not missing.

Follow-ups (not in this PR)

  • The project reference graph and type accessibility across assemblies, to decide names that two assemblies declare.
  • A planner for incremental runs that knows which files mention a name, so that an added name no longer forces a full rebuild.
  • The other languages, one PR each: Java and Kotlin, Go, Rust, TypeScript and JavaScript, PHP, Perl, C (Doxygen).

A reference written in a doc comment is extracted with the definition it
documents (internal/cbm/doclink.c; the C# scanner doclink_cs.c) and
resolved per file in the pipeline (src/pipeline/doc_links.c,
doc_links_cs.c, doc_links_msbuild.c). A reference that names a definition
of the repository with certainty becomes a MENTIONS edge from the
documented definition (the File node for a file-level comment) with the
properties via, syntax, tier, line and count. Any other reference becomes
a row of the new doc_link_unresolved table with its reason: missing,
ambiguous, external, test_only_target, graph_gap, unparseable or
below_bar_tier. Nothing is guessed. index_status reports the layer in a
doc_links block.

C# binds a cref by the compiler's lookup rule: the first scope with the
name decides, a cref never binds an inherited member, System is outside
the repository. Assemblies come from project files, and MSBuild global
usings (<Using>, ImplicitUsings, Directory.Build.props/.targets, imports,
conditions) are evaluated from the project files' content as data the
indexer has already read; the layer opens no file. Two declarations of a
name in different assemblies make the reference ambiguous.

Each file's scope is stored line-free with its LSP surface ("dl"), so an
incremental run re-resolves the files a change can reach and falls back
to a full run for a change that can re-route references anywhere (a
project file, a global using, a deleted file).

The scanner and the resolver treat the repository as untrusted input:
bounded nesting, no input-driven recursion, work that grows linearly
(held by test seams, never by a clock), no silent limits, and a damaged
or foreign stored scope fails closed.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…S18)

- S11: a namespace read back from a stored scope is held to the nesting the
  scanner writes (64 segments of its full name); past it the scope is
  refused like any other stored scope this code did not write.
- S12: index_status lists only the reasons the layer writes; any other
  reason text in the table is counted under one fixed key
  (unrecognized_reason) and the status is error.
- S13: the per-file scratch table of the index build is made anew when a far
  larger file grew it, so emptying it costs about the file's own size.
- S14: the portable copy of a scope keeps an empty line field empty and is
  never longer than the scope it is made from.
- S15: a reference whose brackets do not pair, or that goes on after its
  parameter list or its type arguments, is unparseable instead of being
  resolved by the part that can be read.
- S16: a project file the scan did not read (malformed, or larger than a
  project file is) opens the scope of the projects that evaluate it: what it
  holds is unknown, not absent. It is counted once.
- S18: an incremental repair matches the names a changed file no longer
  declares by the resolver's own identifier rule (bytes of multi-byte
  characters included) and without a length cut.

Tests: cs_portable_scope_bound, cs_unbalanced_brackets,
cs_stored_scope_nesting, index_status_unknown_reason,
msbuild_unread_project_opens, cs_scratch_table_work,
incremental_non_ascii_name; msbuild_gate and index_status_preview_bounds
follow the new S16 and S12 outcomes.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…ails the layer (S4, S17)

S4: the declarators of one C# field declaration share one doc comment. Its
references are taken once, from the first declarator (the Field definition
that carries them, or the first variable where there is none); the other
declarators take nothing from it, so tokens, resolutions and rows grow with
the references, not with references x declarators. A shared doc that could
not be collected whole is no doc of any declarator and is not looked up again
per declarator.

S17: memory that runs out while a file's doc-link data is extracted -- a
doc span or doc text, a reference value, a token, the scope scan -- is
recorded on the file's doc links (CBMDocLinkArray.failed), and the resolving
half fails the layer for the run (status error). A MENTIONS edge that could
not be stored fails the layer too and is not counted. Seams: the scope scan
(CBM_DOCLINK_ALLOC_SCOPE) and the edge insert.

Tests: cs_shared_doc_once, alloc_failure_status; the shared-doc tests of the
checkpoint follow the once rule.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…its file; blobs are UTF-8 (S8, S10)

S8: a C# scope blob written in this run that the resolver cannot parse no
longer fails the layer for the whole repository. What the file had put into
the shared index is taken back (the namespaces it created and their keys,
the declared and production marks it set, the quarantine keys it added), the
file is kept as an empty scope marked rejected, and each of its references
becomes a graph_gap row. The rejected files and references are counted
(stats, and the log line doc_links.cs.rejected_scopes); the layer status
stays ok. A STORED scope that does not parse still fails the run. The
scanner's output passes the reader's record checks for the shapes the review
names: a namespace name with an empty dotted segment is a namespace the scan
cannot name. Seams: spoil the scope of one file in the run, and parse a blob
with the resolver's reader.

S10: every byte a scan writes into a scope blob is well-formed UTF-8, where
it is written. In a C# file a name that is not UTF-8 is not placed (an
identifier, a namespace or type name) and a text that is not is written as
"?" (a kept text, a signature); in a project file every byte that starts no
well-formed sequence is written as U+FFFD, in element names, attribute
values and text, and a numeric reference to a surrogate code point is
U+FFFD too. The stored bytes stay a function of the file alone, and the
surface writer no longer fails the run on such a file.

Tests: cs_rejected_scope_contained, cs_scanner_output_reads_back,
cs_utf8_surfaces (the former observation test now requires the writer to
succeed and the blob to be UTF-8).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…ild (S1, evaluator side)

The checkpoint writes an <ImportGroup>'s condition once into the project
blob (S1), but the evaluator still evaluated it again for every import of
the group, and an <ItemGroup>'s condition again for every <Using> of the
group: the product imports (or items) x condition length that S1 removed
from the blob came back as evaluator work.

An <ImportGroup>'s condition is now evaluated once, at the group's first
import, which is where MSBuild evaluates it: a property the group's first
import sets no longer takes the group's later imports away. A legacy blob
(the condition copied into every import) is still evaluated per import. An
<ItemGroup>'s condition is evaluated once per group for the items of one
pass; the final properties the items see do not change while they are read,
so the result is the same.

Tests: msbuild_import_group_condition_work and
msbuild_item_group_condition_work (doubling both the children and the
condition length at most doubles the evaluator's work counter; it
quadrupled), msbuild_import_group_condition_once.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The review names two scans whose failure returned no scope: the C# scope
scan and the project-file scan. Both record the failure on the file's doc
links since e775dcd9, but only the C# scan had a failure seam. The
project-file scan gets one (CBM_DOCLINK_ALLOC_PROJECT), and the status test
gets a seventh point: a project file whose scan ran out of memory fails the
layer (one error row), like the other six.

Test: alloc_failure_status (its repository now has a project file).
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…e check type)

make -f Makefile.cbm lint-ci failed on code of the checkpoint:

- doc_links_msbuild.c: outside a seam build, msb_fail_value_alloc and
  msb_fail_prop_insert were stubs that return false, so cppcheck reported
  every condition that calls them as always false or always true
  (knownConditionTrueFalse, five sites). The calls are now compiled only
  with CBM_ENABLE_TEST_SEAMS, as the layer's other seams are, and the stubs
  are gone.
- mcp.c: in one configuration of vendored yyjson.h's fallback typedef chain
  cppcheck takes uint32_t for unsigned short and reports the bounds of the
  U+E0000 tag block as out of range (compareValueOutOfTypeRangeError). The
  preview's escape check takes unsigned long, which the standard makes at
  least 32 bits wide.

The product binary is byte-identical before and after.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…e decides (S5)

A name that no enclosing type or namespace has is looked up in the using
directives of each scope level, and a name some namespaces declare walked
the min(postings, directives) of the level for every reference. A file with
r references to such a name under k directives cost r x k.

The using step of a lookup -- what one region's directives, and at the
global namespace the unit's, give one query -- is now remembered for the
pass: the references of one file (a per-file working state the doc-link
layer hands to the resolver: file_begin / file_end), and the build's passes
over using directives and base lists. The key holds everything the step
reads: the file, the region and unit asked, the context's visibility
(product code, unit, assembly) and the query. A memo that runs out of memory
stops remembering; the lookups stay exact.

Inside a lookup, the alias and using walks stop at a level's second
candidate: the result is ambiguous whatever the rest would bring. Which
candidate came first, and whether there was one, never depends on the
directives skipped; only the ambiguity tally of the log line (why) may count
such a reference under another kind.

There is no cap: every outcome is what it was.

Tests: cs_lookup_repeated_name_work (k 60 -> 240 and r 40 -> 160: 236 -> 900
steps, 3.8 x; without the memo 3004 -> 41444, 13.8 x),
cs_lookup_ambiguous_stops (one ambiguous reference under 50 -> 200 imports
that declare it: 17 -> 21 steps; without the stop 65 -> 219).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…ered

A memo is made for every file that has references, and most files ask few
using steps: its table and arena are now opened by the first step it
keeps (8 KB arena block, 16-entry table), not when the file starts.
Outcomes and work counts are unchanged.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
… (S8 across runs)

A C# scope blob written in a run that the reader refuses costs only its own
file in that run (a8c5708d). The blob itself was stored, though, and the next
incremental run read it back as a stored scope the reader refuses, which
fails the layer for the whole run.

The surface row now stores, in place of such a blob, the C# rejected marker
(the tag and one `!` record). Whether the reader takes a blob is decided
where the row is made, by the reader's own record checks on an index of the
blob's own (what makes a blob refused depends on the blob alone); the
doc-link layer asks the language through two new resolver fields,
scope_accepted and rejected_scope (cbm_doclinks_storable_scope). A later run
that reads the marker back treats the file as the run that wrote it: it
declares nothing, its references are graph gaps, it is counted with the
rejected files. A stored scope that does not parse for any other reason
still fails the run. The scope delta of a file whose blob is the marker on
either side is LOCAL only when both are.

The test seam that spoils a file's scope now does it where the scope is
written (internal/cbm/doclink_cs.c), so what is stored carries it too.

Test: cs_rejected_scope_across_runs (a full run with the spoiled file, then
a body-only edit elsewhere and one in the spoiled file: each incremental
step equals a full run, and the stored scope is the marker).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
cbm_doclink_note_doc_line returned silently when the map (or its first
block) could not be allocated. The doc it was noting then keeps its
definition's line, and a doc shared by the declarators of one field
declaration is no longer marked as taken, so its references are parsed
once per declarator -- while the layer reports ok.

Both allocation failures now set the file's doc_links.failed, as every
other lost piece of doc-link data does; the layer then says error.

A test seam, CBM_DOCLINK_ALLOC_DOC_LINE, fails the map's growth.

Test: alloc_failure_status gets the doc-line map as one more point (one
error row).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The scanner names every namespace region it writes (each segment holds at
least one identifier character). The reader took an empty name from the
store all the same: ns_make_path then returns the parent namespace, so the
region nests without deepening it, its types are declared one level up,
and lookup asks the directives of fewer regions than the file has.

parse_region now refuses an empty name field, so such a stored scope fails
the run with bad_scope like any other record this code does not write, and
the next index rebuilds from the files.

Test: cs_damaged_stored_scope runs each damage in turn; the new case blanks
the name of A's region record (one error row, no `missing` row for Target,
then a forced full run that restores the edge).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…root is unknown (R6)

The project-file scan wrote the "could not be read" blob for a *.csproj
that is no readable MSBuild <Project>, and for a *.props or *.targets file
only once its <Project> root was seen. A *.props or *.targets file that is
malformed before its root, or ends before one, got no blob at all: an
import of it counted as outside the repository with a closed scope, and as
the nearest Directory.Build.* it was passed over for the one above, whose
usings were then applied although MSBuild does not read that file.

Such a file now gets the "could not be read" blob too, so the projects that
evaluate it have an open scope. A *.props or *.targets file whose root
element is not <Project> is still no MSBuild file and has no blob.

Test: msbuild_malformed_props_opens (a malformed nearest
Directory.Build.props under a usable one: a name found nowhere is external,
and a type only the upper file's using reaches is not bound; decoy: a
project without a nearer file binds through the upper file and keeps a
closed scope).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…layer in both routes (R7)

The re-check suspected that the worker pipeline never reads the doc-link
failure flag of a C# file holding only directives, because its resolve
loop skips a file with no definitions, calls, usages, throws, reads or
writes, or impl traits. That case does not occur: every extraction pushes
the file's Module definition first (cbm_extract_definitions, before the
doc-link driver runs, with no return in between), so such a file has one
definition and reaches cbm_doclinks_resolve_file, which reads the flag.
No production code changes.

Test: directives_only_failure (a repository of 62 files whose only C# file
holds two global usings, its scope scan failing: one error row with four
workers and with one; none without the failure; the file has nothing but
its Module). Making the skip pass over a Module-only file turns the
four-worker case red (0 error rows), so the test binds that path.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…every run (R4)

A scope written in this run that the reader refuses costs its own file:
what its parse set is taken back and the file declares nothing. Other
files then found none of its declarations either: a reference to one of
its types was `missing`, or bound uniquely to another type of that name
which the full picture would not choose (a type of the enclosing namespace
comes before an imported one). Unplaced declarations are quarantined for
exactly this reason; a refused file's were not.

Now the type names the refused blob's T and Q records carry, read
leniently (a line counts when its name field is there and is a name,
however the rest of it reads), are quarantined, so a reference to one of
them from any file is a graph gap. The stored marker carries the same
names as Q records after its `!` record, so a later run that reads the
marker back quarantines them too and an incremental run equals a full one;
cs_scope_marked_rejected takes the marker followed only by Q records of
names. The resolver's rejected_scope is therefore a function of the
refused blob, and cbm_doclinks_storable_scope hands the marker to its
caller to free (lsp_surface.c).

One S8 expectation changes with this: Healthy.cs's reference to Hidden2,
a name Bad.cs declares where no namespace can be established (a Q record),
is now a graph gap rather than an edge -- as it is when Bad.cs's scope is
taken. The namespace checks of S8 (Shadow, Good.Lurker) are unchanged.

Test: cs_rejected_scope_contained and cs_rejected_scope_across_runs, whose
shared fixture gains Good.Widget in the refused file and Lib.Widget in an
imported namespace: Healthy.cs's Widget is a graph gap with no edge, and
the stored marker holds the five names.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…t once per file (R1)

The using step of a lookup is remembered per file (S5), so the unit's
directives -- its global usings and <Using> items, asked at the global
namespace -- were walked again by every file that asks a name: files x
min(directives, scopes declaring the name), a hang in resolution for a
project with many global usings and a common name.

What a unit directive gives a query reads only the context's unit,
assembly and product flag (level_using and what it calls), never the file.
The run now keeps, per (unit, assembly, product/test, query), the unit
directives that give the query anything at all, in order, up to where they
alone decide it (cs_unit_memo). Each file asks those after its own
directives and stops where the walk of all of them stops. The edges and
rows are unchanged: a directive that leaves an empty step empty makes no
change in any step, and a step that starts with the file's candidates is
decided no later than the unit's directives alone are (a candidate counts
unless it is the step's first).

The memo is made once the index is complete (the build's own lookups see
directives that are not all resolved yet) and is shared by the resolve
workers: each list is found exactly once, by the first worker that asks;
the others that ask for that key take its lock until the list is there, so
the work of a run does not depend on the scheduling. Out of memory, it
stops remembering and the walk is the full one.

A test seam builds the index without the memo.

Tests: cs_unit_usings_across_files_work (k global usings, k scopes
declaring the name, f files naming it once: k 150 -> 300 and f 100 -> 200
give 768 -> 1520 steps, 1.98 x; without the memo 17003 -> 64403, 3.79 x);
cs_unit_usings_unchanged (the same edges and rows with and without the
memo, over the cases of file and unit candidates: none, other, same, same
then other, test-only, another unit).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…o (R1)

The unit memo of the previous commit was asked -- the run's lock taken, an
entry made -- for every unit step, also when the unit has no directive or
its index rules the name out, where the walk returned at once before.
That is lock traffic and memory for most projects, which have no global
usings.

The test that walk starts with is now one function (usings_may_give),
asked before the memo too. The work counts of the doc_mentions cost tests
are those from before the memo again, except where unit directives can
give the name: one memo probe per lookup there.

Test: cs_unit_usings_across_files_work and cs_unit_usings_unchanged as
before (768 -> 1520 steps; without the memo 17003 -> 64403).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…t thread's caches, split suites

The first run of the doc-link suites on the Linux arm64 leg (gcc, ASan with
LeakSanitizer) found three things macOS does not show:

- The node pass bound files in the order the file system listed them, and
  the scratch table's cost depends on that order (the table is emptied per
  file and made anew only after a far larger file). doc_mentions_cs_scratch_table_work
  counted 102,912 steps on Linux against 26,048 on macOS for the same
  files. Files are now bound in path order: a file's bindings do not
  depend on the others, so only the cost changes, and it is the same on
  every system. The test itself had a second dependence: its 20,000-field
  file exceeded the extraction budget on that leg (5 s, defs=0), so the
  test ran without it; 4,000 fields were still cut there beside five other
  suites. It now uses 1,000 fields and 20 small files (3,008 steps against
  a bound of 4,160; 40,512 with the S13 fix reverted) and requires the
  large file to reach the index whole.
- A test thread that extracted C# on its own small stack ended without
  freeing its parser and node-kind caches, as worker threads do: 40 KB
  reported leaked at exit. It frees them now.
- The suite ran 1,014 s against the harness's 900 s per-suite wall clock
  (gcc on aarch64 makes ASan's stack-use-after-return check about 100x
  slower per call). Its 119 tests now run as doc_mentions (72),
  doc_mentions_cost (9, the work-counter tests) and doc_mentions_msbuild
  (38): about 230 s, 210 s and 480 s there.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The doc_links report added two early returns to handle_index_status, each
with its own free of the project name: raw allocator sites the memory-core
ratchet does not allow (mcp.c went from 778 to 779 on this base). A failed
allocation or report now sets a flag and leaves through the existing exit,
which frees the name and answers with the MCP allocation-error response as
before. The baseline of mcp.c is lowered to the 777 the ratchet now
measures.

Tests: doc_mentions (incl. the doc-link report fault injection), mcp pass;
lint-memory-core passes.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant