Skip to content

feat(doc-links): Markdown, ADR, reST, AsciiDoc and PDF sections link to the code they name - #2551

Open
DeusData wants to merge 26 commits into
feat/doc-links-corefrom
feat/doc-links-documents
Open

DeusData wants to merge 26 commits into
feat/doc-links-corefrom
feat/doc-links-documents

Conversation

@DeusData

@DeusData DeusData commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

Summary

Documentation becomes part of the graph, linked by its specific contents to the code it names. Markdown (and MDX), architecture decision records, reStructuredText (Sphinx), AsciiDoc (Antora) and PDF documents get a node per section or page, and each section links to the code files and, where the reference names or covers one, to the exact definition (function, method, class, ...). An agent can go from a section to the code it describes, and from any code node back to the documents that mention it, through the graph alone.

All of it is decided by the indexer, deterministically, from explicit references: links, paths, qualified names in code spans, Sphinx roles and directives, include directives with line ranges and tags, Antora resources and javadoc attributes, and structural mentions in PDF text. Bare names never become edges. A reference that does not resolve with certainty is recorded with a reason instead.

Stacked on #2550, the doc-comment layer (MENTIONS edges, the doc_link_unresolved table, index_status.doc_links), which is stacked on #2543. Upgrade: the semantic index version goes to 6, so an existing index is rebuilt in full once.

What you get

Nodes.

  • Section per heading (Markdown, reST, AsciiDoc), with the section's prose as its docstring. Text before the first heading belongs to the document's File node.
  • Section per PDF page, with the page's text as its docstring (get_code_snippet returns the page text, never PDF bytes).
  • ADR per architecture decision record (Markdown or reST), with id, status, date, deciders and title properties. A record's "supersedes" statement resolves to the other record; a statement (the phrase opens its line or field and names a record) becomes a SUPERSEDES edge; the same words in running prose stay rows.
  • AsciiDoc (.adoc, .asciidoc) and PDF files are discovered and indexed; Sphinx projects' extra source_suffix entries (e.g. Django's .txt) are read as reST.

Edges. MENTIONS from a section (or the document) to a File or a definition, with:

{"via":"markdown","syntax":"code_name","tier":"exact","line":42,"count":1,"target_lines":[10,24]}
  • via: markdown, rst, asciidoc, pdf. syntax: the family (below). line: where the reference is written. target_lines: present when the reference names a line range (an include with lines=/tag=, literalinclude, #L10-L24).
  • A line range binds the innermost definition that holds it, or the definitions it covers, or the file.
Format Families that ship as edges Resolved, kept as rows (not yet past the held-out audit)
Markdown link, path, code_name, code_path, supersedes (an ADR's statement opening its line or field: SUPERSEDES edges) bare_path (a file or one directory named alone in a code span); supersedes_prose (the same words in running prose)
reStructuredText role, autodoc, code_name, literalinclude, code_path, object (.. py:class::, .. method::, ...) include, kernel_doc
AsciiDoc include attribute (Antora javadoc/xref resources), code_path, code_name
PDF pdf_path, pdf_file, pdf_qn (bare names are never edges)

Unresolved references go to doc_link_unresolved with the layer's reasons (missing, ambiguous, external, test_only_target, not_indexed, graph_gap, unparseable) and are counted in index_status.

Navigation. trace_path inbound with edge_types:["MENTIONS"] from a symbol lists the sections that mention it; query_graph reads any of it. No new tool.

How references are resolved

  • Paths: relative to the document's directory, then the repository root (no bare-file-name fallback beyond the document's directory); documentation tools' base paths are honoured (Sphinx: /x is relative to the source dir; Antora: component/module/family resource IDs from antora.yml and the playbook).
  • Qualified names: exact definition match first, then a segment-aligned suffix of 2+ pieces; exact case before folded; Python module:attr only in Python; never product documentation to a test target unless the path names it. Five scope rules, each of which can only take a link away (a refused candidate still makes two candidates ambiguous): a name written from std is the standard library's, never a C/C++ definition or header here; a name written from a JVM root package (com., org., ...) matches only where it starts the package, never a relocated copy's tail; a Java/Kotlin/Scala/Groovy/C# type matches in its own case only (debug.log is no member of Debug); a definition in another copy of the document's project (snapshots side by side) never answers the document; a code name never binds a JVM package folder (a dotted config key or cross-reference such as snapshot.mode is not a package; Python packages keep binding).
  • Sphinx: the forms T, C.T, M.T, M.C.T under module/currentmodule/class context, Python re-exports through package __init__.py imports, primary_domain, intersphinx and extlinks from conf.py; C domain names by kind. A one-piece name is searched only for classes, exceptions, functions and data in the declared module, never for bare members.
  • AsciiDoc: include:: with lines= and tag(s)= (tags read from the included indexed file only), attributes resolved from the document and the Antora component, javadoc attributes to Java/Kotlin types and members.
  • PDF: a dependency-free text-layer extractor (zlib only; fonts, encodings, CMaps, Type 3 glyph names, /ActualText, object and cross-reference streams, damaged files repaired as far as the prototype did); structural mentions only.
  • Incremental: a section's edges belong to its document. Documents bound by line ranges into a file re-resolve when that file's body changes (a body-only edit leaves line numbers stale otherwise).

Proof on real repositories (development corpora; the held-out list was checked first)

What Corpus Result
Markdown links/paths/names fastapi, cosmos-sdk, dapr, curl 2,899 / 1,286 / 184 / 52 MENTIONS (before: file-to-file REFERENCES_FILE only); link pairs 408/409 equal to the field-test prototype (the one difference a deliberate self-link skip); 44/45 hand-checked name edges correct, the 45th fixed (exact-case precedence)
ADR detection and facts five corpora detection 182/182 and titles 182/182 identical to the prototype; status 179/182, date 174/182 (differences explained: archived status, braces in a date field)
reST / Sphinx django fb113765 6,695 of 7,238 prototype links identical; the rest are repeated-title attribution (same document→target, file instead of section) and deliberate refusals including the prototype's own audited errors
reST ADRs openedx/course-discovery 34/34 records
AsciiDoc / Antora junit-framework 6c8ca116 javadoc attributes 362/362 identical; includes 292 identical + 60 bound to the enclosing definition + 7 multi-definition ranges to the file
PDF text layer 749 PDFs from 21 repositories 720/720 comparable documents byte-identical to the prototype's text (1,492/1,492 pages), 0.73 s vs 14.7 s; flattened copies identical; 44,070 mutated files without a sanitizer finding
PDF links five corpora identical to the prototype pipeline (53/53)

Extractor fix found by the audits: one Field per class field

The AsciiDoc audit bound a Java field reference to a module-level Variable. Java, C# and Solidity list the same node kind as a field and as a variable, so every class field was emitted twice: as a Field of its class and as a Variable carrying the module's qualified name (demo.query.filters for Query.filters), which collided with package paths and took usages away from the field. A declaration that is already a Field is no longer minted again, and a field with an initializer is named by its declarator's name (it was named PROP = "v"). On junit-framework, Variable nodes drop from 2,899 to 685, and 6,938 of the 9,901 usages that went to the duplicates now reach the field. Swift and GDScript (variables only) and D and Dart (no duplicates) are unchanged.

A field and a method of one name in one class body (executable and executable(Path)) shared one qualified name, so the graph kept only one of them (in junit-framework 107 pairs: the field lost in 90, the method in 17). Such a field is now <Owner>.<name>#field (the base#suffix fence of #macro and the Rust cfg twins; every other field keeps its name), and a read, a write or a plain value reference that resolves to the method goes to its field twin; calls keep the method. junit-framework: Field 2,808 → 2,898 and Method 14,124 → 14,141; WRITES into those methods 66 → 0 and into their fields 1 → 96; USAGE into those methods 97 → 0 and into their fields 2 → 212. Limit: a C# partial class split across files can still collide (the check sees one class body). Both changes ride on this PR's index-format bump (CBM_SEMANTIC_INDEX_VERSION 6 over its base's 5).

Security review

An independent read-only review of the layer found four HIGH, one MEDIUM and one LOW issue, all in the PDF path and all fixed here, each with a test that fails when its fix is reverted, and each re-proven on the PDF corpus (byte-identical text before and after):

  • a chain of objects (a /Length naming a stream, object streams inside object streams) cost a stack frame per object: now constant depth (a /Length is read as a value; ISO 32000-1 7.5.7 for object streams);
  • forms drawing the next form many times, level under level: the first reading of each form stays whole, re-readings are bounded per document (256 MiB, counted);
  • one cross-reference stream named by many sections was decoded each time: now once;
  • objects whose value never ended were read to the end of the file each: an object is read up to the next object;
  • the mention scanner's line joins compared every candidate of the page: now the two lines' candidates;
  • a page's docstring keeps 256 KiB of its text, with a line saying where it was cut.

Tests

New suites doc_links_md (Markdown and ADRs, 10 tests), doc_links_pdf (15), doc_links_rst (7), doc_links_adoc (4); incremental-equals-full steps for every format, including line-bound documents after a body edit. Each PDF security fix has a test that fails when the fix is reverted.

Verification

Local ladder on this PR's tip (the doc-comment layer PR below it included):

  • macOS arm64: 8,619 passed, 0 failed, 10 skipped (153 suites).
  • Linux arm64 (gcc, ASan + LeakSanitizer): 8,457 passed, 0 failed, 9 skipped, and the gcc -O2 production build.
  • Windows: not run on this tip. Hosted CI does not run on this stacked PR (its base is feat/doc-links-core, not main; only DCO runs), so the Windows check comes when the stack is retargeted to main. The local Windows arm64 VM last ran on the previous tip (8,451 passed, 83 skipped); two route tests of edge_types_probe fail there identically on the base of this stack, so they do not come from this PR.
  • Lint: cppcheck, clang-format (Homebrew LLVM), no-suppress and the memory-core ratchet.
  • Held-out audits: below.

Precision audit

Each family ships as edges only if it passes a held-out audit registered before any output existed on the audit repositories: 100 sampled links per family (census below that, families under 15 links pooled), judged blind against the source, bar Wilson 95 % lower bound ≥ 0.90 and point ≥ 0.95. Held-out repositories: three per format, chosen by an objective rule from code search, disjoint from every development corpus.

PDF (done). The C pipeline ran on the eighteen repositories of the field test's two held-out PDF audits. Its links equal the frozen prototype pipeline's on the same graph in every repository (5,217 of 5,217), and the audited prototype links it makes take their labels (255 of 260):

Family Correct Point Wilson 95 % low Ships
pdf_path 61 of 61 1.000 0.941 yes
pdf_file 110 of 114 0.965 0.913 yes
pdf_qn 77 of 80 0.963 0.896 no: its references are below_bar_tier rows

Markdown, reStructuredText, AsciiDoc (done). Nine held-out repositories (three per format, the top-starred results of a code search for the format's markers that hold an ADR log, Sphinx autodoc, or AsciiDoc source includes; none used in development): 782 sampled links judged blind, a second reader on every link not judged correct and on a seeded tenth of the rest (102 rechecks, 100 % agreement).

Family Correct Point Wilson 95 % low Ships
Markdown link 100 of 100 1.000 0.963 yes
Markdown path 50 of 50 (census) 1.000 0.929 yes
Markdown code_name 98 of 100 0.980 0.930 yes
reST role 100 of 100 1.000 0.963 yes
reST autodoc 98 of 98 (census) 1.000 0.962 yes
reST code_name 99 of 100 0.990 0.946 yes
AsciiDoc include 92 of 92 (census) 1.000 0.960 yes
Markdown code_path 93 of 100 0.930 0.863 no: a file name in a code span that means the reader's own file
AsciiDoc code_name 15 of 25 (census) 0.600 0.407 no: Java overloads, a config key and a package read as members
AsciiDoc code_path 12 of 13 (census) 0.923 0.667 no
reST code_path 3 of 3 (census) 1.000 0.439 no: too few held-out links for the bar
ADR supersedes 1 of 1 (census) 1.000 0.207 no: too few held-out links for the bar
reST object, include, literalinclude, kernel_doc; AsciiDoc attribute none held out no: not yet audited

Supplementary audits (same method, each registered before its queries; new held-out repositories chosen by the same kind of objective code-search walk, disjoint from all earlier ones). A family that failed was fixed only from what its sample showed, and a fix is judged on a NEW sample, never on the one that found the error:

Family Correct Point Wilson 95 % low Ships
reST literalinclude 100 of 100 1.000 0.963 yes
reST code_path 99 of 100 0.990 0.946 yes (the one error, a bare ../, is fixed)
PDF pdf_qn, after the owner rule (a Go method whose name lacks its receiver matches only the receiver-qualified form) 97 of 99 0.980 0.929 yes
Markdown code_path, with a directory (src/app/main.go, docs/decisions/) 94 of 100 0.940 0.875 no: all six errors named one directory alone (doc/ of another project)
Markdown code_path without the bare part (two or more path segments) 97 of 100 0.970 0.916 yes
Markdown bare_path (a file or one directory alone) 28 of 41 judged across the samples above 0.683 0.530 no: config.yaml or doc/ in prose is as often the reader's or another project's
reST object 17 of 22 (census) 0.773 0.566 no: a function directive bound an example script's module; fixed, needs a new sample
ADR supersedes 37 of 38 (census, 12 ADR repositories) 0.974 0.865 no: a template's combined label Supersedes / Depends on: read as a statement; fixed
ADR supersedes, after that fix 144 of 163 (census, 8 repositories) 0.883 0.825 no: Supersedes: none (it extends ADR 31), another record as the subject; fixed
ADR supersedes, after those fixes 82 of 87 (census, 40 repositories, at most 25 per repository) 0.943 0.872 no: the remaining errors were in running prose; split into statements (the phrase opens its line or field and names a record: 44 of 45 on this census) and supersedes_prose (38 of 42, held)
ADR supersedes statements 80 of 81 (fresh census, 27 repositories, at most 25 per repository) 0.988 0.933 yes; its one error, a wrapped paragraph whose subject was the line before, is fixed (a statement may not continue a paragraph)
reST kernel_doc 2 of 2 (census) 1.000 0.342 no: too few held-out links
reST object, after the fix (a function or class directive never binds a module) 105 of 105 (census, 6 of 12 new repositories with links, at most 25 per repository) 1.000 0.965 yes
AsciiDoc code_path 66 of 95 (census, 10 of 18 Antora repositories with links) 0.695 0.596 no: a path in product documentation means the reader's own file (.circleci/config.yml in CircleCI's docs: 25 of the 29 errors; ~/.m2/settings.xml, "your pom.xml"), and .. as range syntax
AsciiDoc attribute none in 48 Antora repositories no: no held-out links
AsciiDoc code_name, after the field and file-name fixes, overloads judged at method level 43 of 80 (census, 7 of 36 new Antora repositories with links, at most 25 per repository) 0.538 0.429 no: two repositories that carry copies of other projects (Boost vendored into a game, several HBase/Cassandra snapshots in a research repository) gave 17 of 49 correct: a std:: name bound to Boost's specialization, a class bound to another snapshot's copy; the documentation-first Antora repositories gave 26 of 28
AsciiDoc code_name, after four scope rules (std::, rooted JVM names, type case, project copies) 61 of 83 (census, 5 new repositories with links) 0.735 0.631 no: 21 of 22 errors bound a Java package folder to a dotted name that is no code (Debezium's snapshot.mode, OSGi configuration ids, JMX domains, a Maven groupId)
AsciiDoc code_name, after a fifth rule (no JVM package folders) 76 of 82 (census, 4 new repositories with links) 0.927 0.849 no: Ruby predicate methods (sections? bound to sections), a method bound to its file, and name() where a zero-argument overload exists (one node per method name)

A family that does not ship still resolves: its references are doc_link_unresolved rows with reason below_bar_tier, and its tests run it switched on, so a later audit can switch it on without other changes.

Known limits

  • Repeated section titles in one document share one node (the existing Markdown rule), so their links attach to the document.
  • A changed Python re-export forces a full re-index (as any added definition name already does).
  • PDFs: no OCR; encrypted PDFs are not read.
  • Links through symlinked directories are not followed (the walk skips symlinks).

Third-party data

The PDF tables carry the Adobe Glyph List subset (BSD-3-Clause), the core-14 font widths of Helvetica and Times-Roman (Adobe), and Unicode Character Database data (Unicode License v3); notices are in THIRD_PARTY.md and the SBOM. scripts/gen-pdf-unicode.py regenerates pdf_unicode.c byte for byte.

…ns they name

A Markdown document is documentation in itself: its sections now link to
the code they reference explicitly, as MENTIONS edges (via "markdown") from
the Section node (the File node for text before the first heading, and for a
heading whose title repeats in the file, whose Section node it does not
own). An agent can go from a document to the code and back through the
graph: trace_path inbound with edge_types ["MENTIONS"] lists the sections
that name a definition.

Families (the "syntax" property; each with its ship gate):
- link: the destination of an inline link or a link reference definition
  (no URL, no in-page anchor, no image);
- path: a bare path in prose (a slash and a file extension);
- code_path: a code span or HTML <code> span holding a path, a file name or
  a directory, with an optional line range (#L3-L9, :3-9) or member
  (path::Name);
- code_name: a code span holding a qualified name (a.b, a::b, A#b, A\B,
  mod:attr).

Resolution is guess-free, with the rules the documentation field tests
measured:
- a path is read relative to the document and, when it has a directory
  part, to the repository root; a file name alone only next to the document;
- a line range binds the innermost code definition holding the whole range,
  else the file with "target_lines" on the edge; path::member the one
  definition of that name in the file;
- a qualified name binds the one code definition whose qualifier chain ends
  with the written qualifier, else the one module or package so named; a
  match as written outranks one that needs case folding; a field binds only
  through its type, never through an instance; module:attr only in Python;
  product documentation never binds test code;
- a bare name is never an edge.
Unresolved references are rows in doc_link_unresolved: missing (the
directory or the repository's own root is there, the name is not),
ambiguous, test_only_target. A URL, an image, a URL route (/docs), a file of
a type the indexer does not read, and code of other projects are no
references to the repository and leave no row.

Extraction scans the document's bytes once (doclink_md.c); every scan of a
line is linear, or n log n where it sorts, using position lists that grow
with the bytes they track, not with the line. Front matter, fenced code,
HTML comments and MDX import/export lines are not scanned.

The incremental route's node proxies now carry their line numbers
(pipeline_delta.c), so a re-resolved document binds a line range exactly as
a full build does; the Markdown resolver reads definitions from its own
by-file list of the run's nodes, never from DEFINES edges, which an
incremental run holds for re-extracted files only.

Measured on four repositories against the main branch (which links file to
file only, REFERENCES_FILE): fastapi 2,899 MENTIONS edges (2,225
REFERENCES_FILE before), cosmos-sdk 1,286 with 205 distinct definitions
reached and 188 edges from 41 ADRs into code, dapr 184, curl 52.
Link pairs agree with the field-test harvester (cosmos-sdk 408 of 409; the
difference is a link of a document to itself, skipped on purpose). A hand
check of 45 sampled qualified-name links found one wrong, fixed by the
exact-case rule above.

Tests: doc_links_md (extraction, repeated headings, the span classifier
table, resolution with every reason, back navigation, incremental == full
for a code edit, a document edit and an added name).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
… facts and SUPERSEDES

A Markdown file that is an architecture decision record now also defines an
ADR node (File DEFINES ADR): its name is the record's canonical id ("ADR-12";
another prefix as written, "DEC-5"; a dated record keeps its file stem), its
qualified name "<module>.__adr__", and its properties carry the record's
facts: adr_id, title, status (normalized: proposed, accepted, implemented,
superseded, deprecated, rejected, abandoned, draft, unknown), status_text,
date, deciders, superseded_by, and the decision itself as the docstring, so
search finds a record by what it decided. Its sections link to code like any
document's (MENTIONS, previous commit).

Detection is per file, with the field tests' rules: a file in a conventional
ADR directory (adr, adrs, decisions, decision-records, ...) whose name carries
an id or whose text has ADR structure; any adr-NNN file; a numbered or dated
file with ADR structure. Never README, index or template files without an id,
never test fixtures. Facts come from front matter, then header fields
(bullets, bold, tables), then the Status section; the date also from a
"Last updated" field, the change log's first date and the file name.

An ADR's own statement that it supersedes another ("Supersedes [ADR-3](...)",
"Supersedes ADR-3"; "Replaces ADR-3" where it opens a line, never "replaces"
as a verb of prose) becomes a SUPERSEDES edge between the two ADR nodes, by
link or by id. "Superseded by" is written in the old record about another
file; an edge from the other node would be owned by a file whose text does
not hold it, so it stays a fact of the old record (status, superseded_by).

The link layer gains a per-family edge type (doclink.h: the family list's
"edge" column), so a family can produce SUPERSEDES instead of MENTIONS; edges
are grouped by (type, source, target). Definitions gain a tail field for a
document's facts (extra_props: key, value pairs), carried through result
compaction and written by both property builders.

Measured against the field-test harvester on five repositories: detection
182 of 182 records, the same files; title 182/182; status 179/182 (cosmos
ADR-050 and its annexes say "Status: ARCHIVED", kept as status_text with
status unknown); date 174/182 (semantic-kernel's "date: {2024-05-02}" is read
as a date where the harvester dropped it). SUPERSEDES: semantic-kernel
0038 -> 0015, log4brains 20201016 -> 20200926, each checked against its
sentence. Markdown MENTIONS unchanged on fastapi, curl, dapr, cosmos-sdk.

Not covered: ADRs in reStructuredText, and ADR-log configuration files
(.adr-dir, log4brains) as a detection source.

Tests: doc_links_md_adr (detection and exclusions, facts from front matter
and from a Status section, SUPERSEDES by link and by id, no edge from prose,
the record's MENTIONS, incremental == full for a status edit and for an
added record named by id).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…names

A PDF is now indexed: its text layer is read by a dependency-free extractor
(internal/cbm/pdf, zlib only) and every page becomes a Section node ("page N",
the page's text as its docstring, property page), so search finds what a page
says and get_code_snippet shows a page's text, never the file's bytes. The code
a page names structurally becomes MENTIONS edges from that page, via "pdf":

  pdf_path   a repository path ("pkg/server/handler.go", "./x/y")
  pdf_file   a file name with a code extension ("models.py", "Pool.sol")
  pdf_qn     a qualified name ("app.models.User.save", "a::b", "x->y")

Bare identifiers never become edges (field test E6c: a type-target rule failed
its pre-registered bar), and only EXACT or UNIQUE resolutions link: the field
test's tiers (path-exact, component-aligned path-suffix, qn-exact, qn-suffix,
the latter also through a node's owner -- a Go method's QN has no receiver --
and with the same one-entity collapses). A mention split across a line break is
proposed as one reference and replaces its fragments only when it resolves.
Hygiene as measured: no links from PDFs under test or fixture directories (R1),
none to data files (R2), none for keyword arguments (R3); a test, mock or
fixture target is a test_only_target row (R4, with the test directories the
field test found missing). A path or name whose first component is not of this
repository is no row at all.

The extractor is a port of the field-tested prototype: xref tables and
streams, /Prev chains, repair near a wrong offset and reconstruction from the
object headers, object streams, Flate (with partial recovery), LZW, ASCII85,
ASCIIHex, RunLength and PNG/TIFF predictors, simple fonts (base encodings,
/Differences, built-in Type1 encodings, standard-14 widths, Type3), ToUnicode
and encoding CMaps, Type0/CID fonts with the embedded TrueType cmap read
backwards, Form XObjects, inline images skipped, word and line breaks from
glyph geometry, right-to-left runs reversed. Its geometry is bit-exact on every
platform (no FMA contraction; CPython's hypot). Untrusted input: every read is
bounds-checked, recursion is depth-limited, integers from the file saturate, a
decoded stream is cut at 512 MiB, and every lookup that could rescan the file
per object or per glyph is answered from an index built once. Not read:
encrypted documents (reported as such); scanned pages have no text layer.

The mention scanner is the field test's (idents.py) with CPython's semantics:
NFKC per code point plus canonical composition and the str predicates, from
generated Unicode 16.0 tables (pdf_unicode.c), and re's leftmost-first
matching.

Pipeline: CBM_LANG_PDF (no grammar; the registry ledger gets a DOCUMENT
capability, the call-node manifest exempts it like the other grammar-free
classification), extraction before any grammar lookup, a resolver hook
file_prepare (a file's state from all its references before any resolves).
Incremental: a proxy node carries parent_class where owner + name differs
from its QN (Go methods, class-body variables), so a re-resolved document binds
a Go method as a full build does.

Proof against the field-test prototype on its corpus (749 PDFs from 21
repositories, restored at the recorded commits, sha256-verified): text
byte-identical on all 1,492 pages of the 720 comparable PDFs (0.73 s vs 14.7 s);
mention tokens identical on all 960 page-text sets of E6, E6b and E6c (26,435
tokens); links identical to the frozen Python pipeline run over the same graph
on semantic-kernel 5/5, cosmos-sdk 30/30, Polly 11/11, mediastore3 3/3, dask
4/4. 44,070 seeded mutants (truncations, flips, nesting and keyword bombs, on
the originals and on uncompressed copies) under ASan and UBSan: 0 findings
after the first campaign's integer overflows were fixed as a class.

Tests: suite doc_links_pdf (layout, fonts and CMaps, object and xref streams,
xref repair and reconstruction, Form XObjects incl. a self-referencing one,
inline images, filters, hostile input, the scanner against the field test's
output, the pipeline's tiers, hygiene and rows, back navigation, the snippet,
incremental == full).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…arkup names

A reST document now gets a Section per title (the text up to the next title
as its docstring, its lines from the title to the next one) and MENTIONS
edges, via "rst", from each section to the code it names:

  role            :class:`pkg.Model`, :meth:`~x.Y.z`, :func:`.f`, :c:func:`f`,
                  and extlink roles whose URL is a repository blob path
                  (:source:`django/db/models/base.py`)
  object          .. py:class:: / .. method:: / .. function:: ... and their
                  C-domain twins
  autodoc         .. automodule:: / .. autoclass:: ...
  include,        .. include:: / .. kernel-include:: of a code file,
  literalinclude  .. literalinclude:: (its :lines: bind the innermost
                  definition holding them, kept as target_lines; its
                  :pyobject: the named definition)
  kernel_doc      .. kernel-doc:: path and its :identifiers: / :functions:
  code_path,      inline literals naming a path or a qualified name (the
  code_name       Markdown code-span rules)

Structure is the field test's line model (H3): titles by adornment style,
literal blocks after `::` and in code directives, comments, directive
arguments and options, Sphinx's module / currentmodule / class contexts,
CPython's str whitespace and \w. Text before the first title belongs to the
file. A Sphinx documentation set's conf.py is read as text, never run: its
source_suffix makes discovery index e.g. Django's `.txt` documents as reST in
that set only, its primary_domain decides undomained roles, its extlinks and
intersphinx names are known to the resolver.

Python names resolve the way Sphinx imports them: the longest prefix that is
a module's import name (named from its topmost package directory), the rest
by qualified name, a package's re-exports (`from .base import Model` in its
__init__.py) followed. The re-exports and the conf.py settings travel as
Python scope blobs (doclink_py.c), so an incremental run resolves as a full
one; one resolver serves reST and Python. Forms in Sphinx's order (T, C.T,
M.T, M.C.T; a leading `.` reverses it). EXACT when a form names a node,
otherwise UNIQUE by name with the field test's fixes: a bare member
(:meth:`save`, :attr:`x`) is never searched (37 of 40 such results were
wrong), an object directive's target lies in its declared module, a module
never answers a member's role (`:meth:`dispatch`` is not tests/dispatch),
test code is no target of a search but binds when the document writes its
name in full (`django.test.TestCase`). C identifiers are whole names: the one
C/C++ definition of that name and kind. Unresolved: missing when the name or
path is the repository's, ambiguous, test_only_target; another project's
name (stdlib, intersphinx) is no row.

reST architecture decision records: the ADR detector reads reST records
like Markdown ones (titles underlined with = or -): course-discovery's 34
decision records become ADR nodes.

Proof on Django's documentation (the field test's corpus and commit), C
against the frozen H3 harvester over the same graph, per (document,
section, target): 6,695 of H3's 7,238 links identical; H3-only 543 = links
C attributes to the document because a section title repeats in it (the
Markdown rule; the same document -> code link exists), deliberate refusals
(bare members, directive searches without a module) and H3's own audited
error class (router protocol methods -> ConnectionRouter, forms Field ->
model Field); C-only = the same attribution mirrored, plus 1,053 inline
literal links (sample of 30: 30 correct). 6,574 reST sections; 12 s for the
whole index.

Tests: suite doc_links_rst (structure and tokens, the scope blobs, the
pipeline's rules, source_suffix in discovery, reST ADRs, incremental ==
full). Revert check: without the module-kind guard and without the
source_suffix mapping exactly their assertions fail (observed). Two
expectations of the old behaviour now name the new one: the reST grammar
golden and probe have the document's Sections (they recorded "reST titles
are no Sections", the gap this closes).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
… edit

The closure route treats a body-only edit (the file's surface, which holds
no line numbers, unchanged) as "the file re-resolves, nobody else does".
That holds for references that bind names, but a document can bind lines:
Markdown `app/routing.py#L8-L9` and reST `.. literalinclude:: :lines:` bind
the innermost definition holding those lines. A body edit that moves code
left such a document on its old target on an incremental run, while a full
index bound the new one (observed: the doc_links_md and doc_links_rst
"lines moved" steps, incremental != full).

The documents (Markdown, reST: cbm_doclinks_binds_lines) with edges into a
body-edited file now join the repair closure, from one more dependent-files
query over the body-edited files. Code files stay out: their references
bind names, which a body edit keeps. The query shares the existing
dependents query's block and failure path (no new allocation site).

Tests: a "lines moved" step in doc_links_md_incremental and in
rst_links_incremental; with the predicate disabled exactly these two steps
fail (observed). doc_links_md/rst/pdf, doc_mentions, incremental, pipeline
pass.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…e and name

AsciiDoc (.adoc, .asciidoc) is now indexed, without a grammar like PDF: a
Section per heading (`=` .. `======`, or `#`), the text up to the next
heading as its docstring, delimited blocks respected (listing, literal,
comment, passthrough, example, sidebar, quote, table, fences). Its
references become MENTIONS edges, via "asciidoc", from the section they are
written in:

  include     include::target[...] -- Asciidoctor's preprocessor directive,
              read in every block but a comment block; the page's own
              attributes, then the Antora component's and the playbook's,
              substituted
  attribute   an {attribute} whose value is an Antora javadoc link
              ({javadoc-root}/<module>/<package path>/<Type>.html[#member])
              binds the Java or Kotlin type, or its one member, that the
              repository's main sources declare under that path
  code_path,  monospace spans naming a path or a qualified name (the
  code_name   Markdown code-span rules)

The rules are the field test's (H8 Asciidoctor/Antora harvester, audited
~100 % precise): an Antora resource ID ([component:][module:]family$path)
names <component>/modules/<module>/<family dir>/<path>, or the file a
collector scan contributes there (a scan of build output is generated: no
row); any other target is relative to the including file. `lines=` and
`tag=`/`tags=` regions bind by H8's rule: edge blank and comment lines
trimmed, the innermost enclosing definition, unless that is a type (or there
is none) and the definitions inside cover every other line (an annotation
right above one counts) -- then that definition; tag regions come from the
included file's own `tag::x[]` / `end::x[]` comments, read from that indexed
file; a tag the file lacks binds the file. One reference makes one edge: a
region holding several top-level definitions binds the file.

Antora settings travel as YAML scope blobs (antora.yml: component name,
asciidoc attributes, collector scans; antora-playbook*.yml: attributes),
read as text, so an incremental run resolves as a full one; a change to them
is GLOBAL. AsciiDoc joins the documents bound by lines: a body edit of an
included file re-resolves its pages.

Also: the Markdown resolver's module index no longer takes AsciiDoc or PDF
files for code modules.

Proof on junit-team/junit-framework at 6c8ca116 (H8's corpus and commit),
C against H8's recorded links: javadoc attributes 362/362 identical;
includes 292 identical, 60 bind the type enclosing the several definitions
H8 names, 7 regions with several top-level definitions bind the file. 317
sections, 668 edges, no include rows (9 generated partials, 1 file type the
indexer does not read); monospace spans 30/30 correct by hand. 8.3 s for the
whole index.

Tests: suite doc_links_adoc (structure and tokens, the Antora blob, the
pipeline: resource IDs through a collector scan, tag regions, a region with
its imports binding its class, lines=, javadoc attributes, spans; incremental
== full: a moved tag region, a changed component attribute). Revert check:
without H8's region rule and without AsciiDoc among the line-bound documents
exactly their assertions fail (observed).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…het measures

With the line-bound documents' re-resolution sharing the dependents block
of an incremental run (no raw allocation site of its own), the file holds
143 raw allocator sites; the memory-core baseline follows (a file may only
go down).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The PDF text layer carries data tables (D3). Their sources and notices are
now in THIRD_PARTY.md (shipped in every archive through
THIRD_PARTY_NOTICES.md) and in the SBOM:

  - the Adobe Glyph List subset (BSD-3-Clause, Copyright 2002-2019 Adobe),
  - the printable-ASCII advance widths of Helvetica and Times-Roman from
    Adobe's core-14 font metrics (the AFM files are not distributed),
  - the Unicode Character Database data of pdf_unicode.c (Unicode
    License v3, full text included).

scripts/gen-pdf-unicode.py (stdlib CPython only) regenerates
pdf_unicode.c byte for byte with CPython 3.14.6 (Unicode 16.0.0);
pdf_tables.c is maintained data from here on. License texts were taken
from their sources (adobe-type-tools/agl-aglfn, unicode.org).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Markdown sections linked to code, ADR nodes, reStructuredText and
AsciiDoc sections, PDF pages and the files newly discovered for them
change the graph of documents that did not change. An index written
before (version 5, the doc-comment layer) keeps those documents as they
were, so it is rebuilt in full once.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Loading a stream whose /Length referred to another stream loaded that
stream first, and its /Length the next one: one nested call per object.
Loading an object of an object stream whose /Length (or filter) lay in
another object stream opened that one inside the first, again one call
per stream. A crafted chain overflowed the stack (D1 of the document
layer security review).

A /Length is now read as a value only: an object that is a stream there
gives no length (the stream is then read up to its endstream, as for any
wrong length) and is marked, so a second reference costs nothing. The
objects that open an object stream are no objects of an object stream
(ISO 32000-1, 7.5.7); one that is, is refused and counted
(objstm_nested). Both rules leave the nesting constant; nothing is cut
by a limit.

Test: pdf_extract_object_chains extracts a 20,000-deep /Length chain and
a 3,000-deep object-stream chain on a 256 KiB thread stack (stack
overflow before; each half fails alone when its rule is reverted). The
749-PDF dev corpus extracts byte-identically to the prototype as before
(720 of 720 documents, 1,492 of 1,492 pages).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Every classic section that named an /XRefStm parsed and decoded that
stream afresh, and each decoded copy was kept until the document closed:
thousands of tiny sections naming one highly compressible stream asked
for thousands of copies (D3 of the document layer security review).

A target already merged is skipped: a later (older) section naming it
again adds nothing, since every number it holds is in the table already.
A cross-reference stream's decoded bytes are freed once its rows are
read (a stream parsed there is no cached object), so distinct streams
cost one at a time. The filter pipeline of pdf_stream_decode is split
out for that (decode_filters); cached streams decode as before.

Test: pdf_extract_xrefstm_once: a thousand sections naming one stream
decode it once (xref_streams == 1; 1,000 with the skip reverted). The
749-PDF dev corpus extracts byte-identically as before (720 of 720
documents, 1,492 of 1,492 pages).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Objects whose value began a token that never ended (a string without its
closing parenthesis) were each read to the end of the file, and the
cross-reference rebuild reads every object: n such objects cost n times
the file in time and in arena memory (D4 of the document layer security
review). The same held for scanned trailers and for the members of an
object stream.

The objects of a file do not overlap, so an indirect object's value is
now read up to the next object header (the header scan, built once per
document), a scanned trailer up to the next "trailer", and an object
stream member up to the next member. A stream's bytes are still placed by
its /Length. The total work is linear in the file; nothing is cut that a
well-formed file holds.

Test: pdf_extract_open_strings_linear: a thousand objects with open
strings and no cross-reference grow the arena and extraction classes by
at most 16 x the file + 1 MiB (67 MB for a 112 KB file with the window
reverted). The 749-PDF dev corpus extracts byte-identically as before
(720 of 720 documents, 1,492 of 1,492 pages), and its 720 flattened
copies give the same output before and after.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
A Form XObject that drew the next one many times, level under level,
was read (draws ^ depth) times: a few kilobytes asked for an unbounded
reading, and the page text grew with every leaf (D2 of the document
layer security review). The forms' text cannot be read faithfully; it is
exponential in the file.

The first reading of each form in a document stays whole. Reading one
again (the same content drawn elsewhere: a logo, a header on every page)
is counted in decoded bytes; beyond 256 MiB of such re-readings the
document reads no form again, and each skip is counted
(form_rereads_cut), as a decoded stream above 512 MiB already is. Real
documents re-read small forms: the 749-PDF dev corpus extracts
byte-identically as before (720 of 720 documents, 1,492 of 1,492 pages;
flattened copies too). Reading each form only once per page was tried
first and changed one real page, so it was not taken.

Test: pdf_extract_form_ladder_bounded: seven levels of forms drawing the
next four times each, 64 KiB per form: re-readings are cut and the page
keeps its glyphs bounded (with the bound reverted nothing is cut).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
For each line break that joins a name, the scanner listed the
page's candidates overlapping the joined span by reading every candidate of
the page: (joins x candidates), quadratic in a page's text, and the
geometry deciding the line breaks is the file's (D5 of the document
layer security review).

A join's span covers the end of one line and the start of the next, so
the emitted candidates are indexed by line once per page and a join
merges the two lines' lists: the same candidates in the same order, so
the same tokens.

Test: pdf_scan_join_linear: a page twice as long costs at most 2.5 x the
join work (seam counter; 717,600 vs 2,875,200 with the full scan
restored). The scanner's expected-token test is unchanged, and the PDF
links on three dev corpora (Polly, mediastore3, dask) equal the frozen
field-test pipeline's as before (0 links only in C, 0 only in Python).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
A PDF page's Section carried the whole page text as its docstring, which
is also what get_code_snippet returns for it: a crafted text layer
stored megabytes per node and handed them to an agent whole (D6 of the
document layer security review). Markdown, reStructuredText and AsciiDoc
sections keep a short docstring already.

A page's docstring is now its text up to 256 KiB, cut on a character
boundary and followed by a line that says where it was cut and how long
the text is. Pages hold a few kilobytes; the mentions are still read
from the whole text.

Test: pdf_page_doc_bounded: a 315 KB page keeps 256 KiB plus the note,
and a path written at its end is still a mention (the whole text is
stored with the bound reverted).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The pre-registered held-out check of the PDF families: the C pipeline was
run on the eighteen repositories of the field test's two held-out PDF
audits (its links equal the frozen prototype pipeline's on the same
graph: 5,217 of 5,217, none only in one), and the audited prototype
links it makes take their labels (255 of 260 audited links). Bar per
family: Wilson 95 % lower bound >= 0.90 and point >= 0.95.

  pdf_path   61 of 61 correct    point 1.000   low 0.941   ships
  pdf_file  110 of 114           point 0.965   low 0.913   ships
  pdf_qn     77 of 80            point 0.963   low 0.896   does not ship

pdf_qn misses the bar: its references now become below_bar_tier rows
(the family's ships flag), as the pre-registration says, with no rule
tuned on the held-out data. Its three errors, all in one repository, are
names of an external API matched to a same-named local definition. The
pipeline test now
asserts the rows, and back navigation through the file a page names.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The pre-registered held-out audit of the document families: nine
repositories chosen by an objective rule and disjoint from every
development corpus (three each for Markdown, reStructuredText and
AsciiDoc), 782 sampled links judged blind against the source, a second
reader on every link not judged correct and on a seeded tenth of the
rest (102; agreement 100 %). Bar per family: Wilson 95 % lower bound
>= 0.90 and point >= 0.95.

  ships   markdown link 100/100, path 50/50 (census), code_name 98/100;
          rst role 100/100, autodoc 98/98 (census), code_name 99/100;
          asciidoc include 92/92 (census)
  rows    markdown code_path 93/100 (a file name in a code span that
          means the reader's own file, not the repository's);
          asciidoc code_name 15/25 (Java overloads, a config key and a
          package read as members, a binary name read as a manifest);
          asciidoc code_path 12/13; rst code_path 3/3 and supersedes 1/1
          (censuses too small for the bar)
  rows    not audited, no held-out occurrences: rst object, include,
          literalinclude, kernel_doc; asciidoc attribute

A family that is not shipped still resolves; a resolved reference is a
doc_link_unresolved row with reason below_bar_tier. The suites of these
families test their resolution with the families forced on (the test
seam), and doc_links_md_ship_gate tests the production gate (red with
one held family switched back on).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The Linux arm64 production build (gcc -O2 -Werror) stopped on
maybe-uninitialized in the TrueType cmap readers and the cross-reference
stream rows: ttf_fmt4 and ttf_fmt12 ignored the results of rd16/rd32 after a
bounds check that makes every read succeed, and xref_rows read pairs[] that
only the default /Index branch fills. The ASan test build did not warn.

The reads now fail the map on a failed read (the prototype's short-slice
rule; never taken, the bounds check stands before them), and pairs[] starts
at zero where it is declared. Text output is unchanged.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…milies ship

Supplementary held-out audits, each registered before its queries and run on
repositories chosen by an objective code-search walk, disjoint from every
earlier audit and development corpus (blind readers, a second reader on
every link not judged correct and on a seeded tenth of the rest; bar: Wilson
95 % lower bound >= 0.90 and point >= 0.95). A fix is judged on a NEW sample,
never on the one that found the error.

Ship now:
- Markdown code_path, refined to code spans with two or more path segments:
  97 of 100 (lower bound 0.916). A file or one directory named alone in a
  code span is its own family, bare_path, held: across the judged samples it
  was 28 of 41 correct (`config.yaml` the reader's own file, `doc/` another
  project's directory). A span without a word character (`../`, `./`) is no
  path at all.
- reST literalinclude: 100 of 100; an open range (`:lines: 5-`, `-5`) now
  reads the included file for its end and binds like a closed one.
- reST code_path: 99 of 100; its one error, a bare `../`, is fixed above.
- PDF qualified names: 97 of 99 after the owner rule: a definition whose
  qualified name lacks its owner (a Go method under its package) matches
  only owner-qualified text, so `time.Now` no longer binds a receiver's Now.

Fixed, still held:
- reST object directives: a function or class directive never binds a
  Module, Folder or File (an example script's module took them); a new
  sample is due.
- ADR supersedes: three censuses failed (37/38, 144/163, 82/87). Fixed from
  them: a combined template label (`Supersedes / Depends on:`), a value of
  "none" or "n/a" followed by records it extends, another record as the
  subject (`ADR 0239 supersedes ...`, a table cell `0259 (supersedes ...)`),
  a full stop after a number ending the target sentence while a
  `[text](destination)` link stays one reference (adr-tools'
  `Supersedes [3. Title](0003-...md)`). The family is split: a statement that
  opens its line or field and names a record directly stays `supersedes`
  (44 of 45 on the last census); the same words in running prose are
  `supersedes_prose` (38 of 42), held. Both still held in this commit; a
  fresh census of statements passed afterwards and is acted on next.
- An ADR log in a repository without code definitions lost every supersedes
  link by file: the ADR table was filled after an early return. It is now
  filled first.

Every fix has a test that fails with the fix reverted.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
A fresh held-out census of ADR supersedes statements (the phrase opens its
line or field and names a record directly) passed: 80 of 81 correct (point
0.988, Wilson 95 % lower bound 0.933; 27 ADR repositories from the same
objective code-search walk, at most 25 links per repository, blind readers,
second-reader agreement 1.0). The family ships; the same words in running
prose (`supersedes_prose`) stay rows.

Its one error was a paragraph wrapped before the phrase:
`[ADR-0240](0240-x.md)` / `supersedes ADR-0134 and ...`, the subject on the
line before. A statement may no longer continue a paragraph: a line that no
list, quote or table mark opens, after a line of running text (not blank, a
heading or a `Key: value` field), is prose unless the phrase is itself a
field (`Supersedes:`). On the census this moves exactly that link to prose;
the 80 correct ones stay statements and none is added. Tests cover the
wrapped paragraph, consecutive field lines and a field under running text,
each failing with its rule reverted; the ship gate checks a statement edge
and a prose row in production.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Java, C# and Solidity list the same node kind as a field and as a
variable, so every class field was emitted twice: as a Field of its
class and as a Variable carrying the module's qualified name
(`demo.query.filters` for field Query.filters). The Variable collided
with package paths, and usages bound to it instead of the Field.
extract_class_variables now skips a declaration that extract_class_fields
already made a Field (walked alongside the body's Fields, in order).

Fields with an initializer were named by the whole declarator text
(`PROP = "v"`), so the duplicate Variable had been the only findable
node for them; resolve_field_name_node now takes a variable_declarator's
own name.

Measured on junit-framework: Variables 2,899 -> 685; of 9,901 usages
into the removed duplicates, 6,938 now reach the Field. Swift and
GDScript (Variables only) and D/Dart (no duplicates) are unchanged.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
A code span naming `OliveTin.exe` was read as member `exe` of the
module stem of OliveTin.exe.manifest. Extensions of built and shipped
files (exe, msi, deb, rpm, dmg, tar, xz, bz2, 7z, apk, ipa, manifest,
csv, tsv) now make a dotted token a file name; none of them is a
common member name.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Sphinx object directives (`.. py:class::`, `.. method::`,
`.. function::`, ...) bind the definition they document. The family
was held back until its held-out audit: a census of 105 links in
repositories never used for development, judged blind in two rounds,
found all 105 correct (Wilson 95% lower bound 0.965).

On django's documentation (a development corpus) the 2,046 held rows
become 1,895 edges: 1,640 new section-to-definition pairs, and 255
pairs that a role or code span already linked now carry the directive.
No other edge or row changes.

include and kernel-doc stay held. rst_ship_gate tests the production
gate itself, after the suite's forced families are reset.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…field`

A Java class may declare a field and a method of one name (`executable`
and `executable(Path)`), and the graph keeps one node per qualified name,
so one of the two was lost: the field in 90 and the method in 17 of the
107 such pairs in junit-framework. With fields named correctly (the
previous commit) every initialized field joined these collisions.

When the same class body declares a method of that name, the field now
takes the qualified name `<Owner>.<name>#field`, the `base#suffix`
fence of `#macro` and the Rust cfg twins; every other field keeps its
name. The registry already files a fenced symbol under its plain name,
so name lookups still reach it. A read, a write or a plain value
reference that resolves to such a method goes to its field twin instead
(cbm_registry_value_target); calls and callable references keep the
method.

junit-framework (development corpus), before -> after: Field 2,808 ->
2,898, Method 14,124 -> 14,141 (the 107 pairs, both nodes kept); WRITES
into these methods 66 -> 0 and into their fields 1 -> 96; USAGE into
these methods 97 -> 0 and into their fields 2 -> 212 (accessors reading
their own field were self-edges before); CALLS into the methods 1,158 +
24 (then landing on the field) -> 1,215.

A C# partial class split across files can still collide: the check sees
one class body.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
AsciiDoc code names (resolved like Markdown's) failed their held-out
census at 43 of 80. Four of its error classes have a rule that holds
for every format:

- a name written from the root `std` never binds a C/C++ definition or
  header of the repository: `std::complex` is the standard library's,
  not Boost.Typeof's `typeof/std/complex.hpp`;
- a name written from a JVM root package (`com`, `org`, `java`, ...)
  binds a JVM definition only where it starts the package (after a
  source-root piece such as `src/main/java`), never as the tail of a
  relocated copy (`...hbase.shaded.com.google.protobuf.Message`);
- a capitalized type of a Java, Kotlin, Scala, Groovy or C# definition
  matches in its own case only: `debug.log` is a log file, not the
  method `log` of class `Debug`;
- a definition in another copy of the document's project never answers
  the document: below the directory where the two paths part, the
  document's own branch holds the same rest of the path, at least two
  directories deep, as a file with definitions (snapshots side by side;
  a shared `__init__.py` is no copy).

Each rule only takes a link away: a refused candidate still counts, so
two candidates stay ambiguous, and it is never the one linked.

Proof: six development corpora (junit-framework, fastapi, dask,
pydantic, curl, django) keep all 1,730 code_name links. Replayed on the
judged held-out samples: Markdown and reST code names keep all 195 links
judged correct and add none; on the failed AsciiDoc census the rules
remove judged-false links only (7 in the vendored-Boost repository,
which loses 127 code_name links and gains none). The first version of the copy rule dropped
one (a package's `__init__.py` taken for a copy) and turned ambiguity
between copies into new links; both are fixed here. The AsciiDoc family
stays off until a new census.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
A second held-out census of AsciiDoc code names (61 of 83) failed on
one class: 21 of its 22 errors bound a Java package folder to a dotted
name that is no code at all. Debezium's `snapshot.mode` and
`debezium.sink.type` are configuration keys, OpenNMS's
`org.opennms.features.amqp.eventreceiver` is an OSGi configuration id,
`org.apache.cassandra.auth` there is a JMX domain, `sk.ainet.vendors` a
Maven groupId; each spells the path of a real package directory.

A code name now never binds a folder whose first code definition is in
a Java, Kotlin, Scala or Groovy file. A Python package is what its
dotted name says and keeps binding. Like the four rules before it, this
one only takes links away.

Proof: development corpora: fastapi, dask and django unchanged;
junit-framework loses 2 links to package folders (the docs name
`org.junit.jupiter.api.condition` as a package, so those two were
right: a recall cost of the rule). Judged held-out replays: Markdown
and reST code names keep all 195 links judged correct and lose none;
the failed AsciiDoc census loses its 21 judged-false folder links and
its 3 judged-correct ones. The AsciiDoc family stays off until a new
census.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant