Repository navigation
Conversation
…ns they name A Markdown document is documentation in itself: its sections now link to the code they reference explicitly, as MENTIONS edges (via "markdown") from the Section node (the File node for text before the first heading, and for a heading whose title repeats in the file, whose Section node it does not own). An agent can go from a document to the code and back through the graph: trace_path inbound with edge_types ["MENTIONS"] lists the sections that name a definition. Families (the "syntax" property; each with its ship gate): - link: the destination of an inline link or a link reference definition (no URL, no in-page anchor, no image); - path: a bare path in prose (a slash and a file extension); - code_path: a code span or HTML <code> span holding a path, a file name or a directory, with an optional line range (#L3-L9, :3-9) or member (path::Name); - code_name: a code span holding a qualified name (a.b, a::b, A#b, A\B, mod:attr). Resolution is guess-free, with the rules the documentation field tests measured: - a path is read relative to the document and, when it has a directory part, to the repository root; a file name alone only next to the document; - a line range binds the innermost code definition holding the whole range, else the file with "target_lines" on the edge; path::member the one definition of that name in the file; - a qualified name binds the one code definition whose qualifier chain ends with the written qualifier, else the one module or package so named; a match as written outranks one that needs case folding; a field binds only through its type, never through an instance; module:attr only in Python; product documentation never binds test code; - a bare name is never an edge. Unresolved references are rows in doc_link_unresolved: missing (the directory or the repository's own root is there, the name is not), ambiguous, test_only_target. A URL, an image, a URL route (/docs), a file of a type the indexer does not read, and code of other projects are no references to the repository and leave no row. Extraction scans the document's bytes once (doclink_md.c); every scan of a line is linear, or n log n where it sorts, using position lists that grow with the bytes they track, not with the line. Front matter, fenced code, HTML comments and MDX import/export lines are not scanned. The incremental route's node proxies now carry their line numbers (pipeline_delta.c), so a re-resolved document binds a line range exactly as a full build does; the Markdown resolver reads definitions from its own by-file list of the run's nodes, never from DEFINES edges, which an incremental run holds for re-extracted files only. Measured on four repositories against the main branch (which links file to file only, REFERENCES_FILE): fastapi 2,899 MENTIONS edges (2,225 REFERENCES_FILE before), cosmos-sdk 1,286 with 205 distinct definitions reached and 188 edges from 41 ADRs into code, dapr 184, curl 52. Link pairs agree with the field-test harvester (cosmos-sdk 408 of 409; the difference is a link of a document to itself, skipped on purpose). A hand check of 45 sampled qualified-name links found one wrong, fixed by the exact-case rule above. Tests: doc_links_md (extraction, repeated headings, the span classifier table, resolution with every reason, back navigation, incremental == full for a code edit, a document edit and an added name). Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
… facts and SUPERSEDES
A Markdown file that is an architecture decision record now also defines an
ADR node (File DEFINES ADR): its name is the record's canonical id ("ADR-12";
another prefix as written, "DEC-5"; a dated record keeps its file stem), its
qualified name "<module>.__adr__", and its properties carry the record's
facts: adr_id, title, status (normalized: proposed, accepted, implemented,
superseded, deprecated, rejected, abandoned, draft, unknown), status_text,
date, deciders, superseded_by, and the decision itself as the docstring, so
search finds a record by what it decided. Its sections link to code like any
document's (MENTIONS, previous commit).
Detection is per file, with the field tests' rules: a file in a conventional
ADR directory (adr, adrs, decisions, decision-records, ...) whose name carries
an id or whose text has ADR structure; any adr-NNN file; a numbered or dated
file with ADR structure. Never README, index or template files without an id,
never test fixtures. Facts come from front matter, then header fields
(bullets, bold, tables), then the Status section; the date also from a
"Last updated" field, the change log's first date and the file name.
An ADR's own statement that it supersedes another ("Supersedes [ADR-3](...)",
"Supersedes ADR-3"; "Replaces ADR-3" where it opens a line, never "replaces"
as a verb of prose) becomes a SUPERSEDES edge between the two ADR nodes, by
link or by id. "Superseded by" is written in the old record about another
file; an edge from the other node would be owned by a file whose text does
not hold it, so it stays a fact of the old record (status, superseded_by).
The link layer gains a per-family edge type (doclink.h: the family list's
"edge" column), so a family can produce SUPERSEDES instead of MENTIONS; edges
are grouped by (type, source, target). Definitions gain a tail field for a
document's facts (extra_props: key, value pairs), carried through result
compaction and written by both property builders.
Measured against the field-test harvester on five repositories: detection
182 of 182 records, the same files; title 182/182; status 179/182 (cosmos
ADR-050 and its annexes say "Status: ARCHIVED", kept as status_text with
status unknown); date 174/182 (semantic-kernel's "date: {2024-05-02}" is read
as a date where the harvester dropped it). SUPERSEDES: semantic-kernel
0038 -> 0015, log4brains 20201016 -> 20200926, each checked against its
sentence. Markdown MENTIONS unchanged on fastapi, curl, dapr, cosmos-sdk.
Not covered: ADRs in reStructuredText, and ADR-log configuration files
(.adr-dir, log4brains) as a detection source.
Tests: doc_links_md_adr (detection and exclusions, facts from front matter
and from a Status section, SUPERSEDES by link and by id, no edge from prose,
the record's MENTIONS, incremental == full for a status edit and for an
added record named by id).
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…names
A PDF is now indexed: its text layer is read by a dependency-free extractor
(internal/cbm/pdf, zlib only) and every page becomes a Section node ("page N",
the page's text as its docstring, property page), so search finds what a page
says and get_code_snippet shows a page's text, never the file's bytes. The code
a page names structurally becomes MENTIONS edges from that page, via "pdf":
pdf_path a repository path ("pkg/server/handler.go", "./x/y")
pdf_file a file name with a code extension ("models.py", "Pool.sol")
pdf_qn a qualified name ("app.models.User.save", "a::b", "x->y")
Bare identifiers never become edges (field test E6c: a type-target rule failed
its pre-registered bar), and only EXACT or UNIQUE resolutions link: the field
test's tiers (path-exact, component-aligned path-suffix, qn-exact, qn-suffix,
the latter also through a node's owner -- a Go method's QN has no receiver --
and with the same one-entity collapses). A mention split across a line break is
proposed as one reference and replaces its fragments only when it resolves.
Hygiene as measured: no links from PDFs under test or fixture directories (R1),
none to data files (R2), none for keyword arguments (R3); a test, mock or
fixture target is a test_only_target row (R4, with the test directories the
field test found missing). A path or name whose first component is not of this
repository is no row at all.
The extractor is a port of the field-tested prototype: xref tables and
streams, /Prev chains, repair near a wrong offset and reconstruction from the
object headers, object streams, Flate (with partial recovery), LZW, ASCII85,
ASCIIHex, RunLength and PNG/TIFF predictors, simple fonts (base encodings,
/Differences, built-in Type1 encodings, standard-14 widths, Type3), ToUnicode
and encoding CMaps, Type0/CID fonts with the embedded TrueType cmap read
backwards, Form XObjects, inline images skipped, word and line breaks from
glyph geometry, right-to-left runs reversed. Its geometry is bit-exact on every
platform (no FMA contraction; CPython's hypot). Untrusted input: every read is
bounds-checked, recursion is depth-limited, integers from the file saturate, a
decoded stream is cut at 512 MiB, and every lookup that could rescan the file
per object or per glyph is answered from an index built once. Not read:
encrypted documents (reported as such); scanned pages have no text layer.
The mention scanner is the field test's (idents.py) with CPython's semantics:
NFKC per code point plus canonical composition and the str predicates, from
generated Unicode 16.0 tables (pdf_unicode.c), and re's leftmost-first
matching.
Pipeline: CBM_LANG_PDF (no grammar; the registry ledger gets a DOCUMENT
capability, the call-node manifest exempts it like the other grammar-free
classification), extraction before any grammar lookup, a resolver hook
file_prepare (a file's state from all its references before any resolves).
Incremental: a proxy node carries parent_class where owner + name differs
from its QN (Go methods, class-body variables), so a re-resolved document binds
a Go method as a full build does.
Proof against the field-test prototype on its corpus (749 PDFs from 21
repositories, restored at the recorded commits, sha256-verified): text
byte-identical on all 1,492 pages of the 720 comparable PDFs (0.73 s vs 14.7 s);
mention tokens identical on all 960 page-text sets of E6, E6b and E6c (26,435
tokens); links identical to the frozen Python pipeline run over the same graph
on semantic-kernel 5/5, cosmos-sdk 30/30, Polly 11/11, mediastore3 3/3, dask
4/4. 44,070 seeded mutants (truncations, flips, nesting and keyword bombs, on
the originals and on uncompressed copies) under ASan and UBSan: 0 findings
after the first campaign's integer overflows were fixed as a class.
Tests: suite doc_links_pdf (layout, fonts and CMaps, object and xref streams,
xref repair and reconstruction, Form XObjects incl. a self-referencing one,
inline images, filters, hostile input, the scanner against the field test's
output, the pipeline's tiers, hygiene and rows, back navigation, the snippet,
incremental == full).
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…arkup names
A reST document now gets a Section per title (the text up to the next title
as its docstring, its lines from the title to the next one) and MENTIONS
edges, via "rst", from each section to the code it names:
role :class:`pkg.Model`, :meth:`~x.Y.z`, :func:`.f`, :c:func:`f`,
and extlink roles whose URL is a repository blob path
(:source:`django/db/models/base.py`)
object .. py:class:: / .. method:: / .. function:: ... and their
C-domain twins
autodoc .. automodule:: / .. autoclass:: ...
include, .. include:: / .. kernel-include:: of a code file,
literalinclude .. literalinclude:: (its :lines: bind the innermost
definition holding them, kept as target_lines; its
:pyobject: the named definition)
kernel_doc .. kernel-doc:: path and its :identifiers: / :functions:
code_path, inline literals naming a path or a qualified name (the
code_name Markdown code-span rules)
Structure is the field test's line model (H3): titles by adornment style,
literal blocks after `::` and in code directives, comments, directive
arguments and options, Sphinx's module / currentmodule / class contexts,
CPython's str whitespace and \w. Text before the first title belongs to the
file. A Sphinx documentation set's conf.py is read as text, never run: its
source_suffix makes discovery index e.g. Django's `.txt` documents as reST in
that set only, its primary_domain decides undomained roles, its extlinks and
intersphinx names are known to the resolver.
Python names resolve the way Sphinx imports them: the longest prefix that is
a module's import name (named from its topmost package directory), the rest
by qualified name, a package's re-exports (`from .base import Model` in its
__init__.py) followed. The re-exports and the conf.py settings travel as
Python scope blobs (doclink_py.c), so an incremental run resolves as a full
one; one resolver serves reST and Python. Forms in Sphinx's order (T, C.T,
M.T, M.C.T; a leading `.` reverses it). EXACT when a form names a node,
otherwise UNIQUE by name with the field test's fixes: a bare member
(:meth:`save`, :attr:`x`) is never searched (37 of 40 such results were
wrong), an object directive's target lies in its declared module, a module
never answers a member's role (`:meth:`dispatch`` is not tests/dispatch),
test code is no target of a search but binds when the document writes its
name in full (`django.test.TestCase`). C identifiers are whole names: the one
C/C++ definition of that name and kind. Unresolved: missing when the name or
path is the repository's, ambiguous, test_only_target; another project's
name (stdlib, intersphinx) is no row.
reST architecture decision records: the ADR detector reads reST records
like Markdown ones (titles underlined with = or -): course-discovery's 34
decision records become ADR nodes.
Proof on Django's documentation (the field test's corpus and commit), C
against the frozen H3 harvester over the same graph, per (document,
section, target): 6,695 of H3's 7,238 links identical; H3-only 543 = links
C attributes to the document because a section title repeats in it (the
Markdown rule; the same document -> code link exists), deliberate refusals
(bare members, directive searches without a module) and H3's own audited
error class (router protocol methods -> ConnectionRouter, forms Field ->
model Field); C-only = the same attribution mirrored, plus 1,053 inline
literal links (sample of 30: 30 correct). 6,574 reST sections; 12 s for the
whole index.
Tests: suite doc_links_rst (structure and tokens, the scope blobs, the
pipeline's rules, source_suffix in discovery, reST ADRs, incremental ==
full). Revert check: without the module-kind guard and without the
source_suffix mapping exactly their assertions fail (observed). Two
expectations of the old behaviour now name the new one: the reST grammar
golden and probe have the document's Sections (they recorded "reST titles
are no Sections", the gap this closes).
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
… edit The closure route treats a body-only edit (the file's surface, which holds no line numbers, unchanged) as "the file re-resolves, nobody else does". That holds for references that bind names, but a document can bind lines: Markdown `app/routing.py#L8-L9` and reST `.. literalinclude:: :lines:` bind the innermost definition holding those lines. A body edit that moves code left such a document on its old target on an incremental run, while a full index bound the new one (observed: the doc_links_md and doc_links_rst "lines moved" steps, incremental != full). The documents (Markdown, reST: cbm_doclinks_binds_lines) with edges into a body-edited file now join the repair closure, from one more dependent-files query over the body-edited files. Code files stay out: their references bind names, which a body edit keeps. The query shares the existing dependents query's block and failure path (no new allocation site). Tests: a "lines moved" step in doc_links_md_incremental and in rst_links_incremental; with the predicate disabled exactly these two steps fail (observed). doc_links_md/rst/pdf, doc_mentions, incremental, pipeline pass. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…e and name
AsciiDoc (.adoc, .asciidoc) is now indexed, without a grammar like PDF: a
Section per heading (`=` .. `======`, or `#`), the text up to the next
heading as its docstring, delimited blocks respected (listing, literal,
comment, passthrough, example, sidebar, quote, table, fences). Its
references become MENTIONS edges, via "asciidoc", from the section they are
written in:
include include::target[...] -- Asciidoctor's preprocessor directive,
read in every block but a comment block; the page's own
attributes, then the Antora component's and the playbook's,
substituted
attribute an {attribute} whose value is an Antora javadoc link
({javadoc-root}/<module>/<package path>/<Type>.html[#member])
binds the Java or Kotlin type, or its one member, that the
repository's main sources declare under that path
code_path, monospace spans naming a path or a qualified name (the
code_name Markdown code-span rules)
The rules are the field test's (H8 Asciidoctor/Antora harvester, audited
~100 % precise): an Antora resource ID ([component:][module:]family$path)
names <component>/modules/<module>/<family dir>/<path>, or the file a
collector scan contributes there (a scan of build output is generated: no
row); any other target is relative to the including file. `lines=` and
`tag=`/`tags=` regions bind by H8's rule: edge blank and comment lines
trimmed, the innermost enclosing definition, unless that is a type (or there
is none) and the definitions inside cover every other line (an annotation
right above one counts) -- then that definition; tag regions come from the
included file's own `tag::x[]` / `end::x[]` comments, read from that indexed
file; a tag the file lacks binds the file. One reference makes one edge: a
region holding several top-level definitions binds the file.
Antora settings travel as YAML scope blobs (antora.yml: component name,
asciidoc attributes, collector scans; antora-playbook*.yml: attributes),
read as text, so an incremental run resolves as a full one; a change to them
is GLOBAL. AsciiDoc joins the documents bound by lines: a body edit of an
included file re-resolves its pages.
Also: the Markdown resolver's module index no longer takes AsciiDoc or PDF
files for code modules.
Proof on junit-team/junit-framework at 6c8ca116 (H8's corpus and commit),
C against H8's recorded links: javadoc attributes 362/362 identical;
includes 292 identical, 60 bind the type enclosing the several definitions
H8 names, 7 regions with several top-level definitions bind the file. 317
sections, 668 edges, no include rows (9 generated partials, 1 file type the
indexer does not read); monospace spans 30/30 correct by hand. 8.3 s for the
whole index.
Tests: suite doc_links_adoc (structure and tokens, the Antora blob, the
pipeline: resource IDs through a collector scan, tag regions, a region with
its imports binding its class, lines=, javadoc attributes, spans; incremental
== full: a moved tag region, a changed component attribute). Revert check:
without H8's region rule and without AsciiDoc among the line-bound documents
exactly their assertions fail (observed).
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…het measures With the line-bound documents' re-resolution sharing the dependents block of an incremental run (no raw allocation site of its own), the file holds 143 raw allocator sites; the memory-core baseline follows (a file may only go down). Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The PDF text layer carries data tables (D3). Their sources and notices are
now in THIRD_PARTY.md (shipped in every archive through
THIRD_PARTY_NOTICES.md) and in the SBOM:
- the Adobe Glyph List subset (BSD-3-Clause, Copyright 2002-2019 Adobe),
- the printable-ASCII advance widths of Helvetica and Times-Roman from
Adobe's core-14 font metrics (the AFM files are not distributed),
- the Unicode Character Database data of pdf_unicode.c (Unicode
License v3, full text included).
scripts/gen-pdf-unicode.py (stdlib CPython only) regenerates
pdf_unicode.c byte for byte with CPython 3.14.6 (Unicode 16.0.0);
pdf_tables.c is maintained data from here on. License texts were taken
from their sources (adobe-type-tools/agl-aglfn, unicode.org).
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Markdown sections linked to code, ADR nodes, reStructuredText and AsciiDoc sections, PDF pages and the files newly discovered for them change the graph of documents that did not change. An index written before (version 5, the doc-comment layer) keeps those documents as they were, so it is rebuilt in full once. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Loading a stream whose /Length referred to another stream loaded that stream first, and its /Length the next one: one nested call per object. Loading an object of an object stream whose /Length (or filter) lay in another object stream opened that one inside the first, again one call per stream. A crafted chain overflowed the stack (D1 of the document layer security review). A /Length is now read as a value only: an object that is a stream there gives no length (the stream is then read up to its endstream, as for any wrong length) and is marked, so a second reference costs nothing. The objects that open an object stream are no objects of an object stream (ISO 32000-1, 7.5.7); one that is, is refused and counted (objstm_nested). Both rules leave the nesting constant; nothing is cut by a limit. Test: pdf_extract_object_chains extracts a 20,000-deep /Length chain and a 3,000-deep object-stream chain on a 256 KiB thread stack (stack overflow before; each half fails alone when its rule is reverted). The 749-PDF dev corpus extracts byte-identically to the prototype as before (720 of 720 documents, 1,492 of 1,492 pages). Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Every classic section that named an /XRefStm parsed and decoded that stream afresh, and each decoded copy was kept until the document closed: thousands of tiny sections naming one highly compressible stream asked for thousands of copies (D3 of the document layer security review). A target already merged is skipped: a later (older) section naming it again adds nothing, since every number it holds is in the table already. A cross-reference stream's decoded bytes are freed once its rows are read (a stream parsed there is no cached object), so distinct streams cost one at a time. The filter pipeline of pdf_stream_decode is split out for that (decode_filters); cached streams decode as before. Test: pdf_extract_xrefstm_once: a thousand sections naming one stream decode it once (xref_streams == 1; 1,000 with the skip reverted). The 749-PDF dev corpus extracts byte-identically as before (720 of 720 documents, 1,492 of 1,492 pages). Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Objects whose value began a token that never ended (a string without its closing parenthesis) were each read to the end of the file, and the cross-reference rebuild reads every object: n such objects cost n times the file in time and in arena memory (D4 of the document layer security review). The same held for scanned trailers and for the members of an object stream. The objects of a file do not overlap, so an indirect object's value is now read up to the next object header (the header scan, built once per document), a scanned trailer up to the next "trailer", and an object stream member up to the next member. A stream's bytes are still placed by its /Length. The total work is linear in the file; nothing is cut that a well-formed file holds. Test: pdf_extract_open_strings_linear: a thousand objects with open strings and no cross-reference grow the arena and extraction classes by at most 16 x the file + 1 MiB (67 MB for a 112 KB file with the window reverted). The 749-PDF dev corpus extracts byte-identically as before (720 of 720 documents, 1,492 of 1,492 pages), and its 720 flattened copies give the same output before and after. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
A Form XObject that drew the next one many times, level under level, was read (draws ^ depth) times: a few kilobytes asked for an unbounded reading, and the page text grew with every leaf (D2 of the document layer security review). The forms' text cannot be read faithfully; it is exponential in the file. The first reading of each form in a document stays whole. Reading one again (the same content drawn elsewhere: a logo, a header on every page) is counted in decoded bytes; beyond 256 MiB of such re-readings the document reads no form again, and each skip is counted (form_rereads_cut), as a decoded stream above 512 MiB already is. Real documents re-read small forms: the 749-PDF dev corpus extracts byte-identically as before (720 of 720 documents, 1,492 of 1,492 pages; flattened copies too). Reading each form only once per page was tried first and changed one real page, so it was not taken. Test: pdf_extract_form_ladder_bounded: seven levels of forms drawing the next four times each, 64 KiB per form: re-readings are cut and the page keeps its glyphs bounded (with the bound reverted nothing is cut). Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
For each line break that joins a name, the scanner listed the page's candidates overlapping the joined span by reading every candidate of the page: (joins x candidates), quadratic in a page's text, and the geometry deciding the line breaks is the file's (D5 of the document layer security review). A join's span covers the end of one line and the start of the next, so the emitted candidates are indexed by line once per page and a join merges the two lines' lists: the same candidates in the same order, so the same tokens. Test: pdf_scan_join_linear: a page twice as long costs at most 2.5 x the join work (seam counter; 717,600 vs 2,875,200 with the full scan restored). The scanner's expected-token test is unchanged, and the PDF links on three dev corpora (Polly, mediastore3, dask) equal the frozen field-test pipeline's as before (0 links only in C, 0 only in Python). Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
A PDF page's Section carried the whole page text as its docstring, which is also what get_code_snippet returns for it: a crafted text layer stored megabytes per node and handed them to an agent whole (D6 of the document layer security review). Markdown, reStructuredText and AsciiDoc sections keep a short docstring already. A page's docstring is now its text up to 256 KiB, cut on a character boundary and followed by a line that says where it was cut and how long the text is. Pages hold a few kilobytes; the mentions are still read from the whole text. Test: pdf_page_doc_bounded: a 315 KB page keeps 256 KiB plus the note, and a path written at its end is still a mention (the whole text is stored with the bound reverted). Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The pre-registered held-out check of the PDF families: the C pipeline was run on the eighteen repositories of the field test's two held-out PDF audits (its links equal the frozen prototype pipeline's on the same graph: 5,217 of 5,217, none only in one), and the audited prototype links it makes take their labels (255 of 260 audited links). Bar per family: Wilson 95 % lower bound >= 0.90 and point >= 0.95. pdf_path 61 of 61 correct point 1.000 low 0.941 ships pdf_file 110 of 114 point 0.965 low 0.913 ships pdf_qn 77 of 80 point 0.963 low 0.896 does not ship pdf_qn misses the bar: its references now become below_bar_tier rows (the family's ships flag), as the pre-registration says, with no rule tuned on the held-out data. Its three errors, all in one repository, are names of an external API matched to a same-named local definition. The pipeline test now asserts the rows, and back navigation through the file a page names. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The pre-registered held-out audit of the document families: nine
repositories chosen by an objective rule and disjoint from every
development corpus (three each for Markdown, reStructuredText and
AsciiDoc), 782 sampled links judged blind against the source, a second
reader on every link not judged correct and on a seeded tenth of the
rest (102; agreement 100 %). Bar per family: Wilson 95 % lower bound
>= 0.90 and point >= 0.95.
ships markdown link 100/100, path 50/50 (census), code_name 98/100;
rst role 100/100, autodoc 98/98 (census), code_name 99/100;
asciidoc include 92/92 (census)
rows markdown code_path 93/100 (a file name in a code span that
means the reader's own file, not the repository's);
asciidoc code_name 15/25 (Java overloads, a config key and a
package read as members, a binary name read as a manifest);
asciidoc code_path 12/13; rst code_path 3/3 and supersedes 1/1
(censuses too small for the bar)
rows not audited, no held-out occurrences: rst object, include,
literalinclude, kernel_doc; asciidoc attribute
A family that is not shipped still resolves; a resolved reference is a
doc_link_unresolved row with reason below_bar_tier. The suites of these
families test their resolution with the families forced on (the test
seam), and doc_links_md_ship_gate tests the production gate (red with
one held family switched back on).
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The Linux arm64 production build (gcc -O2 -Werror) stopped on maybe-uninitialized in the TrueType cmap readers and the cross-reference stream rows: ttf_fmt4 and ttf_fmt12 ignored the results of rd16/rd32 after a bounds check that makes every read succeed, and xref_rows read pairs[] that only the default /Index branch fills. The ASan test build did not warn. The reads now fail the map on a failed read (the prototype's short-slice rule; never taken, the bounds check stands before them), and pairs[] starts at zero where it is declared. Text output is unchanged. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…milies ship Supplementary held-out audits, each registered before its queries and run on repositories chosen by an objective code-search walk, disjoint from every earlier audit and development corpus (blind readers, a second reader on every link not judged correct and on a seeded tenth of the rest; bar: Wilson 95 % lower bound >= 0.90 and point >= 0.95). A fix is judged on a NEW sample, never on the one that found the error. Ship now: - Markdown code_path, refined to code spans with two or more path segments: 97 of 100 (lower bound 0.916). A file or one directory named alone in a code span is its own family, bare_path, held: across the judged samples it was 28 of 41 correct (`config.yaml` the reader's own file, `doc/` another project's directory). A span without a word character (`../`, `./`) is no path at all. - reST literalinclude: 100 of 100; an open range (`:lines: 5-`, `-5`) now reads the included file for its end and binds like a closed one. - reST code_path: 99 of 100; its one error, a bare `../`, is fixed above. - PDF qualified names: 97 of 99 after the owner rule: a definition whose qualified name lacks its owner (a Go method under its package) matches only owner-qualified text, so `time.Now` no longer binds a receiver's Now. Fixed, still held: - reST object directives: a function or class directive never binds a Module, Folder or File (an example script's module took them); a new sample is due. - ADR supersedes: three censuses failed (37/38, 144/163, 82/87). Fixed from them: a combined template label (`Supersedes / Depends on:`), a value of "none" or "n/a" followed by records it extends, another record as the subject (`ADR 0239 supersedes ...`, a table cell `0259 (supersedes ...)`), a full stop after a number ending the target sentence while a `[text](destination)` link stays one reference (adr-tools' `Supersedes [3. Title](0003-...md)`). The family is split: a statement that opens its line or field and names a record directly stays `supersedes` (44 of 45 on the last census); the same words in running prose are `supersedes_prose` (38 of 42), held. Both still held in this commit; a fresh census of statements passed afterwards and is acted on next. - An ADR log in a repository without code definitions lost every supersedes link by file: the ADR table was filled after an early return. It is now filled first. Every fix has a test that fails with the fix reverted. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
A fresh held-out census of ADR supersedes statements (the phrase opens its line or field and names a record directly) passed: 80 of 81 correct (point 0.988, Wilson 95 % lower bound 0.933; 27 ADR repositories from the same objective code-search walk, at most 25 links per repository, blind readers, second-reader agreement 1.0). The family ships; the same words in running prose (`supersedes_prose`) stay rows. Its one error was a paragraph wrapped before the phrase: `[ADR-0240](0240-x.md)` / `supersedes ADR-0134 and ...`, the subject on the line before. A statement may no longer continue a paragraph: a line that no list, quote or table mark opens, after a line of running text (not blank, a heading or a `Key: value` field), is prose unless the phrase is itself a field (`Supersedes:`). On the census this moves exactly that link to prose; the 80 correct ones stay statements and none is added. Tests cover the wrapped paragraph, consecutive field lines and a field under running text, each failing with its rule reverted; the ship gate checks a statement edge and a prose row in production. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Java, C# and Solidity list the same node kind as a field and as a variable, so every class field was emitted twice: as a Field of its class and as a Variable carrying the module's qualified name (`demo.query.filters` for field Query.filters). The Variable collided with package paths, and usages bound to it instead of the Field. extract_class_variables now skips a declaration that extract_class_fields already made a Field (walked alongside the body's Fields, in order). Fields with an initializer were named by the whole declarator text (`PROP = "v"`), so the duplicate Variable had been the only findable node for them; resolve_field_name_node now takes a variable_declarator's own name. Measured on junit-framework: Variables 2,899 -> 685; of 9,901 usages into the removed duplicates, 6,938 now reach the Field. Swift and GDScript (Variables only) and D/Dart (no duplicates) are unchanged. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
A code span naming `OliveTin.exe` was read as member `exe` of the module stem of OliveTin.exe.manifest. Extensions of built and shipped files (exe, msi, deb, rpm, dmg, tar, xz, bz2, 7z, apk, ipa, manifest, csv, tsv) now make a dotted token a file name; none of them is a common member name. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Sphinx object directives (`.. py:class::`, `.. method::`, `.. function::`, ...) bind the definition they document. The family was held back until its held-out audit: a census of 105 links in repositories never used for development, judged blind in two rounds, found all 105 correct (Wilson 95% lower bound 0.965). On django's documentation (a development corpus) the 2,046 held rows become 1,895 edges: 1,640 new section-to-definition pairs, and 255 pairs that a role or code span already linked now carry the directive. No other edge or row changes. include and kernel-doc stay held. rst_ship_gate tests the production gate itself, after the suite's forced families are reset. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…field` A Java class may declare a field and a method of one name (`executable` and `executable(Path)`), and the graph keeps one node per qualified name, so one of the two was lost: the field in 90 and the method in 17 of the 107 such pairs in junit-framework. With fields named correctly (the previous commit) every initialized field joined these collisions. When the same class body declares a method of that name, the field now takes the qualified name `<Owner>.<name>#field`, the `base#suffix` fence of `#macro` and the Rust cfg twins; every other field keeps its name. The registry already files a fenced symbol under its plain name, so name lookups still reach it. A read, a write or a plain value reference that resolves to such a method goes to its field twin instead (cbm_registry_value_target); calls and callable references keep the method. junit-framework (development corpus), before -> after: Field 2,808 -> 2,898, Method 14,124 -> 14,141 (the 107 pairs, both nodes kept); WRITES into these methods 66 -> 0 and into their fields 1 -> 96; USAGE into these methods 97 -> 0 and into their fields 2 -> 212 (accessors reading their own field were self-edges before); CALLS into the methods 1,158 + 24 (then landing on the field) -> 1,215. A C# partial class split across files can still collide: the check sees one class body. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
AsciiDoc code names (resolved like Markdown's) failed their held-out census at 43 of 80. Four of its error classes have a rule that holds for every format: - a name written from the root `std` never binds a C/C++ definition or header of the repository: `std::complex` is the standard library's, not Boost.Typeof's `typeof/std/complex.hpp`; - a name written from a JVM root package (`com`, `org`, `java`, ...) binds a JVM definition only where it starts the package (after a source-root piece such as `src/main/java`), never as the tail of a relocated copy (`...hbase.shaded.com.google.protobuf.Message`); - a capitalized type of a Java, Kotlin, Scala, Groovy or C# definition matches in its own case only: `debug.log` is a log file, not the method `log` of class `Debug`; - a definition in another copy of the document's project never answers the document: below the directory where the two paths part, the document's own branch holds the same rest of the path, at least two directories deep, as a file with definitions (snapshots side by side; a shared `__init__.py` is no copy). Each rule only takes a link away: a refused candidate still counts, so two candidates stay ambiguous, and it is never the one linked. Proof: six development corpora (junit-framework, fastapi, dask, pydantic, curl, django) keep all 1,730 code_name links. Replayed on the judged held-out samples: Markdown and reST code names keep all 195 links judged correct and add none; on the failed AsciiDoc census the rules remove judged-false links only (7 in the vendored-Boost repository, which loses 127 code_name links and gains none). The first version of the copy rule dropped one (a package's `__init__.py` taken for a copy) and turned ambiguity between copies into new links; both are fixed here. The AsciiDoc family stays off until a new census. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
A second held-out census of AsciiDoc code names (61 of 83) failed on one class: 21 of its 22 errors bound a Java package folder to a dotted name that is no code at all. Debezium's `snapshot.mode` and `debezium.sink.type` are configuration keys, OpenNMS's `org.opennms.features.amqp.eventreceiver` is an OSGi configuration id, `org.apache.cassandra.auth` there is a JMX domain, `sk.ainet.vendors` a Maven groupId; each spells the path of a real package directory. A code name now never binds a folder whose first code definition is in a Java, Kotlin, Scala or Groovy file. A Python package is what its dotted name says and keeps binding. Like the four rules before it, this one only takes links away. Proof: development corpora: fastapi, dask and django unchanged; junit-framework loses 2 links to package folders (the docs name `org.junit.jupiter.api.condition` as a package, so those two were right: a recall cost of the rule). Judged held-out replays: Markdown and reST code names keep all 195 links judged correct and lose none; the failed AsciiDoc census loses its 21 judged-false folder links and its 3 judged-correct ones. The AsciiDoc family stays off until a new census. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Documentation becomes part of the graph, linked by its specific contents to the code it names. Markdown (and MDX), architecture decision records, reStructuredText (Sphinx), AsciiDoc (Antora) and PDF documents get a node per section or page, and each section links to the code files and, where the reference names or covers one, to the exact definition (function, method, class, ...). An agent can go from a section to the code it describes, and from any code node back to the documents that mention it, through the graph alone.
All of it is decided by the indexer, deterministically, from explicit references: links, paths, qualified names in code spans, Sphinx roles and directives, include directives with line ranges and tags, Antora resources and javadoc attributes, and structural mentions in PDF text. Bare names never become edges. A reference that does not resolve with certainty is recorded with a reason instead.
What you get
Nodes.
Sectionper heading (Markdown, reST, AsciiDoc), with the section's prose as its docstring. Text before the first heading belongs to the document'sFilenode.Sectionper PDF page, with the page's text as its docstring (get_code_snippetreturns the page text, never PDF bytes).ADRper architecture decision record (Markdown or reST), withid,status,date,decidersandtitleproperties. A record's "supersedes" statement resolves to the other record; a statement (the phrase opens its line or field and names a record) becomes aSUPERSEDESedge; the same words in running prose stay rows..adoc,.asciidoc) and PDF files are discovered and indexed; Sphinx projects' extrasource_suffixentries (e.g. Django's.txt) are read as reST.Edges.
MENTIONSfrom a section (or the document) to aFileor a definition, with:{"via":"markdown","syntax":"code_name","tier":"exact","line":42,"count":1,"target_lines":[10,24]}via:markdown,rst,asciidoc,pdf.syntax: the family (below).line: where the reference is written.target_lines: present when the reference names a line range (an include withlines=/tag=,literalinclude,#L10-L24).link,path,code_name,code_path,supersedes(an ADR's statement opening its line or field:SUPERSEDESedges)bare_path(a file or one directory named alone in a code span);supersedes_prose(the same words in running prose)role,autodoc,code_name,literalinclude,code_path,object(.. py:class::,.. method::, ...)include,kernel_docincludeattribute(Antora javadoc/xref resources),code_path,code_namepdf_path,pdf_file,pdf_qnUnresolved references go to
doc_link_unresolvedwith the layer's reasons (missing,ambiguous,external,test_only_target,not_indexed,graph_gap,unparseable) and are counted inindex_status.Navigation.
trace_pathinbound withedge_types:["MENTIONS"]from a symbol lists the sections that mention it;query_graphreads any of it. No new tool.How references are resolved
/xis relative to the source dir; Antora: component/module/family resource IDs fromantora.ymland the playbook).module:attronly in Python; never product documentation to a test target unless the path names it. Five scope rules, each of which can only take a link away (a refused candidate still makes two candidates ambiguous): a name written fromstdis the standard library's, never a C/C++ definition or header here; a name written from a JVM root package (com.,org., ...) matches only where it starts the package, never a relocated copy's tail; a Java/Kotlin/Scala/Groovy/C# type matches in its own case only (debug.logis no member ofDebug); a definition in another copy of the document's project (snapshots side by side) never answers the document; a code name never binds a JVM package folder (a dotted config key or cross-reference such assnapshot.modeis not a package; Python packages keep binding).T,C.T,M.T,M.C.Tundermodule/currentmodule/class context, Python re-exports through package__init__.pyimports,primary_domain, intersphinx andextlinksfromconf.py; C domain names by kind. A one-piece name is searched only for classes, exceptions, functions and data in the declared module, never for bare members.include::withlines=andtag(s)=(tags read from the included indexed file only), attributes resolved from the document and the Antora component, javadoc attributes to Java/Kotlin types and members./ActualText, object and cross-reference streams, damaged files repaired as far as the prototype did); structural mentions only.Proof on real repositories (development corpora; the held-out list was checked first)
MENTIONS(before: file-to-fileREFERENCES_FILEonly); link pairs 408/409 equal to the field-test prototype (the one difference a deliberate self-link skip); 44/45 hand-checked name edges correct, the 45th fixed (exact-case precedence)fb1137656c8ca116Extractor fix found by the audits: one Field per class field
The AsciiDoc audit bound a Java field reference to a module-level
Variable. Java, C# and Solidity list the same node kind as a field and as a variable, so every class field was emitted twice: as aFieldof its class and as aVariablecarrying the module's qualified name (demo.query.filtersforQuery.filters), which collided with package paths and took usages away from the field. A declaration that is already aFieldis no longer minted again, and a field with an initializer is named by its declarator's name (it was namedPROP = "v"). On junit-framework,Variablenodes drop from 2,899 to 685, and 6,938 of the 9,901 usages that went to the duplicates now reach the field. Swift and GDScript (variables only) and D and Dart (no duplicates) are unchanged.A field and a method of one name in one class body (
executableandexecutable(Path)) shared one qualified name, so the graph kept only one of them (in junit-framework 107 pairs: the field lost in 90, the method in 17). Such a field is now<Owner>.<name>#field(thebase#suffixfence of#macroand the Rust cfg twins; every other field keeps its name), and a read, a write or a plain value reference that resolves to the method goes to its field twin; calls keep the method. junit-framework:Field2,808 → 2,898 andMethod14,124 → 14,141;WRITESinto those methods 66 → 0 and into their fields 1 → 96;USAGEinto those methods 97 → 0 and into their fields 2 → 212. Limit: a C# partial class split across files can still collide (the check sees one class body). Both changes ride on this PR's index-format bump (CBM_SEMANTIC_INDEX_VERSION6 over its base's 5).Security review
An independent read-only review of the layer found four HIGH, one MEDIUM and one LOW issue, all in the PDF path and all fixed here, each with a test that fails when its fix is reverted, and each re-proven on the PDF corpus (byte-identical text before and after):
/Lengthnaming a stream, object streams inside object streams) cost a stack frame per object: now constant depth (a/Lengthis read as a value; ISO 32000-1 7.5.7 for object streams);Tests
New suites
doc_links_md(Markdown and ADRs, 10 tests),doc_links_pdf(15),doc_links_rst(7),doc_links_adoc(4); incremental-equals-full steps for every format, including line-bound documents after a body edit. Each PDF security fix has a test that fails when the fix is reverted.Verification
Local ladder on this PR's tip (the doc-comment layer PR below it included):
feat/doc-links-core, notmain; only DCO runs), so the Windows check comes when the stack is retargeted tomain. The local Windows arm64 VM last ran on the previous tip (8,451 passed, 83 skipped); two route tests ofedge_types_probefail there identically on the base of this stack, so they do not come from this PR.Precision audit
Each family ships as edges only if it passes a held-out audit registered before any output existed on the audit repositories: 100 sampled links per family (census below that, families under 15 links pooled), judged blind against the source, bar Wilson 95 % lower bound ≥ 0.90 and point ≥ 0.95. Held-out repositories: three per format, chosen by an objective rule from code search, disjoint from every development corpus.
PDF (done). The C pipeline ran on the eighteen repositories of the field test's two held-out PDF audits. Its links equal the frozen prototype pipeline's on the same graph in every repository (5,217 of 5,217), and the audited prototype links it makes take their labels (255 of 260):
pdf_pathpdf_filepdf_qnbelow_bar_tierrowsMarkdown, reStructuredText, AsciiDoc (done). Nine held-out repositories (three per format, the top-starred results of a code search for the format's markers that hold an ADR log, Sphinx autodoc, or AsciiDoc source includes; none used in development): 782 sampled links judged blind, a second reader on every link not judged correct and on a seeded tenth of the rest (102 rechecks, 100 % agreement).
linkpathcode_nameroleautodoccode_nameincludecode_pathcode_namecode_pathcode_pathsupersedesobject,include,literalinclude,kernel_doc; AsciiDocattributeSupplementary audits (same method, each registered before its queries; new held-out repositories chosen by the same kind of objective code-search walk, disjoint from all earlier ones). A family that failed was fixed only from what its sample showed, and a fix is judged on a NEW sample, never on the one that found the error:
literalincludecode_path../, is fixed)pdf_qn, after the owner rule (a Go method whose name lacks its receiver matches only the receiver-qualified form)code_path, with a directory (src/app/main.go,docs/decisions/)doc/of another project)code_pathwithout the bare part (two or more path segments)bare_path(a file or one directory alone)config.yamlordoc/in prose is as often the reader's or another project'sobjectsupersedesSupersedes / Depends on:read as a statement; fixedsupersedes, after that fixSupersedes: none (it extends ADR 31), another record as the subject; fixedsupersedes, after those fixessupersedes_prose(38 of 42, held)supersedesstatementskernel_docobject, after the fix (a function or class directive never binds a module)code_path.circleci/config.ymlin CircleCI's docs: 25 of the 29 errors;~/.m2/settings.xml, "yourpom.xml"), and..as range syntaxattributecode_name, after the field and file-name fixes, overloads judged at method levelstd::name bound to Boost's specialization, a class bound to another snapshot's copy; the documentation-first Antora repositories gave 26 of 28code_name, after four scope rules (std::, rooted JVM names, type case, project copies)snapshot.mode, OSGi configuration ids, JMX domains, a Maven groupId)code_name, after a fifth rule (no JVM package folders)sections?bound tosections), a method bound to its file, andname()where a zero-argument overload exists (one node per method name)A family that does not ship still resolves: its references are
doc_link_unresolvedrows with reasonbelow_bar_tier, and its tests run it switched on, so a later audit can switch it on without other changes.Known limits
Third-party data
The PDF tables carry the Adobe Glyph List subset (BSD-3-Clause), the core-14 font widths of Helvetica and Times-Roman (Adobe), and Unicode Character Database data (Unicode License v3); notices are in
THIRD_PARTY.mdand the SBOM.scripts/gen-pdf-unicode.pyregeneratespdf_unicode.cbyte for byte.