From 44efa6281b4f6fa188ef095668e975346f66a619 Mon Sep 17 00:00:00 2001 From: Martin Vogel Date: Fri, 2 Oct 2026 06:37:36 +0200 Subject: [PATCH 1/7] fix(extract): C typedefs become nodes; macros and enumerators get their own QNs In C-family code several different entities competed for one node, and some had no node at all: - A typedef name was never a definition: `type_definition` has no `name` field, so the class path dropped it. With an anonymous `typedef struct/enum {...} X;` the fields and enumerators went with it. - A function behind an unknown leading macro (`API RetT name(...)`) was named after its return type, became a Method `RetT.name`, or was dropped, depending on the error-recovery shape. - Every bodyless `struct X` / `enum X` (a parameter, a field type, a forward declaration) minted a Class/Enum node. In one file, or in a same-stem .h/.c pair, such a reference displaced the real definition. - A `#define NAME` shared the QN of the function, type or enumerator called NAME, and the later line won: a namespace-rename macro replaced the typedef it renames, an `#else` stub macro replaced the function. - Enumerators were named `..`, although C and unscoped C++ enums put them in the enclosing scope. - A definition repeated in several `#if` branches kept one span; the others left no trace. - Braces split across `#if` branches lost whole functions, and a C keyword could surface as a function name (`if`). Extraction: - typedef names are definitions: `Type`, or `Class`/`Enum` with members for an anonymous `typedef struct/enum {...} X;`; an alias is dropped when the same file defines the tag of that name - macro-prefixed functions get their real name in all three recovery shapes, consistently for defs, call scopes, parameters and the C LSP - a bodyless tag is a reference and creates no node; a definition head left as loose ERROR tokens (`struct NAME {`) is recovered - a C keyword is never a function name - a first-branch projection (one branch per `#if` group, positions unchanged) is parsed when the raw parse lost structure wholesale, and only functions missing from the raw result are adopted from it - macro-wrapped enumerator lists yield one constant per slot Identity (semantic index version 4; an existing index is rebuilt once): - a C-preprocessor macro's QN is `.#macro` (C, C++, CUDA, GLSL, Objective-C, ISPC); name and label are unchanged. Resolvers prefer a definition and use the macro only when no definition of that name is visible (LSP target lookup, trace_path, get_code_snippet) - enumerators of C enums and of unscoped C++/Objective-C enums are `.`; `enum class` stays nested; `parent_class` names the enum - a node whose file defines its QN more than once under one label carries `variants`, the line spans of all those definitions (`#if` twins, per-platform macro redefinitions, C++ overloads) Full-text search indexes a macro's QN without the fence, so macro rows score as before. Measured on redis and curl against the parent commit: nodes 38,625 -> 40,484 and 28,180 -> 29,750; Type nodes 0 -> 573 and 0 -> 193; phantom Class nodes from references 202 and 821 removed; functions lost to a same-named macro 45 and 41 restored; curl functions with no node 11 -> 0; nodes of non-C-family files unchanged by QN. Of 174 audited Doxygen references in xxhash, 174 now reach the right entity (64 before). Index time is unchanged within noise (+2% pipeline time for 5% more nodes). Signed-off-by: Martin Vogel --- internal/cbm/cbm.c | 448 ++++++++++++++++++ internal/cbm/cbm.h | 21 + internal/cbm/extract_defs.c | 615 +++++++++++++++++++++++-- internal/cbm/helpers.c | 131 +++++- internal/cbm/helpers.h | 31 ++ internal/cbm/lsp/c_lsp.c | 28 +- internal/cbm/result_compact.c | 1 + src/graph_buffer/graph_buffer.c | 31 +- src/mcp/mcp.c | 24 + src/pipeline/lsp_resolve.h | 47 +- src/pipeline/pass_definitions.c | 40 +- src/pipeline/pass_parallel.c | 38 +- src/pipeline/pipeline_internal.h | 10 +- src/pipeline/registry.c | 4 +- src/store/store.c | 15 +- tests/test_extraction.c | 763 ++++++++++++++++++++++++++++++- tests/test_graph_buffer.c | 67 +++ tests/test_mcp.c | 119 +++++ tests/test_pipeline.c | 582 +++++++++++++++++++++++ tests/test_store_search.c | 62 +++ 20 files changed, 2997 insertions(+), 80 deletions(-) diff --git a/internal/cbm/cbm.c b/internal/cbm/cbm.c index 082bca1d6d..ea6124a3ad 100644 --- a/internal/cbm/cbm.c +++ b/internal/cbm/cbm.c @@ -2098,6 +2098,444 @@ static const char *cbm_error_ranges_str(CBMArena *a, const cbm_error_regions_t * return buf; } +/* ── Same-file duplicate definitions: the `variants` property ────────────── + * + * The graph keeps one node per qualified name. When one file defines a QN more + * than once under one label -- both branches of an #if, a macro redefined per + * platform, C++ overloads -- the node shows the winner's span and the other + * definitions leave no trace. The member the graph keeps gets the group's + * spans as CBMDefinition.variants (format in cbm.h), so the surviving node + * carries them. + * + * C-preprocessor languages only (cbm_is_c_preprocessor_lang, the caller's + * gate): conditional compilation is what the property describes. A repeated + * QN elsewhere is a different matter -- a JSON file repeats one key thousands + * of times -- and gets no list. + * + * Computed here from this file's defs alone: a pure function of the file, + * hence deterministic, identical on the sequential and the parallel path, and + * owned by the file under incremental re-extraction. Two defs with the same + * span are one definition seen twice and count once; a group left with a + * single span gets nothing. */ + +typedef struct { + uint64_t key; /* hash of (QN, label) */ + uint32_t start; + uint32_t end; + int idx; /* index into result->defs */ +} cbm_variant_ref_t; + +static int cbm_variant_ref_cmp(const void *pa, const void *pb) { + const cbm_variant_ref_t *a = (const cbm_variant_ref_t *)pa; + const cbm_variant_ref_t *b = (const cbm_variant_ref_t *)pb; + if (a->key != b->key) { + return a->key < b->key ? -1 : 1; + } + if (a->start != b->start) { + return a->start < b->start ? -1 : 1; + } + if (a->end != b->end) { + return a->end < b->end ? -1 : 1; + } + return (a->idx > b->idx) - (a->idx < b->idx); +} + +/* FNV-1a over the QN, a separator and the label. */ +static uint64_t cbm_variant_key(const char *qn, const char *label) { + uint64_t h = 1469598103934665603ULL; + for (const char *p = qn; *p; p++) { + h = (h ^ (unsigned char)*p) * 1099511628211ULL; + } + h = (h ^ 0xffU) * 1099511628211ULL; + for (const char *p = label; *p; p++) { + h = (h ^ (unsigned char)*p) * 1099511628211ULL; + } + return h; +} + +/* refs[0..n) is one (QN, label) group in span order: give every member the + * list of the group's distinct spans. */ +static void cbm_variant_assign(CBMFileResult *result, const cbm_variant_ref_t *refs, int n) { + enum { VARIANT_ENTRY_MAX = 40 }; /* ,{"start":4294967295,"end":4294967295} */ + int distinct = 0; + for (int i = 0; i < n; i++) { + if (i == 0 || refs[i].start != refs[i - 1].start || refs[i].end != refs[i - 1].end) { + distinct++; + } + } + if (distinct < 2) { + return; + } + size_t cap = (size_t)distinct * VARIANT_ENTRY_MAX + 3; + char *json = (char *)cbm_arena_alloc(&result->arena, cap); + if (!json) { + return; + } + size_t pos = 0; + json[pos++] = '['; + for (int i = 0; i < n; i++) { + if (i > 0 && refs[i].start == refs[i - 1].start && refs[i].end == refs[i - 1].end) { + continue; + } + int w = snprintf(json + pos, cap - pos, "%s{\"start\":%u,\"end\":%u}", pos > 1 ? "," : "", + refs[i].start, refs[i].end); + if (w <= 0 || (size_t)w >= cap - pos) { + return; /* cannot happen: cap covers the longest entry */ + } + pos += (size_t)w; + } + json[pos++] = ']'; + json[pos] = '\0'; + /* Only the definition the graph keeps for this file carries the list: the + * one with the largest start line (cbm_gbuf_upsert_node's same-file rule; + * every member starting on that line, since arrival order breaks that + * tie). Giving it to all n members would serialize an n-entry list n + * times -- quadratic on a file that defines one name very often. */ + uint32_t last_start = refs[n - 1].start; + for (int i = n - 1; i >= 0 && refs[i].start == last_start; i--) { + result->defs.items[refs[i].idx].variants = json; + } +} + +static void cbm_mark_def_variants(CBMFileResult *result) { + enum { VARIANT_STACK_REFS = 128 }; + int n = result->defs.count; + if (n < 2) { + return; + } + cbm_variant_ref_t stack[VARIANT_STACK_REFS]; + cbm_variant_ref_t *refs = stack; + if (n > VARIANT_STACK_REFS) { + refs = (cbm_variant_ref_t *)cbm_alloc(CBM_MEM_CLASS_EXTRACT, (size_t)n * sizeof(*refs)); + if (!refs) { + return; + } + } + const CBMDefinition *defs = result->defs.items; + int m = 0; + for (int i = 0; i < n; i++) { + if (defs[i].qualified_name && defs[i].label) { + refs[m].key = cbm_variant_key(defs[i].qualified_name, defs[i].label); + refs[m].start = defs[i].start_line; + refs[m].end = defs[i].end_line; + refs[m].idx = i; + m++; + } + } + qsort(refs, (size_t)m, sizeof(*refs), cbm_variant_ref_cmp); + for (int i = 0; i < m;) { + int j = i + 1; + while (j < m && refs[j].key == refs[i].key) { + j++; + } + /* [i, j) shares one key. That is one (QN, label) group unless two + * keys collided, so split it by real equality, keeping span order. */ + int lo = i; + while (j - i >= 2 && lo < j) { + const CBMDefinition *lead = &defs[refs[lo].idx]; + int hi = lo + 1; + for (int k = lo + 1; k < j; k++) { + const CBMDefinition *d = &defs[refs[k].idx]; + if (strcmp(lead->qualified_name, d->qualified_name) == 0 && + strcmp(lead->label, d->label) == 0) { + cbm_variant_ref_t moved = refs[k]; + memmove(&refs[hi + 1], &refs[hi], (size_t)(k - hi) * sizeof(*refs)); + refs[hi++] = moved; + } + } + if (hi - lo >= 2) { + cbm_variant_assign(result, &refs[lo], hi - lo); + } + lo = hi; + } + i = j; + } + if (refs != stack) { + cbm_free(CBM_MEM_CLASS_EXTRACT, refs); + } +} + +/* ── First-branch projection rescue (C/C++/CUDA) ─────────────────────────── + * + * The raw parse sees every #if branch at once. When branches split a brace + * (`#ifndef X (a)) { #else (b)) { #endif`) or a directive sits inside an + * expression, the braces stop balancing and an ERROR region swallows whole + * functions: they get no node at all (curl ossl_connect_step2). The #961 + * rescue re-parses the preprocessed source, but with no build defines a + * whole-file guard (`#ifdef USE_OPENSSL`) leaves nothing to parse. The + * projection below needs no defines: every conditional group keeps exactly one + * branch (the first; for `#if 0` the #else), so the kept text is one + * consistent configuration the author wrote. Directive lines and dropped + * branches become spaces with newlines kept, so every byte keeps its line and + * column and a def found there needs no remapping. */ + +enum { CBM_PROJ_MAX_DEPTH = 64 }; + +typedef struct { + uint8_t branch; /* index of the branch being scanned (0 = the #if part) */ + uint8_t keep; /* branch that survives: 0, or 1 for `#if 0` */ +} cbm_proj_frame_t; + +/* Directive name after '#' and optional blanks, as [*name, *name + return). */ +static int cbm_proj_directive(const char *p, const char *e, const char **name) { + while (p < e && (*p == ' ' || *p == '\t')) { + p++; + } + if (p >= e || *p != '#') { + return 0; + } + p++; + while (p < e && (*p == ' ' || *p == '\t')) { + p++; + } + const char *n = p; + while (p < e && isalpha((unsigned char)*p)) { + p++; + } + *name = n; + return (int)(p - n); +} + +static bool cbm_proj_word_is(const char *n, int len, const char *w) { + return (size_t)len == strlen(w) && strncmp(n, w, (size_t)len) == 0; +} + +/* `#if 0` (the dead-code idiom): the condition is the single token 0. */ +static bool cbm_proj_if_zero(const char *after, const char *e) { + while (after < e && (*after == ' ' || *after == '\t')) { + after++; + } + if (after >= e || *after != '0') { + return false; + } + after++; + while (after < e && (*after == ' ' || *after == '\t' || *after == '\r')) { + after++; + } + return after >= e || (after + 1 < e && after[0] == '/' && (after[1] == '/' || after[1] == '*')); +} + +/* Track block-comment state across one line, skipping string and char + * literals, so a `#if` inside a block comment is not taken for a directive. */ +static bool cbm_proj_scan_comments(const char *p, const char *e, bool in_comment) { + while (p < e) { + if (in_comment) { + if (p + 1 < e && p[0] == '*' && p[1] == '/') { + in_comment = false; + p += 2; + continue; + } + p++; + continue; + } + if (*p == '"' || *p == '\'') { + char q = *p++; + while (p < e && *p != q) { + p += (*p == '\\' && p + 1 < e) ? 2 : 1; + } + p++; + continue; + } + if (p + 1 < e && p[0] == '/' && p[1] == '/') { + return false; + } + if (p + 1 < e && p[0] == '/' && p[1] == '*') { + in_comment = true; + p += 2; + continue; + } + p++; + } + return in_comment; +} + +static void cbm_proj_blank(char *p, const char *e) { + for (; p < e; p++) { + if (*p != '\n') { + *p = ' '; + } + } +} + +/* Heap copy of `src` with the first-branch projection applied, or NULL when + * the source has no conditional group (nothing to project), nests deeper than + * CBM_PROJ_MAX_DEPTH, or allocation fails. Caller frees with cbm_free. */ +static char *cbm_first_branch_projection(const char *src, int len) { + char *out = cbm_alloc(CBM_MEM_CLASS_EXTRACT, (size_t)len + 1); + if (!out) { + return NULL; + } + memcpy(out, src, (size_t)len); + out[len] = '\0'; + cbm_proj_frame_t frames[CBM_PROJ_MAX_DEPTH]; + int depth = 0; + bool any_group = false; + bool in_comment = false; + bool continuation = false; /* previous conditional-directive line ended in '\' */ + const char *end = out + len; + for (char *line = out; line < end;) { + char *nl = memchr(line, '\n', (size_t)(end - line)); + char *eol = nl ? nl : (char *)end; + bool active = true; + for (int i = 0; i < depth && active; i++) { + active = frames[i].branch == frames[i].keep; + } + bool blank = !active || continuation; + bool was_continuation = continuation; + continuation = false; + const char *name = NULL; + int nlen = (!in_comment && !was_continuation) ? cbm_proj_directive(line, eol, &name) : 0; + bool conditional = false; + if (nlen > 0) { + if (cbm_proj_word_is(name, nlen, "if") || cbm_proj_word_is(name, nlen, "ifdef") || + cbm_proj_word_is(name, nlen, "ifndef")) { + if (depth >= CBM_PROJ_MAX_DEPTH) { + cbm_free(CBM_MEM_CLASS_EXTRACT, out); + return NULL; + } + bool zero = + cbm_proj_word_is(name, nlen, "if") && cbm_proj_if_zero(name + nlen, eol); + frames[depth].branch = 0; + frames[depth].keep = zero ? 1 : 0; + depth++; + conditional = any_group = true; + } else if (cbm_proj_word_is(name, nlen, "elif") || + cbm_proj_word_is(name, nlen, "elifdef") || + cbm_proj_word_is(name, nlen, "elifndef") || + cbm_proj_word_is(name, nlen, "else")) { + if (depth > 0 && frames[depth - 1].branch < UINT8_MAX) { + frames[depth - 1].branch++; + } + conditional = true; + } else if (cbm_proj_word_is(name, nlen, "endif")) { + if (depth > 0) { + depth--; + } + conditional = true; + } + } + if (conditional) { + blank = true; + char *last = eol; + while (last > line && (last[-1] == '\r' || last[-1] == ' ' || last[-1] == '\t')) { + last--; + } + continuation = last > line && last[-1] == '\\'; + } else if (was_continuation) { + char *last = eol; + while (last > line && (last[-1] == '\r' || last[-1] == ' ' || last[-1] == '\t')) { + last--; + } + continuation = last > line && last[-1] == '\\'; + } + in_comment = cbm_proj_scan_comments(line, eol, in_comment); + if (blank) { + cbm_proj_blank(line, eol); + } + line = nl ? nl + 1 : (char *)end; + } + if (!any_group) { + cbm_free(CBM_MEM_CLASS_EXTRACT, out); + return NULL; + } + return out; +} + +/* Only a parse that lost structure wholesale is worth a second parse: the raw + * root is itself an ERROR, or one top-most error region spans this many lines. + * Small local errors (a macro the grammar cannot read) are the common case — + * about half the kernel's C files have one — and the projection cannot fix + * them; the gate keeps the extra parse to ~1% of kernel C bytes (sampled). */ +enum { CBM_PROJ_MIN_ERROR_LINES = 20 }; + +/* Adopt from the projection only what the raw pass lost: a Function/Method + * def that overlaps a raw ERROR region, whose QN no def already holds, whose + * name is on its first raw line, and whose span has definition syntax (the + * #961 gates). Everything else the projected walk produced is discarded. */ +static void cbm_rescue_defs_from_projection(CBMExtractCtx *raw_ctx, const TSLanguage *ts_lang) { + CBMFileResult *result = raw_ctx->result; + if (!ts_node_has_error(raw_ctx->root)) { + return; + } + cbm_error_regions_t regs = {{0}, {0}, 0, 0}; + bool broken = strcmp(ts_node_type(raw_ctx->root), "ERROR") == 0; + if (broken) { + /* The whole file failed: every line is error territory. */ + regs.starts[0] = 1; + regs.ends[0] = ts_node_end_point(raw_ctx->root).row + 1; + regs.count = 1; + } else { + cbm_collect_error_regions(raw_ctx->root, ®s, raw_ctx->source, raw_ctx->source_len); + for (int r = 0; r < regs.count && !broken; r++) { + broken = regs.ends[r] >= regs.starts[r] + CBM_PROJ_MIN_ERROR_LINES; + } + } + if (!broken) { + return; + } + char *projected = cbm_first_branch_projection(raw_ctx->source, raw_ctx->source_len); + if (!projected) { + return; + } + TSParser *parser = get_thread_parser(ts_lang, raw_ctx->language); + TSTree *tree = NULL; + if (parser) { + ts_parser_reset(parser); + CBMStringInput input = {projected, (uint32_t)raw_ctx->source_len}; + TSInput ts_input = {&input, cbm_string_read, TSInputEncodingUTF8, NULL}; + TSParseOptions opts = {0}; + tree = ts_parser_parse_with_options(parser, NULL, ts_input, opts); + } + CBMHashTable *held = tree ? cbm_ht_create(CBM_SZ_256) : NULL; + if (held) { + for (int i = 0; i < result->defs.count; i++) { + if (result->defs.items[i].qualified_name) { + cbm_ht_set(held, result->defs.items[i].qualified_name, &result->defs.items[i]); + } + } + CBMExtractCtx pctx = { + .arena = raw_ctx->arena, + .scratch = raw_ctx->scratch, + .result = result, + .source = projected, + .source_len = raw_ctx->source_len, + .language = raw_ctx->language, + .project = raw_ctx->project, + .rel_path = raw_ctx->rel_path, + .module_qn = result->module_qn, + .root = ts_tree_root_node(tree), + }; + int before = result->defs.count; + cbm_extract_definitions_without_module(&pctx); + int w = before; + for (int i = before; i < result->defs.count; i++) { + CBMDefinition *d = &result->defs.items[i]; + bool adopt = d->name && d->qualified_name && cbm_def_label_is_callable(d->label) && + !cbm_ht_has(held, d->qualified_name); + bool overlaps = false; + for (int r = 0; adopt && r < regs.count && !overlaps; r++) { + overlaps = d->start_line <= regs.ends[r] && d->end_line >= regs.starts[r]; + } + adopt = + adopt && overlaps && + cbm_line_contains(raw_ctx->source, raw_ctx->source_len, d->start_line, d->name) && + cbm_span_contains_callable_def(raw_ctx->source, raw_ctx->source_len, d->start_line, + d->end_line, d->name); + if (adopt) { + result->defs.items[w++] = *d; + cbm_ht_set(held, result->defs.items[w - 1].qualified_name, + &result->defs.items[w - 1]); + } + } + result->defs.count = w; + cbm_ht_free(held); + } + if (tree) { + ts_tree_delete(tree); + } + cbm_free(CBM_MEM_CLASS_EXTRACT, projected); +} + /* Public entry: run the extraction and journal completion. The DONE mark on * every ordinary return (including error/timeout results) tells the crash * supervisor this file did NOT kill the worker — only a file whose S has no @@ -2798,6 +3236,11 @@ static CBMFileResult *extract_file_ex_body(const char *source, int source_len, C cbm_preprocessed_source_free(preprocessed); } atomic_fetch_add(&total_preprocess_ns, now_ns() - pp_start); + + /* Functions both the raw parse and the #961 preprocessed rescue lost + * (whole-file feature guards, directives inside expressions). Runs + * only when the raw tree has errors and the file has #if groups. */ + cbm_rescue_defs_from_projection(&ctx, ts_lang); } // Bottleneck call-context metrics. Each call is attributed to the INNERMOST @@ -2950,6 +3393,11 @@ static CBMFileResult *extract_file_ex_body(const char *source, int source_len, C } } + /* The def list is final here (raw walk, preprocessed and projection rescues). */ + if (cbm_is_c_preprocessor_lang(language)) { + cbm_mark_def_variants(result); + } + result->imports_count = result->imports.count; // Accumulate profiling counters diff --git a/internal/cbm/cbm.h b/internal/cbm/cbm.h index c0bba77923..4bb02a22a2 100644 --- a/internal/cbm/cbm.h +++ b/internal/cbm/cbm.h @@ -246,6 +246,16 @@ typedef struct { * HTTP_CALLS edge to base + path. Tail fields: zero-init stays valid. */ const char *http_client; const char *http_base_url; + /* Set when this FILE holds two or more definitions with this QN and label + * (`#if`/`#else` twins, a macro redefined per platform, overloads): the + * graph keeps one node per QN, and this lists every one of those + * definitions' line spans, the surviving one included, as a JSON array + * sorted by start line: [{"start":10,"end":14},{"start":20,"end":26}]. + * Carried by the definition the graph keeps for the file (the last by + * start line), which makes it the node's `variants` property. NULL on + * every other definition, and in every language outside + * cbm_is_c_preprocessor_lang. */ + const char *variants; } CBMDefinition; /* Argument captured from a call expression */ @@ -258,6 +268,17 @@ typedef struct { #define CBM_MAX_CALL_ARGS 8 +/* The qualified name of a C-preprocessor macro (`#define NAME ...` in C, C++, + * CUDA, Objective-C, GLSL, ISPC) is `.#macro`. C keeps macros + * apart from ordinary identifiers, so a macro and a function, type, variable or + * enumerator may legally share a name in one file; the fence gives the macro + * an identity of its own, and the plain `.` always belongs to + * the definition. `name` stays NAME and the label stays "Macro". + * + * Tie rule for anything that resolves a NAME or a plain QN to a node: the + * definition is the target, and the macro only when no definition is visible. */ +#define CBM_MACRO_QN_SUFFIX "#macro" + /* Byte offsets are meaningful only within the source buffer that produced * them. C/C++/CUDA run both raw and preprocessed extraction passes, and those * buffers can contain unrelated occurrences at the same numeric span. */ diff --git a/internal/cbm/extract_defs.c b/internal/cbm/extract_defs.c index b3b6e6acc3..388166c099 100644 --- a/internal/cbm/extract_defs.c +++ b/internal/cbm/extract_defs.c @@ -3,6 +3,7 @@ #include "helpers.h" #include "lang_specs.h" #include "foundation/constants.h" +#include "foundation/hash_table.h" #include "foundation/log.h" // cbm_log_error #include "foundation/mem_core.h" // cbm_realloc/cbm_free -- walk_defs stack #include "extract_node_stack.h" @@ -207,6 +208,9 @@ enum { RT_PAIR_SIZE = 2 }; // Forward declarations static void extract_func_def(CBMExtractCtx *ctx, TSNode node, const CBMLangSpec *spec); static void extract_class_def(CBMExtractCtx *ctx, TSNode node, const CBMLangSpec *spec); +static void emit_class_def(CBMExtractCtx *ctx, TSNode node, const CBMLangSpec *spec, + const char *kind, char *name); +static void extract_c_typedef(CBMExtractCtx *ctx, TSNode node, const CBMLangSpec *spec); static void walk_defs(CBMExtractCtx *ctx, TSNode root, const CBMLangSpec *spec, int depth_unused); static void extract_variables(CBMExtractCtx *ctx, TSNode root, const CBMLangSpec *spec); static void extract_var_names(CBMExtractCtx *ctx, TSNode node, const CBMLangSpec *spec); @@ -537,6 +541,13 @@ char *cbm_cpp_out_of_line_parent_class(CBMArena *a, TSNode node, const char *sou for (int depth = 0; depth < DECLARATOR_DEPTH_LIMIT && !ts_node_is_null(decl); depth++) { const char *dk = ts_node_type(decl); if (strcmp(dk, "qualified_identifier") == 0 || strcmp(dk, "scoped_identifier") == 0) { + if (cbm_c_qualifier_is_recovered(decl)) { + /* `API RetT name(...)` misread as `RetT::name` (MISSING "::"): + * RetT is the return type, not a class. Keep looking for a real + * qualifier on the name side. */ + decl = ts_node_child_by_field_name(decl, TS_FIELD("name")); + continue; + } qid = decl; break; } @@ -4961,6 +4972,11 @@ static TSNode find_c_params(TSNode func_node) { return params; } TSNode nested = ts_node_child_by_field_name(decl, TS_FIELD("declarator")); + /* `API RetT *name(...)` misread as `RetT::*name(...)` (MISSING "::"): + * the function declarator is on the qualifier's name side. */ + if (ts_node_is_null(nested) && cbm_c_qualifier_is_recovered(decl)) { + nested = ts_node_child_by_field_name(decl, TS_FIELD("name")); + } /* tree-sitter-cpp and tree-sitter-cuda do not assign the nested * function declarator a `declarator` field when a reference return * wraps it (`Item& operator[](int)`). That wrapper has one named child; @@ -5071,6 +5087,63 @@ static bool is_c_declarator_lang(CBMLanguage lang) { lang == CBM_LANG_SLANG || lang == CBM_LANG_OBJC; } +/* C-family `struct X` / `union X` / `enum X` / `class X` WITHOUT a body is a + * reference to the type — a forward declaration, a variable, parameter or field + * type, a sizeof operand, the aliased side of `typedef struct X Y;` — never its + * definition. Minting a def for it put a one-line Class/Enum node on the type's + * QN: when the definition shares that QN (same file, or a same-stem .h/.c pair, + * whose module QNs coincide), the later or smaller-path reference displaced the + * definition, and elsewhere it left a phantom type node in every file that + * merely mentions the type (`struct timeval` in redis-cli.c). */ +static bool is_c_tag_reference(CBMLanguage lang, TSNode node, const char *kind) { + if (!is_c_declarator_lang(lang)) { + return false; + } + if (strcmp(kind, "struct_specifier") != 0 && strcmp(kind, "union_specifier") != 0 && + strcmp(kind, "enum_specifier") != 0 && strcmp(kind, "class_specifier") != 0) { + return false; + } + return ts_node_is_null(ts_node_child_by_field_name(node, TS_FIELD("body"))); +} + +/* Languages whose plain `enum` is UNSCOPED: its enumerators are names of the + * scope that holds the enum, not of the enum (C, C++ and its CUDA dialect, + * Objective-C). `XXH_OK` of `typedef enum { XXH_OK, XXH_ERROR } XXH_errorcode;` + * is written and looked up as plain XXH_OK, so its QN is `.XXH_OK`. */ +static bool c_enum_lang(CBMLanguage lang) { + return lang == CBM_LANG_C || lang == CBM_LANG_CPP || lang == CBM_LANG_CUDA || + lang == CBM_LANG_OBJC; +} + +/* Is a struct, union or class around an unscoped enum the scope of its + * enumerators? In C++ (and CUDA; every .h is parsed as C++) it is: the + * enumerator is `Brush::ROUND`. C and Objective-C have no struct scope for + * ordinary identifiers: an enum declared inside a struct puts its enumerators + * in the scope around the struct and the code names them unqualified, so the + * struct is no segment of their QN. */ +static bool c_enum_scope_is_class(CBMLanguage lang) { + return lang == CBM_LANG_CPP || lang == CBM_LANG_CUDA; +} + +/* C++11 `enum class X` / `enum struct X`: the scoping keyword is an anonymous + * token child between `enum` and the body. Its enumerators stay `X::A`. */ +static bool c_enum_is_scoped(TSNode enum_node) { + TSNode body = ts_node_child_by_field_name(enum_node, TS_FIELD("body")); + uint32_t stop = ts_node_is_null(body) ? UINT32_MAX : ts_node_start_byte(body); + uint32_t nc = ts_node_child_count(enum_node); + for (uint32_t i = 0; i < nc; i++) { + TSNode c = ts_node_child(enum_node, i); + if (ts_node_start_byte(c) >= stop) { + break; + } + if (!ts_node_is_named(c) && + (strcmp(ts_node_type(c), "class") == 0 || strcmp(ts_node_type(c), "struct") == 0)) { + return true; + } + } + return false; +} + /* Render the canonical return type into `out`; returns how many qualifiers and * markers were added around the base type (0 = the base text alone is already * the whole type). The declarator walk is one strict child chain, so it is @@ -5957,6 +6030,26 @@ static void extract_class_def(CBMExtractCtx *ctx, TSNode node, const CBMLangSpec if (extract_sql_ddl_class_def(ctx, node, kind)) { return; } + if (is_c_tag_reference(ctx->language, node, kind)) { + return; + } + /* A C/C++ typedef carries no `name` field (the alias names live in its + * declarators), so the generic path below always dropped it. */ + if (is_c_declarator_lang(ctx->language) && strcmp(kind, "type_definition") == 0) { + extract_c_typedef(ctx, node, spec); + return; + } + /* An anonymous C-family enum (`enum { A, B };`, also as a variable's or + * member's type) has no Enum def to mint, but its enumerators are names of + * the enclosing scope all the same. Inside a typedef, extract_c_typedef + * names the enum and emits them. */ + if (c_enum_lang(ctx->language) && strcmp(kind, "enum_specifier") == 0 && + ts_node_is_null(ts_node_child_by_field_name(node, TS_FIELD("name")))) { + if (!doc_kind_is(doc_parent(ctx, node), "type_definition")) { + extract_enum_members(ctx, node, NULL); + } + return; + } TSNode name_node = ts_node_child_by_field_name(node, TS_FIELD("name")); // ObjC: class name is first identifier child @@ -6221,6 +6314,16 @@ static void extract_class_def(CBMExtractCtx *ctx, TSNode node, const CBMLangSpec if (!name || !name[0]) { return; } + emit_class_def(ctx, node, spec, kind, name); +} + +/* Emit the class-like def for `node` under `name`, plus its members (enum + * members, methods, fields, class variables). Split from extract_class_def so a + * C `typedef struct { ... } Name;` can name its anonymous struct after the + * typedef (extract_c_typedef). */ +static void emit_class_def(CBMExtractCtx *ctx, TSNode node, const CBMLangSpec *spec, + const char *kind, char *name) { + CBMArena *a = ctx->arena; // For nested classes, prefix with enclosing class QN (e.g., Outer.Inner). // Top-level classes use the language-aware module QN so Java/Go don't double @@ -6383,6 +6486,139 @@ static void extract_class_def(CBMExtractCtx *ctx, TSNode node, const CBMLangSpec } } +/* The alias a C typedef declarator introduces: the type_identifier at the end + * of its pointer / array / function / parenthesized declarator chain + * (`typedef int (*cmp_fn)(const void *, const void *);` names cmp_fn). The C + * grammar lexes stdint-style names as primitive_type (`typedef unsigned + * uint32_t;` in a compat header), so that leaf counts as well. */ +static TSNode c_typedef_alias_node(TSNode decl) { + for (int depth = 0; depth < DECLARATOR_DEPTH_LIMIT && !ts_node_is_null(decl); depth++) { + const char *dk = ts_node_type(decl); + if (strcmp(dk, "type_identifier") == 0 || strcmp(dk, "primitive_type") == 0) { + return decl; + } + TSNode inner = ts_node_child_by_field_name(decl, TS_FIELD("declarator")); + if (ts_node_is_null(inner) && ts_node_named_child_count(decl) > 0) { + inner = ts_node_named_child(decl, 0); + } + decl = inner; + } + TSNode null_node = {0}; + return null_node; +} + +/* C/C++ `typedef`: one def per alias name its declarators introduce. + * typedef struct { ... } Name; -> the anonymous struct IS Name: Class/Enum + * Name with its fields / enum members + * typedef struct Tag { ... } Tag; -> nothing extra: the struct def Tag (the + * walk reaches the specifier) is the entity + * typedef struct Tag Name; -> Type Name (alias; the struct is a + * typedef struct Tag Tag; reference, see is_c_tag_reference). An + * typedef unsigned int Name; identity alias of an incomplete struct is + * typedef int (*Name)(int); often the only declaration of an opaque + * handle type in the repo. + * The type_definition span and the doc comment above it belong to every alias. + * QN scheme unchanged: ., or . inside a + * C++ class body. drop_c_typedefs_shadowed_by_tags then removes an alias whose + * QN a struct/union/enum definition in the same file holds. */ +static void extract_c_typedef(CBMExtractCtx *ctx, TSNode node, const CBMLangSpec *spec) { + CBMArena *a = ctx->arena; + TSNode type = ts_node_child_by_field_name(node, TS_FIELD("type")); + const char *tk = ts_node_is_null(type) ? "" : ts_node_type(type); + bool tag_kind = strcmp(tk, "struct_specifier") == 0 || strcmp(tk, "union_specifier") == 0 || + strcmp(tk, "enum_specifier") == 0 || strcmp(tk, "class_specifier") == 0; + bool has_body = + tag_kind && !ts_node_is_null(ts_node_child_by_field_name(type, TS_FIELD("body"))); + TSNode tag_node = {0}; + if (tag_kind) { + tag_node = ts_node_child_by_field_name(type, TS_FIELD("name")); + } + /* Only a tag DEFINED here makes a same-named alias redundant. */ + const char *tag = + (ts_node_is_null(tag_node) || !has_body) ? NULL : cbm_node_text(a, tag_node, ctx->source); + bool anonymous_body = has_body && ts_node_is_null(tag_node); + + TSTreeCursor cursor = ts_tree_cursor_new(node); + bool more = ts_tree_cursor_goto_first_child(&cursor); + for (; more; more = ts_tree_cursor_goto_next_sibling(&cursor)) { + const char *field = ts_tree_cursor_current_field_name(&cursor); + if (!field || strcmp(field, "declarator") != 0) { + continue; + } + TSNode decl = ts_tree_cursor_current_node(&cursor); + TSNode alias_node = c_typedef_alias_node(decl); + char *name = ts_node_is_null(alias_node) ? NULL : cbm_node_text(a, alias_node, ctx->source); + if (!name || !name[0] || (tag && strcmp(tag, name) == 0)) { + continue; + } + if (anonymous_body && ts_node_eq(decl, alias_node)) { + /* its doc is the comment above the typedef (doc_anchor_c) */ + emit_class_def(ctx, type, spec, tk, name); + anonymous_body = false; /* a second plain name is an alias of the first */ + continue; + } + CBMDefinition def; + memset(&def, 0, sizeof(def)); + def.name = name; + def.qualified_name = + ctx->enclosing_class_qn + ? cbm_arena_sprintf(a, "%s.%s", ctx->enclosing_class_qn, name) + : cbm_fqn_compute_source_lang(a, ctx->project, ctx->rel_path, name, ctx->language); + def.label = "Type"; + def.file_path = ctx->rel_path; + def.start_line = ts_node_start_point(node).row + TS_LINE_OFFSET; + def.end_line = ts_node_end_point(node).row + TS_LINE_OFFSET; + def.lines = (int)(def.end_line - def.start_line + TS_LINE_OFFSET); + def.is_exported = cbm_is_exported(name, ctx->language); + def.docstring = extract_docstring(ctx, node, name); + cbm_defs_push(&ctx->result->defs, a, def); + } + ts_tree_cursor_delete(&cursor); + /* No plain declarator named the anonymous enum (`typedef enum {A, B} *p;`): + * its enumerators are still names of the enclosing scope. */ + if (anonymous_body && c_enum_lang(ctx->language) && strcmp(tk, "enum_specifier") == 0) { + extract_enum_members(ctx, type, NULL); + } +} + +/* `typedef struct X X;` plus `struct X { ... };` in one file give the alias and + * the definition one QN, and the later line would win the graph upsert. The + * definition (fields, enum members, span) is the entity: drop the alias. Only + * defs from `first` on (this extraction pass) are considered, so the + * preprocessed rescue pass never filters the raw pass's defs. */ +static void drop_c_typedefs_shadowed_by_tags(CBMExtractCtx *ctx, int first) { + CBMDefArray *defs = &ctx->result->defs; + bool any_alias = false; + for (int i = first; i < defs->count && !any_alias; i++) { + any_alias = defs->items[i].label && strcmp(defs->items[i].label, "Type") == 0; + } + if (!any_alias) { + return; + } + CBMHashTable *tags = cbm_ht_create(CBM_SZ_64); + if (!tags) { + return; + } + for (int i = first; i < defs->count; i++) { + const CBMDefinition *d = &defs->items[i]; + if (d->label && d->qualified_name && + (strcmp(d->label, "Class") == 0 || strcmp(d->label, "Enum") == 0)) { + cbm_ht_set(tags, d->qualified_name, (void *)d); + } + } + int w = first; + for (int i = first; i < defs->count; i++) { + const CBMDefinition *d = &defs->items[i]; + if (d->label && d->qualified_name && strcmp(d->label, "Type") == 0 && + cbm_ht_has(tags, d->qualified_name)) { + continue; + } + defs->items[w++] = defs->items[i]; + } + defs->count = w; + cbm_ht_free(tags); +} + // Find the body/members node inside a class node static TSNode find_class_body(TSNode class_node, CBMLanguage lang) { // Try field names first @@ -7184,40 +7420,144 @@ static bool is_enum_member_kind(const char *kind) { strcmp(kind, "enumerator") == 0; } +/* How an enum's members are named. + * + * The member QN is `.`, with one exception: a C-family + * UNSCOPED enum (c_enum_lang, not `enum class` / `enum struct`). Its enumerators + * live in the scope that holds the enum, so their QN is `.` + * and `class_qn` -- NULL for an anonymous enum -- is not a segment. That scope + * is the module, or in C++ the enclosing namespace or class; a C or + * Objective-C struct is none (c_enum_scope_is_class). Every C-family + * enumerator of a named enum carries `parent_class` = the enum's QN, which is + * what ties a flattened enumerator to its enum. */ +typedef struct { + const char *class_qn; /* the enum's QN; NULL for an anonymous enum */ + bool c_enum; /* a C-family enum_specifier */ + bool flat; /* the member QN leaves the enum's name out */ +} enum_member_scope_t; + +/* One enum member -> a Variable def. `doc_node` is the node whose leading + * comments document it: the member itself, or the ERROR that opens its slot + * in a macro-wrapped C list (see extract_c_enum_members). */ +static void push_enum_member(CBMExtractCtx *ctx, TSNode member, TSNode doc_node, + const enum_member_scope_t *scope) { + CBMArena *a = ctx->arena; + TSNode mname = ts_node_child_by_field_name(member, TS_FIELD("name")); + if (ts_node_is_null(mname)) { + mname = cbm_find_child_by_kind(member, "identifier"); + } + if (ts_node_is_null(mname)) { + return; + } + char *member_name = cbm_node_text(a, mname, ctx->source); + if (!member_name || !member_name[0]) { + return; + } + CBMDefinition mdef; + memset(&mdef, 0, sizeof(mdef)); + mdef.name = member_name; + if (!scope->flat) { + mdef.qualified_name = cbm_arena_sprintf(a, "%s.%s", scope->class_qn, member_name); + } else if (ctx->enclosing_class_qn && c_enum_scope_is_class(ctx->language)) { + mdef.qualified_name = cbm_arena_sprintf(a, "%s.%s", ctx->enclosing_class_qn, member_name); + } else { + mdef.qualified_name = + cbm_fqn_compute_source_lang(a, ctx->project, ctx->rel_path, member_name, ctx->language); + } + if (scope->c_enum && scope->class_qn) { + mdef.parent_class = scope->class_qn; + } + mdef.label = "Variable"; + mdef.file_path = ctx->rel_path; + mdef.start_line = ts_node_start_point(member).row + TS_LINE_OFFSET; + mdef.end_line = ts_node_end_point(member).row + TS_LINE_OFFSET; + mdef.docstring = extract_member_docstring(ctx, doc_node); + cbm_defs_push(&ctx->result->defs, a, mdef); +} + +typedef struct { + int depth; /* parentheses an ERROR child opened and none closed yet */ + bool open; /* the current slot has no constant yet */ + bool has_head; /* an ERROR opened the current slot: `head` is that node */ + TSNode head; +} c_enum_slot_t; + +/* The parentheses and slot separators an ERROR child of the list carries. */ +static void c_enum_slot_scan_error(TSNode err, c_enum_slot_t *slot) { + if (slot->open && !slot->has_head) { + slot->head = err; + slot->has_head = true; + } + uint32_t n = ts_node_child_count(err); + for (uint32_t i = 0; i < n; i++) { + const char *k = ts_node_type(ts_node_child(err, i)); + if (strcmp(k, "(") == 0) { + slot->depth++; + } else if (strcmp(k, ")") == 0) { + if (slot->depth > 0) { + slot->depth--; + } + } else if (strcmp(k, ",") == 0 && slot->depth == 0) { + slot->open = true; + slot->has_head = false; + } + } +} + +/* A C-family enumerator_list under error recovery. A macro in the list -- + * CURLOPT(CURLOPT_URL, CURLOPTTYPE_STRINGPOINT, 2), + * CURLINFO_SPEED CURL_DEPRECATED(7.55.0, "...") = CURLINFO_DOUBLE + 9, + * -- is not enumerator grammar: the parser keeps every bare identifier it can + * as an `enumerator` and wraps the rest in ERROR nodes, so macro arguments and + * value operands come back as constants that do not exist (and, flattened, + * would take the plain QN of the macro they really are). The constant is the + * enumerator that OPENS a slot: the first one after `{` or after a `,` outside + * parentheses. A macro call that opens the slot hands that role to its first + * argument (the X-macro form), and the comment above the call documents it. + * A well-formed list has one enumerator per slot, so nothing changes there. */ +static void extract_c_enum_members(CBMExtractCtx *ctx, TSNode body, + const enum_member_scope_t *scope) { + c_enum_slot_t slot = {.depth = 0, .open = true, .has_head = false, .head = body}; + TSTreeCursor cur = ts_tree_cursor_new(body); + if (ts_tree_cursor_goto_first_child(&cur)) { + do { + TSNode child = ts_tree_cursor_current_node(&cur); + const char *ck = ts_node_type(child); + if (strcmp(ck, ",") == 0) { + if (slot.depth == 0) { + slot.open = true; + slot.has_head = false; + } + } else if (strcmp(ck, "ERROR") == 0) { + c_enum_slot_scan_error(child, &slot); + } else if (strcmp(ck, "enumerator") == 0 && slot.open) { + push_enum_member(ctx, child, slot.has_head ? slot.head : child, scope); + slot.open = false; + } + } while (ts_tree_cursor_goto_next_sibling(&cur)); + } + ts_tree_cursor_delete(&cur); +} + /* Extract enum members as Variable nodes (C#, Java, TypeScript, C++). */ static void extract_enum_members(CBMExtractCtx *ctx, TSNode node, const char *class_qn) { - CBMArena *a = ctx->arena; TSNode body = find_class_body(node, ctx->language); if (ts_node_is_null(body)) { return; } + enum_member_scope_t scope = {.class_qn = class_qn, .c_enum = false, .flat = false}; + scope.c_enum = c_enum_lang(ctx->language) && strcmp(ts_node_type(node), "enum_specifier") == 0; + scope.flat = scope.c_enum && (!class_qn || !c_enum_is_scoped(node)); + if (scope.c_enum) { + extract_c_enum_members(ctx, body, &scope); + return; + } uint32_t mc = ts_node_named_child_count(body); for (uint32_t mi = 0; mi < mc; mi++) { TSNode member = ts_node_named_child(body, mi); - if (!is_enum_member_kind(ts_node_type(member))) { - continue; - } - TSNode mname = ts_node_child_by_field_name(member, TS_FIELD("name")); - if (ts_node_is_null(mname)) { - mname = cbm_find_child_by_kind(member, "identifier"); - } - if (ts_node_is_null(mname)) { - continue; - } - char *member_name = cbm_node_text(a, mname, ctx->source); - if (!member_name || !member_name[0]) { - continue; + if (is_enum_member_kind(ts_node_type(member))) { + push_enum_member(ctx, member, member, &scope); } - CBMDefinition mdef; - memset(&mdef, 0, sizeof(mdef)); - mdef.name = member_name; - mdef.qualified_name = cbm_arena_sprintf(a, "%s.%s", class_qn, member_name); - mdef.label = "Variable"; - mdef.file_path = ctx->rel_path; - mdef.start_line = ts_node_start_point(member).row + TS_LINE_OFFSET; - mdef.end_line = ts_node_end_point(member).row + TS_LINE_OFFSET; - mdef.docstring = extract_member_docstring(ctx, member); - cbm_defs_push(&ctx->result->defs, a, mdef); } } @@ -8982,6 +9322,202 @@ static void wd_push_children_reverse(wd_stack_t *s, TSNode node, const char *enc free(kids); } +/* 1-based line of the `}` that closes the `{` at `open`, or 0 when it does not + * close before `limit`. Counts only the first branch of every #if group (the + * first-branch projection rule of the C rescue in cbm.c): `#ifdef X struct { + * #else union { #endif` opens one brace, not two. Comments, string and char + * literals are skipped; dropped branches are ignored wholesale. */ +enum { C_BRACE_PP_DEPTH = 64 }; + +static bool c_brace_directive_is(const char *p, const char *e, const char *w) { + size_t n = strlen(w); + return (size_t)(e - p) >= n && strncmp(p, w, n) == 0 && + ((size_t)(e - p) == n || !isalpha((unsigned char)p[n])); +} + +static uint32_t c_matching_brace_line(const char *src, uint32_t open, uint32_t limit, + uint32_t open_row) { + uint8_t branch[C_BRACE_PP_DEPTH]; + uint8_t keep[C_BRACE_PP_DEPTH]; + int pp = 0; + int depth = 0; + uint32_t row = open_row; + bool line_start = false; + for (uint32_t i = open; i < limit; i++) { + char c = src[i]; + if (c == '\n') { + row++; + line_start = true; + continue; + } + if (line_start && (c == ' ' || c == '\t')) { + continue; + } + if (line_start && c == '#') { + const char *p = src + i + 1; + const char *e = src + limit; + while (p < e && (*p == ' ' || *p == '\t')) { + p++; + } + if (c_brace_directive_is(p, e, "if") || c_brace_directive_is(p, e, "ifdef") || + c_brace_directive_is(p, e, "ifndef")) { + if (pp >= C_BRACE_PP_DEPTH) { + return 0; + } + const char *q = p + 2; + while (q < e && isalpha((unsigned char)*q)) { + q++; + } + while (q < e && (*q == ' ' || *q == '\t')) { + q++; + } + bool zero = c_brace_directive_is(p, e, "if") && q < e && *q == '0' && + (q + 1 >= e || !isalnum((unsigned char)q[1])); + branch[pp] = 0; + keep[pp] = zero ? 1 : 0; + pp++; + } else if (c_brace_directive_is(p, e, "elif") || c_brace_directive_is(p, e, "else") || + c_brace_directive_is(p, e, "elifdef") || + c_brace_directive_is(p, e, "elifndef")) { + if (pp > 0) { + branch[pp - 1] = branch[pp - 1] < UINT8_MAX ? branch[pp - 1] + 1 : UINT8_MAX; + } else { + /* the next branch of a group opened before the brace: drop + * everything up to that group's #endif */ + branch[0] = 1; + keep[0] = 0; + pp = 1; + } + } else if (c_brace_directive_is(p, e, "endif") && pp > 0) { + pp--; + } + /* skip the directive's logical line, backslash continuations included */ + for (;;) { + while (i + 1 < limit && src[i + 1] != '\n') { + i++; + } + if (i + 1 < limit && src[i] == '\\') { + i++; /* onto the continued line's newline */ + row++; + continue; + } + break; + } + continue; + } + line_start = false; + bool active = true; + for (int k = 0; k < pp && active; k++) { + active = branch[k] == keep[k]; + } + if (!active) { + continue; + } + if (c == '/' && i + 1 < limit && src[i + 1] == '*') { + for (i += 2; i + 1 < limit && !(src[i] == '*' && src[i + 1] == '/'); i++) { + row += src[i] == '\n'; + } + i++; + continue; + } + if (c == '/' && i + 1 < limit && src[i + 1] == '/') { + while (i + 1 < limit && src[i + 1] != '\n') { + i++; + } + continue; + } + if (c == '"' || c == '\'') { + for (i++; i < limit && src[i] != c && src[i] != '\n'; i++) { + if (src[i] == '\\' && i + 1 < limit) { + i++; + row += src[i] == '\n'; + } + } + if (i < limit && src[i] == '\n') { + i--; /* unterminated literal: let the loop count the newline */ + } + continue; + } + if (c == '{') { + depth++; + } else if (c == '}' && --depth == 0) { + return row + TS_LINE_OFFSET; + } + } + return 0; +} + +/* tree-sitter can give up on a whole region and leave a definition head as loose + * ERROR tokens: `struct task_struct {` in the kernel's include-guarded sched.h is + * never recovered as a struct_specifier, so the type had no definition node (its + * only nodes were references, now dropped by is_c_tag_reference). The tokens + * `struct|union|class|enum NAME {` can only start a definition of NAME; recover it + * with its span up to the matching brace (members stay with the ERROR region). */ +static void recover_c_error_tag_heads(CBMExtractCtx *ctx, TSNode error_node) { + CBMArena *a = ctx->arena; + /* One linear cursor pass with a three-token window: a file-level ERROR can + * hold every top-level token of the file as a flat sibling, where indexed + * child access is quadratic (see wd_collect_children). */ + TSTreeCursor cursor = ts_tree_cursor_new(error_node); + TSNode win[3]; + memset(win, 0, sizeof(win)); + uint32_t seen = 0; + /* Once one head never closes, later heads in this region will not either: + * do not rescan to the region end for each of them. */ + bool unclosed = false; + bool more = ts_tree_cursor_goto_first_child(&cursor); + for (; more; more = ts_tree_cursor_goto_next_sibling(&cursor)) { + win[0] = win[1]; + win[1] = win[2]; + win[2] = ts_tree_cursor_current_node(&cursor); + if (++seen < 3) { + continue; + } + TSNode kw = win[0]; + TSNode name_node = win[1]; + TSNode brace = win[2]; + if (ts_node_is_named(kw)) { + continue; + } + const char *k = ts_node_type(kw); + bool is_enum = strcmp(k, "enum") == 0; + if (!is_enum && strcmp(k, "struct") != 0 && strcmp(k, "union") != 0 && + strcmp(k, "class") != 0) { + continue; + } + if (strcmp(ts_node_type(name_node), "type_identifier") != 0 || ts_node_is_named(brace) || + strcmp(ts_node_type(brace), "{") != 0) { + continue; + } + char *name = cbm_node_text(a, name_node, ctx->source); + if (!name || !name[0]) { + continue; + } + uint32_t end_line = unclosed ? 0 + : c_matching_brace_line(ctx->source, ts_node_start_byte(brace), + ts_node_end_byte(error_node), + ts_node_start_point(brace).row); + unclosed = end_line == 0; + CBMDefinition def; + memset(&def, 0, sizeof(def)); + def.name = name; + def.qualified_name = + ctx->enclosing_class_qn + ? cbm_arena_sprintf(a, "%s.%s", ctx->enclosing_class_qn, name) + : cbm_fqn_compute_source_lang(a, ctx->project, ctx->rel_path, name, ctx->language); + def.label = is_enum ? "Enum" : "Class"; + def.file_path = ctx->rel_path; + def.start_line = ts_node_start_point(kw).row + TS_LINE_OFFSET; + def.end_line = end_line ? end_line : ts_node_end_point(error_node).row + TS_LINE_OFFSET; + def.lines = (int)(def.end_line - def.start_line + TS_LINE_OFFSET); + def.is_exported = cbm_is_exported(name, ctx->language); + /* The doc comment sits before the ERROR node when the head opens it. */ + def.docstring = extract_docstring(ctx, seen == 3 ? error_node : kw, name); + cbm_defs_push(&ctx->result->defs, a, def); + } + ts_tree_cursor_delete(&cursor); +} + // Push nested class nodes from a class body container onto the defs stack. // Iteratively walks into wrapper nodes (field_declaration, template_declaration). static void push_nested_class_nodes(TSNode body, const CBMLangSpec *spec, wd_stack_t *s, @@ -9328,17 +9864,16 @@ static void extract_janet_def(CBMExtractCtx *ctx, TSNode node) { cbm_defs_push(&ctx->result->defs, a, def); } -// Languages that use the C preprocessor and therefore have #define macros. -static bool is_c_preprocessor_lang(CBMLanguage lang) { - return lang == CBM_LANG_C || lang == CBM_LANG_CPP || lang == CBM_LANG_CUDA || - lang == CBM_LANG_GLSL || lang == CBM_LANG_OBJC || lang == CBM_LANG_ISPC; -} - -// C/C++ preprocessor macros become Macro nodes (#375): +// C/C++ preprocessor macros become Macro nodes (#375), in the languages of +// cbm_is_c_preprocessor_lang: // #define SIMPLE 1 -> preproc_def // #define FN(x) (2 * (x)) -> preproc_function_def // The name is the `name` field; a function-like macro's parameter list is kept // as the signature. The macro body (a preproc_arg) is not descended into. +// The QN is `.#macro` (CBM_MACRO_QN_SUFFIX): macros have a +// namespace of their own, so a rename macro (`#define XXH32 XXH_NAME2(...)`) or +// an #else stand-in (`#define match(c, m) TRUE`) never competes with the +// function, type or enumerator of the same name for one node. static void extract_c_macro_def(CBMExtractCtx *ctx, TSNode node) { CBMArena *a = ctx->arena; TSNode name_node = ts_node_child_by_field_name(node, TS_FIELD("name")); @@ -9350,10 +9885,15 @@ static void extract_c_macro_def(CBMExtractCtx *ctx, TSNode node) { return; } + const char *plain_qn = cbm_fqn_compute(a, ctx->project, ctx->rel_path, name); + if (!plain_qn) { + return; + } + CBMDefinition def; memset(&def, 0, sizeof(def)); def.name = name; - def.qualified_name = cbm_fqn_compute(a, ctx->project, ctx->rel_path, name); + def.qualified_name = cbm_arena_sprintf(a, "%s" CBM_MACRO_QN_SUFFIX, plain_qn); def.label = "Macro"; def.file_path = ctx->rel_path; def.start_line = ts_node_start_point(node).row + TS_LINE_OFFSET; @@ -9628,13 +10168,18 @@ static void walk_defs(CBMExtractCtx *ctx, TSNode root, const CBMLangSpec *spec, if (ctx->language == CBM_LANG_KOTLIN && strcmp(kind, "ERROR") == 0) { recover_kotlin_error_classes(ctx, node); } + /* C/C++: a definition head left as loose ERROR tokens (additive, as + * above — descent below still visits the region's parsed subtrees). */ + if (is_c_declarator_lang(ctx->language) && strcmp(kind, "ERROR") == 0) { + recover_c_error_tag_heads(ctx, node); + } if (ctx->language == CBM_LANG_ELIXIR && strcmp(kind, "call") == 0) { extract_elixir_call(ctx, node, spec); continue; } - if (is_c_preprocessor_lang(ctx->language) && + if (cbm_is_c_preprocessor_lang(ctx->language) && (strcmp(kind, "preproc_def") == 0 || strcmp(kind, "preproc_function_def") == 0)) { // Gated to full/advanced index modes — macros dominate extraction on // macro-dense codebases (e.g. the Linux kernel). See #375. @@ -9786,7 +10331,11 @@ void cbm_extract_definitions_without_module(CBMExtractCtx *ctx) { } // Walk AST for function/class definitions + int first = ctx->result->defs.count; walk_defs(ctx, ctx->root, spec, 0); + if (is_c_declarator_lang(ctx->language)) { + drop_c_typedefs_shadowed_by_tags(ctx, first); + } // Extract module-level variables extract_variables(ctx, ctx->root, spec); diff --git a/internal/cbm/helpers.c b/internal/cbm/helpers.c index 4c4d340f0b..03cc8b685d 100644 --- a/internal/cbm/helpers.c +++ b/internal/cbm/helpers.c @@ -988,16 +988,109 @@ static TSNode resolve_qualified_name(TSNode decl) { return null_node; } +bool cbm_c_qualifier_is_recovered(TSNode qid) { + if (ts_node_is_null(qid)) { + return false; + } + uint32_t nc = ts_node_child_count(qid); + for (uint32_t i = 0; i < nc; i++) { + TSNode c = ts_node_child(qid, i); + if (!ts_node_is_named(c) && strcmp(ts_node_type(c), "::") == 0) { + return ts_node_is_missing(c); + } + } + return false; +} + +/* A function DEFINITION's parameter list: empty, `(void)`, or every parameter + * named (variadic `...` allowed). A function-like macro invocation that only + * looks like a declarator — `TEST_BEGIN(test_name)`, whose one "parameter" is + * an untyped identifier read as a type — fails this, so its macro name is never + * mistaken for the defined function's name. */ +static bool c_params_name_every_parameter(TSNode params) { + enum { MAX_PARAMS_CHECKED = 64 }; /* indexed child access: keep it small */ + uint32_t nc = ts_node_named_child_count(params); + if (nc > MAX_PARAMS_CHECKED) { + return false; + } + for (uint32_t i = 0; i < nc; i++) { + TSNode p = ts_node_named_child(params, i); + const char *pk = ts_node_type(p); + if (strcmp(pk, "comment") == 0 || strcmp(pk, "variadic_parameter") == 0) { + continue; + } + if (!ts_node_is_null(ts_node_child_by_field_name(p, TS_FIELD("declarator")))) { + continue; + } + /* `(void)`: the one parameter is an unnamed primitive type. */ + TSNode type = ts_node_child_by_field_name(p, TS_FIELD("type")); + if (nc != 1 || ts_node_is_null(type) || strcmp(ts_node_type(type), "primitive_type") != 0) { + return false; + } + } + return true; +} + +TSNode cbm_c_recovered_func_name(TSNode func_declarator) { + TSNode null_node = {0}; + if (ts_node_is_null(func_declarator) || + strcmp(ts_node_type(func_declarator), "function_declarator") != 0) { + return null_node; + } + TSNode params = ts_node_child_by_field_name(func_declarator, TS_FIELD("parameters")); + if (ts_node_is_null(params) || !c_params_name_every_parameter(params)) { + return null_node; + } + TSNode err = ts_node_prev_sibling(params); + if (ts_node_is_null(err) || strcmp(ts_node_type(err), "ERROR") != 0) { + return null_node; + } + /* Only the macro-prefix shape: the ERROR holds nothing but bare identifiers + * (the real name, possibly after further attribute macros). Any other token + * means the region is not a declaration this rule can vouch for. */ + enum { MAX_PREFIX_TOKENS = 4 }; + uint32_t nc = ts_node_child_count(err); + if (nc == 0 || nc > MAX_PREFIX_TOKENS) { + return null_node; + } + TSNode last = null_node; + for (uint32_t i = 0; i < nc; i++) { + TSNode c = ts_node_child(err, i); + const char *ck = ts_node_type(c); + if (strcmp(ck, "identifier") != 0 && strcmp(ck, "type_identifier") != 0) { + return null_node; + } + last = c; + } + return last; +} + // Resolve function name from C/C++/CUDA/GLSL declarator chain. Shared canonical // implementation — see the header for the full rationale (#438). TSNode cbm_resolve_c_declarator_name_node(TSNode func_node) { TSNode decl = ts_node_child_by_field_name(func_node, TS_FIELD("declarator")); + /* Set once a recovered `RetT::name` qualifier was stepped through: on that + * path the C++ grammar names the function with a type_identifier. */ + bool recovered = false; for (int depth = 0; depth < CBM_DECLARATOR_DEPTH_LIMIT && !ts_node_is_null(decl); depth++) { const char *dk = ts_node_type(decl); - if (is_c_terminal_name(dk)) { + if (is_c_terminal_name(dk) || (recovered && strcmp(dk, "type_identifier") == 0)) { return decl; } + if (strcmp(dk, "function_declarator") == 0) { + TSNode real = cbm_c_recovered_func_name(decl); + if (!ts_node_is_null(real)) { + return real; + } + } if (strcmp(dk, "qualified_identifier") == 0 || strcmp(dk, "scoped_identifier") == 0) { + if (cbm_c_qualifier_is_recovered(decl)) { + /* `API RetT name(...)`: the "scope" is the return type, the name + * side carries the real declarator chain. */ + decl = ts_node_child_by_field_name(decl, TS_FIELD("name")); + recovered = true; + continue; + } return resolve_qualified_name(decl); } TSNode inner = ts_node_child_by_field_name(decl, TS_FIELD("declarator")); @@ -1013,6 +1106,36 @@ TSNode cbm_resolve_c_declarator_name_node(TSNode func_node) { return null_node; } +/* The C-declarator grammars: same set resolve_func_name_c_family routes through + * cbm_resolve_c_declarator_name_node. */ +static bool c_declarator_lang(CBMLanguage lang) { + return lang == CBM_LANG_C || lang == CBM_LANG_CPP || lang == CBM_LANG_CUDA || + lang == CBM_LANG_GLSL || lang == CBM_LANG_HLSL || lang == CBM_LANG_ISPC || + lang == CBM_LANG_SLANG || lang == CBM_LANG_OBJC; +} + +bool cbm_c_reserved_func_name(const char *name) { + /* Statement/operator keywords: error recovery over preprocessor-split code + * reads `else if (a) (b) {` as a definition `else if(...) {...}`. */ + static const char *const kw[] = {"if", "else", "for", "while", "do", + "switch", "case", "default", "return", "break", + "continue", "goto", "sizeof", "typedef", NULL}; + if (!name) { + return false; + } + for (const char *const *k = kw; *k; k++) { + if (strcmp(name, *k) == 0) { + return true; + } + } + return false; +} + +bool cbm_is_c_preprocessor_lang(CBMLanguage lang) { + return lang == CBM_LANG_C || lang == CBM_LANG_CPP || lang == CBM_LANG_CUDA || + lang == CBM_LANG_GLSL || lang == CBM_LANG_OBJC || lang == CBM_LANG_ISPC; +} + // Convert a resolved function/method name node to its name string. Most nodes // map directly to their text, but a C++ conversion-operator's `operator_cast` // node spans the full "operator bool() const" — this grammar folds the parameter @@ -1023,6 +1146,12 @@ TSNode cbm_resolve_c_declarator_name_node(TSNode func_node) { // `if (obj)`) misses. char *cbm_func_name_node_text(CBMArena *a, TSNode name_node, const char *source, CBMLanguage lang) { char *text = cbm_node_text(a, name_node, source); + /* A C keyword is never a function name: the definition is an error-recovery + * artifact. No name means no def and no call scope (calls inside fall back + * to the enclosing scope), the same as any unnamed definition. */ + if (text && c_declarator_lang(lang) && cbm_c_reserved_func_name(text)) { + return NULL; + } if (text && strcmp(ts_node_type(name_node), "operator_cast") == 0) { char *paren = strchr(text, '('); if (paren) { diff --git a/internal/cbm/helpers.h b/internal/cbm/helpers.h index 00e3d30dcc..5426152313 100644 --- a/internal/cbm/helpers.h +++ b/internal/cbm/helpers.h @@ -58,6 +58,37 @@ const char *cbm_enclosing_func_qn_cached(CBMExtractCtx *ctx, TSNode node); // enclosing-function attribution — drift between private copies caused #438. TSNode cbm_resolve_c_declarator_name_node(TSNode func_node); +// tree-sitter error recovery around an unknown leading macro: `API RetT name(...)` +// (`XXH_PUBLIC_API XXH64_hash_t XXH64(...)`, `JEMALLOC_ALWAYS_INLINE T *f(...)`) +// carries one identifier more than the declaration grammar admits, so the parser +// recovers in one of two shapes, and neither names the function: +// - C++ inserts a zero-width MISSING "::" and reads `RetT name` as the +// out-of-line name `RetT::name`; with a pointer return, the name side is a +// pointer_type_declarator that wraps the function_declarator; +// - C and C++ keep RetT as the declarator identifier and put the real name in +// an ERROR node directly before the parameter list. +// cbm_c_qualifier_is_recovered: true when `qid` (a qualified_identifier or +// scoped_identifier) is the first shape, i.e. its "::" token is MISSING. +// cbm_c_recovered_func_name: the real name node of the second shape for a +// function_declarator, or a null node when the declarator is not that shape. +// Shared by every C-family declarator walk (defs, call scope, params, C LSP) so +// the def QN and every caller QN agree on the recovered name. +bool cbm_c_qualifier_is_recovered(TSNode qid); +TSNode cbm_c_recovered_func_name(TSNode func_declarator); + +// True for a C keyword that error recovery can surface as a function "name" +// (`else if (a) (b) {` split by #ifdef reads as a definition named `if`). +// cbm_func_name_node_text returns NULL for these in the C-declarator grammars, +// so neither a def nor a call scope is minted; the C LSP applies the same test. +bool cbm_c_reserved_func_name(const char *name); + +// True for the languages whose `#define`s become Macro nodes: C, C++, CUDA, +// GLSL, Objective-C and ISPC. The one predicate for what follows from the C +// preprocessor: the `#macro` QN namespace (CBM_MACRO_QN_SUFFIX) and the +// `variants` property (CBMDefinition.variants), because #if branches are what +// lets one of these files define a name more than once. +bool cbm_is_c_preprocessor_lang(CBMLanguage lang); + // Convert a resolved function/method name node to its name string, normalizing a // C++ conversion-operator's `operator_cast` node (which spans the full // "operator bool() const") down to "operator bool". Shared by the defs and diff --git a/internal/cbm/lsp/c_lsp.c b/internal/cbm/lsp/c_lsp.c index 74b97d69d9..7ff93bd0d6 100644 --- a/internal/cbm/lsp/c_lsp.c +++ b/internal/cbm/lsp/c_lsp.c @@ -4795,22 +4795,44 @@ static void c_process_function(CLSPContext *ctx, TSNode func_node) { // Navigate declarator to find name and parameters TSNode cur = decl; + // Set after stepping through a recovered `RetT::name` (see + // cbm_c_qualifier_is_recovered): that path names the function with a + // type_identifier, and the "scope" is the return type, not a class. + bool recovered = false; for (int depth = 0; depth < 10 && !ts_node_is_null(cur); depth++) { const char *dk = ts_node_type(cur); if (strcmp(dk, "function_declarator") == 0) { TSNode fdecl = ts_node_child_by_field_name(cur, "declarator", 10); params_node = ts_node_child_by_field_name(cur, "parameters", 10); + // `API RetT name(...)`: the real name sits in an ERROR before the + // parameters; the declarator identifier is the return type. + TSNode real = cbm_c_recovered_func_name(cur); + if (!ts_node_is_null(real)) { + func_name = c_node_text(ctx, real); + break; + } cur = fdecl; continue; } - if (strcmp(dk, "pointer_declarator") == 0 || strcmp(dk, "reference_declarator") == 0) { + if (strcmp(dk, "pointer_declarator") == 0 || strcmp(dk, "reference_declarator") == 0 || + (recovered && strcmp(dk, "pointer_type_declarator") == 0)) { if (ts_node_named_child_count(cur) > 0) cur = ts_node_named_child(cur, ts_node_named_child_count(cur) - 1); else break; continue; } + if (recovered && strcmp(dk, "type_identifier") == 0) { + func_name = c_node_text(ctx, cur); + break; + } + if ((strcmp(dk, "qualified_identifier") == 0 || strcmp(dk, "scoped_identifier") == 0) && + cbm_c_qualifier_is_recovered(cur)) { + cur = ts_node_child_by_field_name(cur, "name", 4); + recovered = true; + continue; + } if (strcmp(dk, "qualified_identifier") == 0 || strcmp(dk, "scoped_identifier") == 0) { func_name = c_node_text(ctx, cur); // Check if this is a method (has Class:: prefix) @@ -4850,7 +4872,9 @@ static void c_process_function(CLSPContext *ctx, TSNode func_node) { break; } - if (!func_name || !func_name[0]) + // A keyword "name" is an error-recovery artifact (see cbm_c_reserved_func_name): + // the def extractor mints no node for it, so no caller QN may name it either. + if (!func_name || !func_name[0] || cbm_c_reserved_func_name(func_name)) return; // Build enclosing function QN diff --git a/internal/cbm/result_compact.c b/internal/cbm/result_compact.c index bd8b9b8a9a..97b16be6e5 100644 --- a/internal/cbm/result_compact.c +++ b/internal/cbm/result_compact.c @@ -275,6 +275,7 @@ static void cr_walk_def(cr_ctx_t *c, CBMDefinition *d) { cr_str(c, &d->impl_trait); cr_str(c, &d->http_client); cr_str(c, &d->http_base_url); + cr_str(c, &d->variants); } static void cr_walk_call(cr_ctx_t *c, CBMCall *call) { diff --git a/src/graph_buffer/graph_buffer.c b/src/graph_buffer/graph_buffer.c index e55eb28f02..5484e1a4fe 100644 --- a/src/graph_buffer/graph_buffer.c +++ b/src/graph_buffer/graph_buffer.c @@ -837,6 +837,20 @@ void cbm_gbuf_set_next_id(cbm_gbuf_t *gb, int64_t next_id) { /* ── Node operations ─────────────────────────────────────────────── */ +/* First key of the canonical same-QN collision order: how strong a claim the + * entity has on a qualified name. A macro is the weakest: when a "Macro" and + * a definition share a QN, the definition, which carries the body, the + * signature and every edge, owns it whichever comes first. + * + * C-preprocessor macros no longer get here: their QN carries its own + * namespace (`...#macro`, CBM_MACRO_QN_SUFFIX) and cannot equal a definition's. + * The case left is Chialisp, whose `defmacro` / `defmac` defs are labelled + * "Macro" under a plain QN: a macro and a `defun` of one name in one file, or + * in a same-stem `.clsp` / `.clib` pair (one module QN). */ +static int gb_qn_claim_rank(const char *label) { + return (label && strcmp(label, "Macro") == 0) ? 0 : 1; +} + int64_t cbm_gbuf_upsert_node(cbm_gbuf_t *gb, const char *label, const char *name, const char *qualified_name, const char *file_path, int start_line, int end_line, const char *properties_json) { @@ -880,8 +894,13 @@ int64_t cbm_gbuf_upsert_node(cbm_gbuf_t *gb, const char *label, const char *name * of the same entity with new content) replaces the earlier one, * deterministically, because intra-file arrival order is fixed. A * full tie is the same entity re-upserted → refresh in place. - * Kind-disambiguated QNs (the real cure) remain a follow-up. */ - int c = strcmp(file_path ? file_path : "", existing->file_path ? existing->file_path : ""); + * Ahead of all of these, a definition beats a macro + * (gb_qn_claim_rank). Kind-disambiguated QNs (the real cure) remain + * a follow-up. */ + int c = gb_qn_claim_rank(existing->label) - gb_qn_claim_rank(label); + if (c == 0) { + c = strcmp(file_path ? file_path : "", existing->file_path ? existing->file_path : ""); + } if (c == 0) { c = existing->start_line - start_line; } @@ -1480,9 +1499,13 @@ static void merge_update_existing(cbm_gbuf_t *dst, cbm_gbuf_node_t *existing, * the node set (and every downstream consumer) run to run. Winner = * smallest file_path, then LARGEST start_line, then largest * name/label — one total order, commutative, scheduling-free; a full - * tie is the same entity → refresh from src. */ - int c = strcmp(sn->file_path ? sn->file_path : "", + * tie is the same entity → refresh from src. A definition beats a + * macro ahead of all of these (gb_qn_claim_rank). */ + int c = gb_qn_claim_rank(existing->label) - gb_qn_claim_rank(sn->label); + if (c == 0) { + c = strcmp(sn->file_path ? sn->file_path : "", existing->file_path ? existing->file_path : ""); + } if (c == 0) { c = existing->start_line - sn->start_line; } diff --git a/src/mcp/mcp.c b/src/mcp/mcp.c index e67cf4aac1..c95cf5b12a 100644 --- a/src/mcp/mcp.c +++ b/src/mcp/mcp.c @@ -8518,6 +8518,17 @@ static long node_resolution_score(const cbm_node_t *n) { label_rank = RES_RANK_OTHER; } } + /* Tie rule (CBM_MACRO_QN_SUFFIX): a name that is both a definition and a C + * macro -- a typedef or enumerator next to its rename macro -- resolves to + * the definition. The macro scores just under every other definition, and + * still above Module/File; without this a one-line typedef and its macro + * tie on span and the name reads as ambiguous. */ + size_t qn_len = n->qualified_name ? strlen(n->qualified_name) : 0; + size_t fence_len = sizeof(CBM_MACRO_QN_SUFFIX) - SKIP_ONE; + if (label_rank == RES_RANK_OTHER && qn_len > fence_len && + strcmp(n->qualified_name + qn_len - fence_len, CBM_MACRO_QN_SUFFIX) == 0) { + return RES_RANK_OTHER * (long)RES_LABEL_WEIGHT - SKIP_ONE; + } long span = (long)n->end_line - (long)n->start_line; if (span < 0) { span = 0; @@ -12871,6 +12882,19 @@ static char *handle_get_code_snippet(cbm_mcp_server_t *srv, const char *args) { int suffix_count = 0; cbm_store_find_nodes_by_qn_suffix(store, effective_project, qn, &suffix_nodes, &suffix_count); + /* Tier 3: the C-macro namespace. A macro's QN is `.#macro` + * (CBM_MACRO_QN_SUFFIX), out of reach of '%.X'. Tried only when no + * definition answered to the name above (tie rule), for a short name, a + * partial QN, or the macro's plain QN. */ + char macro_qn[CBM_SZ_512]; + int macro_len = snprintf(macro_qn, sizeof(macro_qn), "%s" CBM_MACRO_QN_SUFFIX, qn); + if (suffix_count == 0 && macro_len > 0 && (size_t)macro_len < sizeof(macro_qn)) { + cbm_store_free_nodes(suffix_nodes, suffix_count); + suffix_nodes = NULL; + cbm_store_find_nodes_by_qn_suffix(store, effective_project, macro_qn, &suffix_nodes, + &suffix_count); + } + if (suffix_count == SKIP_ONE) { copy_node(&suffix_nodes[0], &node); cbm_store_free_nodes(suffix_nodes, suffix_count); diff --git a/src/pipeline/lsp_resolve.h b/src/pipeline/lsp_resolve.h index 666f9fd552..a8ba015f3d 100644 --- a/src/pipeline/lsp_resolve.h +++ b/src/pipeline/lsp_resolve.h @@ -920,24 +920,39 @@ static inline const cbm_gbuf_node_t *cbm_pipeline_lsp_target_node_policy( if (direct) { return direct; } - if (project_name && project_name[0]) { - size_t proj_len = strlen(project_name); - if (!(strncmp(callee_qn, project_name, proj_len) == 0 && callee_qn[proj_len] == '.')) { - size_t callee_len = strlen(callee_qn); - if (proj_len <= SIZE_MAX - callee_len - 2U) { - size_t prefixed_len = proj_len + 1U + callee_len; - char *prefixed_qn = (char *)malloc(prefixed_len + 1U); - if (prefixed_qn) { - memcpy(prefixed_qn, project_name, proj_len); - prefixed_qn[proj_len] = '.'; - memcpy(prefixed_qn + proj_len + 1U, callee_qn, callee_len + 1U); - const cbm_gbuf_node_t *prefixed = cbm_gbuf_find_by_qn(gbuf, prefixed_qn); - free(prefixed_qn); - if (prefixed) { - return prefixed; - } + /* Exact retries, in tie-rule order (CBM_MACRO_QN_SUFFIX): the definition + * first -- the QN as given, then with the project prefix -- and only when + * neither names a node the C-preprocessor macro of that name, in the same + * two spellings. The macro case is a prototype the C resolver knows whose + * only node in the graph is its rename macro (`extern T f(...);` plus + * `#define f f_v2`). One buffer, ".#macro", serves all + * three retries. */ + size_t proj_len = (project_name && project_name[0]) ? strlen(project_name) : 0; + size_t callee_len = strlen(callee_qn); + size_t suffix_len = sizeof(CBM_MACRO_QN_SUFFIX) - 1U; + bool add_prefix = proj_len > 0 && !(strncmp(callee_qn, project_name, proj_len) == 0 && + callee_qn[proj_len] == '.'); + if (proj_len <= SIZE_MAX - callee_len - suffix_len - 2U) { + char *buf = (char *)malloc(proj_len + 1U + callee_len + suffix_len + 1U); + if (buf) { + char *own = buf + proj_len + 1U; /* callee_qn inside the buffer */ + if (proj_len > 0) { + memcpy(buf, project_name, proj_len); + } + buf[proj_len] = '.'; + memcpy(own, callee_qn, callee_len + 1U); + const cbm_gbuf_node_t *hit = add_prefix ? cbm_gbuf_find_by_qn(gbuf, buf) : NULL; + if (!hit) { + memcpy(own + callee_len, CBM_MACRO_QN_SUFFIX, suffix_len + 1U); + hit = cbm_gbuf_find_by_qn(gbuf, own); + if (!hit && add_prefix) { + hit = cbm_gbuf_find_by_qn(gbuf, buf); } } + free(buf); + if (hit) { + return hit; + } } } if (!allow_tail_match) { diff --git a/src/pipeline/pass_definitions.c b/src/pipeline/pass_definitions.c index f3a2cd948c..c91b8a6c6f 100644 --- a/src/pipeline/pass_definitions.c +++ b/src/pipeline/pass_definitions.c @@ -203,6 +203,24 @@ static void append_json_string(char *buf, size_t bufsize, size_t *pos, const cha *pos = p; } +/* Append an already-serialized JSON value verbatim: ,"key":. Atomic like + * append_json_string. For `variants`, which extraction builds as a JSON array + * of objects (cbm.h). Twin of pass_parallel.c -- keep both in sync. */ +static void append_json_raw(char *buf, size_t bufsize, size_t *pos, const char *key, + const char *json) { + if (!json || json[0] == '\0') { + return; + } + size_t required = strlen(key) + strlen(json) + PD_JSON_FIELD_OVERHEAD; + if (*pos + required + PD_ESC_SPACE > bufsize) { + return; /* whole field would not fit — skip it atomically */ + } + int w = snprintf(buf + *pos, bufsize - *pos, ",\"%s\":%s", key, json); + if (w > 0 && (size_t)w < bufsize - *pos) { + *pos += (size_t)w; + } +} + /* Append a JSON array of strings: ,"key":["a","b","c"]. Atomic like * append_json_string: emitted only if the whole array fits. */ static void append_json_str_array(char *buf, size_t bufsize, size_t *pos, const char *key, @@ -286,6 +304,9 @@ static void build_def_props(char *buf, size_t bufsize, const CBMDefinition *def) } size_t pos = (size_t)n; append_json_string(buf, bufsize, &pos, "docstring", def->docstring); + /* Right after the docstring: the buffer is sized for exactly these two + * uncapped fields (pd_props_buf), so neither can be squeezed out. */ + append_json_raw(buf, bufsize, &pos, "variants", def->variants); append_json_string(buf, bufsize, &pos, "signature", def->signature); append_json_string(buf, bufsize, &pos, "return_type", def->return_type); append_json_string(buf, bufsize, &pos, "parent_class", def->parent_class); @@ -323,16 +344,21 @@ static void build_def_props(char *buf, size_t bufsize, const CBMDefinition *def) } /* A def's properties buffer: CBM_SZ_2K for every other field plus the whole - * serialized docstring field, which has no length cap (a field that does not - * fit is dropped whole). Returns `stack` for a def without a docstring, or - * when the larger buffer cannot be allocated. Twin of pass_parallel.c -- keep - * both in sync. */ + * serialized docstring and variants fields, which have no length cap (a field + * that does not fit is dropped whole). Returns `stack` for a def with neither, + * or when the larger buffer cannot be allocated. Twin of pass_parallel.c -- + * keep both in sync. */ static char *pd_props_buf(const CBMDefinition *def, char *stack, size_t *size) { - if (!def->docstring || !def->docstring[0]) { + size_t need = *size; + if (def->docstring && def->docstring[0]) { + need += strlen("docstring") + def_json_escaped_len(def->docstring) + PD_JSON_FIELD_OVERHEAD; + } + if (def->variants && def->variants[0]) { + need += strlen("variants") + strlen(def->variants) + PD_JSON_FIELD_OVERHEAD; + } + if (need == *size) { return stack; } - size_t need = - *size + strlen("docstring") + def_json_escaped_len(def->docstring) + PD_JSON_FIELD_OVERHEAD; char *buf = cbm_alloc(CBM_MEM_CLASS_GBUF_STRING, need); if (!buf) { return stack; diff --git a/src/pipeline/pass_parallel.c b/src/pipeline/pass_parallel.c index 0c12e6c4ab..03c065d96e 100644 --- a/src/pipeline/pass_parallel.c +++ b/src/pipeline/pass_parallel.c @@ -477,6 +477,24 @@ static void append_json_str_array(char *buf, size_t bufsize, size_t *pos, const *pos = p; } +/* Append an already-serialized JSON value verbatim: ,"key":. Atomic like + * append_json_string. For `variants`, which extraction builds as a JSON array + * of objects (cbm.h). Twin of pass_definitions.c -- keep both in sync. */ +static void append_json_raw(char *buf, size_t bufsize, size_t *pos, const char *key, + const char *json) { + if (!json || json[0] == '\0') { + return; + } + size_t required = strlen(key) + strlen(json) + PP_JSON_FIELD_OVERHEAD; + if (*pos + required + PP_ESC_SPACE > bufsize) { + return; /* whole field would not fit — skip it atomically */ + } + int w = snprintf(buf + *pos, bufsize - *pos, ",\"%s\":%s", key, json); + if (w > 0 && (size_t)w < bufsize - *pos) { + *pos += (size_t)w; + } +} + static void build_def_props(char *buf, size_t bufsize, const CBMDefinition *def) { /* Complexity/loop/recursion metrics are meaningful only for Function/Method. * Gate the block so the millions of Macro/Field/Variable/Class/Enum nodes @@ -513,6 +531,9 @@ static void build_def_props(char *buf, size_t bufsize, const CBMDefinition *def) } size_t pos = (size_t)n; append_json_string(buf, bufsize, &pos, "docstring", def->docstring); + /* Right after the docstring: the buffer is sized for exactly these two + * uncapped fields (pp_props_buf), so neither can be squeezed out. */ + append_json_raw(buf, bufsize, &pos, "variants", def->variants); append_json_string(buf, bufsize, &pos, "signature", def->signature); append_json_string(buf, bufsize, &pos, "return_type", def->return_type); append_json_string(buf, bufsize, &pos, "parent_class", def->parent_class); @@ -698,16 +719,21 @@ typedef struct { enum { PP_OVERSIZED_WARN_MAX = 32 }; /* A def's properties buffer: CBM_SZ_2K for every other field plus the whole - * serialized docstring field, which has no length cap (a field that does not - * fit is dropped whole). Returns `stack` for a def without a docstring, or - * when the larger buffer cannot be allocated. Twin of pass_definitions.c -- + * serialized docstring and variants fields, which have no length cap (a field + * that does not fit is dropped whole). Returns `stack` for a def with neither, + * or when the larger buffer cannot be allocated. Twin of pass_definitions.c -- * keep both in sync. */ static char *pp_props_buf(const CBMDefinition *def, char *stack, size_t *size) { - if (!def->docstring || !def->docstring[0]) { + size_t need = *size; + if (def->docstring && def->docstring[0]) { + need += strlen("docstring") + pp_json_escaped_len(def->docstring) + PP_JSON_FIELD_OVERHEAD; + } + if (def->variants && def->variants[0]) { + need += strlen("variants") + strlen(def->variants) + PP_JSON_FIELD_OVERHEAD; + } + if (need == *size) { return stack; } - size_t need = - *size + strlen("docstring") + pp_json_escaped_len(def->docstring) + PP_JSON_FIELD_OVERHEAD; char *buf = cbm_alloc(CBM_MEM_CLASS_GBUF_STRING, need); if (!buf) { return stack; diff --git a/src/pipeline/pipeline_internal.h b/src/pipeline/pipeline_internal.h index 28daa54346..ae741fbbea 100644 --- a/src/pipeline/pipeline_internal.h +++ b/src/pipeline/pipeline_internal.h @@ -806,8 +806,14 @@ int cbm_pipeline_build_fresh_semantic_manifest(cbm_pipeline_t *p, const char *pr cbm_file_hash_t **out, int *out_count); /* Compatibility contract persisted in coverage metadata. Increment when a - * graph/manifest semantic change makes prior exact-input indexes unsafe. */ -enum { CBM_SEMANTIC_INDEX_VERSION = 3 }; + * graph/manifest semantic change makes prior exact-input indexes unsafe. + * 4: C-family node identities changed. A preprocessor macro's QN ends in + * "#macro"; unscoped C/C++/Objective-C enumerators are `.` + * (the enum name is no longer a segment); typedef names, anonymous-enum + * constants and macro-prefixed functions are nodes; a bodyless + * `struct X` is no node. An index written before this holds the old + * QNs for every unchanged file, so it is rebuilt in full once. */ +enum { CBM_SEMANTIC_INDEX_VERSION = 4 }; typedef struct { cbm_gbuf_t *gbuf; diff --git a/src/pipeline/registry.c b/src/pipeline/registry.c index bdf6f5944d..5b5e79df08 100644 --- a/src/pipeline/registry.c +++ b/src/pipeline/registry.c @@ -947,7 +947,9 @@ void cbm_registry_add(cbm_registry_t *r, const char *name, const char *qualified const char *derived = simple_name(qualified_name); index_under_name(r, derived, owned_qn); /* '#' is a QN fence, and extract_defs.c's rust_cfg_qualified_name is the - * only thing in the tree that mints one today. A grammar that starts + * only thing in the tree that mints one on a REGISTERED symbol today (a C + * macro's QN ends in "#macro", but Macro is no registry label, so it never + * gets here). A grammar that starts * minting a '#' opts into this second key by doing so, whatever it means by * the fence: its symbols become reachable under the passed name as well, * and they share that name's bucket with everything else filed under it. */ diff --git a/src/store/store.c b/src/store/store.c index d66ef0ffbb..6680135444 100644 --- a/src/store/store.c +++ b/src/store/store.c @@ -421,7 +421,18 @@ static int init_schema(cbm_store_t *s) { "CASE WHEN properties LIKE '%\"docstring\"%' AND json_valid(properties) " \ "THEN json_extract(properties, '$.docstring') END" -enum { FTS_SQL_BUF = 512 }; +/* nodes_fts.qualified_name: the QN without the C-macro fence. A macro's QN is + * `.NAME#macro` (CBM_MACRO_QN_SUFFIX in internal/cbm/cbm.h), and `#` + * separates tokens, so the fence would add the word "macro" to every macro + * node a second time (the label column already says Macro). On a C repository + * that is thousands of rows outscoring the few nodes that carry the word in + * their NAME, enough to push those out of the BM25 candidate window. */ +#define FTS_QN_EXPR \ + "CASE WHEN label = 'Macro' AND substr(qualified_name, -6) = '#macro' " \ + "THEN substr(qualified_name, 1, length(qualified_name) - 6) " \ + "ELSE qualified_name END" + +enum { FTS_SQL_BUF = 768 }; /* Does nodes_fts carry the `body` column? A database created by an older * build has only the four identifier columns, and CREATE VIRTUAL TABLE IF NOT @@ -446,7 +457,7 @@ static int fts_backfill_try(cbm_store_t *s, const char *project, int64_t after_i char sql[FTS_SQL_BUF]; int n = snprintf(sql, sizeof(sql), "INSERT INTO nodes_fts (rowid, name, qualified_name, label, file_path%s)" - " SELECT id, %s, qualified_name, label, file_path%s FROM nodes%s;", + " SELECT id, %s, " FTS_QN_EXPR ", label, file_path%s FROM nodes%s;", with_body ? ", body" : "", camel ? "cbm_camel_split(name)" : "name", with_body ? ", " FTS_BODY_EXPR : "", project ? " WHERE project = ?1 AND id > ?2" : ""); diff --git a/tests/test_extraction.c b/tests/test_extraction.c index fe46d7246f..55e2031440 100644 --- a/tests/test_extraction.c +++ b/tests/test_extraction.c @@ -183,12 +183,13 @@ TEST(extract_ts_factory_object_methods_issue341) { * * The bound is derived, not tuned. Measured on this source, total_alloc was * 365984 before the scratch arena and is 87456 after, exactly that difference. - * #1916 added two pointers to CBMDefinition (240 -> 256 bytes), so it now - * measures 87968: 8192 is the defs item array at GROW_ARRAY's starting - * capacity of 32 times sizeof(CBMDefinition) 256, and the other 79776 is - * everything else this file's extraction interns; none of it is traversal - * scratch. So the bound sits above 87968 with room and a factor of four below - * the 365984 the scratch stacks cost. + * #1916 added two pointers to CBMDefinition (240 -> 256 bytes) and the C + * `variants` list one more (264), so it now measures 88224: 8448 is the defs + * item array at GROW_ARRAY's starting capacity of 32 times + * sizeof(CBMDefinition) 264, and the other 79776 is everything else this + * file's extraction interns; none of it is traversal scratch. So the bound + * sits above 88224 with room and a factor of four below the 365984 the + * scratch stacks cost. * * It is a byte budget, not a proof of lifetime; that is * extract_traversal_stacks_come_from_ctx_scratch_issue2010 in test_mem.c. */ @@ -6658,6 +6659,740 @@ TEST(extract_c_export_macro_recovery_issue1989) { PASS(); } +/* ── PR C1: C graph defects behind the C/Doxygen doc-link audit ──────── */ + +static const CBMDefinition *c1_def(CBMFileResult *r, const char *label, const char *name) { + for (int i = 0; i < r->defs.count; i++) { + const CBMDefinition *d = &r->defs.items[i]; + if (d->label && d->name && strcmp(d->label, label) == 0 && strcmp(d->name, name) == 0) { + return d; + } + } + return NULL; +} + +/* Caller QN of the first raw-pass call to `callee`, or "" when there is none. */ +static const char *c1_call_scope(CBMFileResult *r, const char *callee) { + for (int i = 0; i < r->calls.count; i++) { + const CBMCall *c = &r->calls.items[i]; + if (c->callee_name && strcmp(c->callee_name, callee) == 0 && + c->source_origin == CBM_SOURCE_ORIGIN_RAW) { + return c->enclosing_func_qn ? c->enclosing_func_qn : ""; + } + } + return ""; +} + +/* Caller QN of the first C-LSP resolution whose callee QN ends in ".". */ +static const char *c1_lsp_caller(CBMFileResult *r, const char *leaf) { + size_t ll = strlen(leaf); + for (int i = 0; i < r->resolved_calls.count; i++) { + const CBMResolvedCall *rc = &r->resolved_calls.items[i]; + size_t cl = rc->callee_qn ? strlen(rc->callee_qn) : 0; + if (cl > ll && rc->callee_qn[cl - ll - 1] == '.' && + strcmp(rc->callee_qn + cl - ll, leaf) == 0) { + return rc->caller_qn ? rc->caller_qn : ""; + } + } + return ""; +} + +/* xxhash.h shape: `XXH_PUBLIC_API name(...)` has one leading identifier more than + * the grammar admits. tree-sitter-cpp (every .h is C++) recovers it as `RetT::name` with a + * zero-width MISSING "::" (a Method of the return type), with a pointer return as + * `RetT::*name` (no name at all: the function was dropped), or keeps RetT as the name and + * puts the real one in an ERROR before the parameters. The implementation sits behind a + * guard the preprocessed pass does not take, as in xxhash.h, so only the raw parse sees + * it. Def QN, call-scope QN and C-LSP caller QN must all name the real function. */ +static const char *C1_MACRO_PREFIX_SRC = + "#if defined(XXH_IMPLEMENTATION)\n" + "XXH_PUBLIC_API XXH_errorcode XXH32_freeState(XXH32_state_t* statePtr)\n" + "{\n" + " XXH_free(statePtr);\n" + " return XXH_OK;\n" + "}\n" + "\n" + "XXH_PUBLIC_API XXH32_state_t* XXH32_createState(void)\n" + "{\n" + " return (XXH32_state_t*)XXH_malloc(sizeof(XXH32_state_t));\n" + "}\n" + "\n" + "XXH_PUBLIC_API XXH32_hash_t XXH32 (const void* input, size_t len, XXH32_hash_t seed)\n" + "{\n" + " return XXH32_endian_align(input, len, seed);\n" + "}\n" + "#endif\n"; + +TEST(extract_cpp_macro_prefixed_function_names_c1) { + CBMFileResult *r = extract(C1_MACRO_PREFIX_SRC, CBM_LANG_CPP, "p", "xxhash.h"); + ASSERT_NOT_NULL(r); + char want[256]; + + const CBMDefinition *free_state = c1_def(r, "Function", "XXH32_freeState"); + ASSERT_NOT_NULL(free_state); + ASSERT_EQ(count_defs_named(r, "Method", "XXH32_freeState"), 0); + snprintf(want, sizeof(want), "%s.XXH32_freeState", r->module_qn); + ASSERT_STR_EQ(free_state->qualified_name, want); + ASSERT_STR_EQ(c1_call_scope(r, "XXH_free"), want); + + const CBMDefinition *create = c1_def(r, "Function", "XXH32_createState"); + ASSERT_NOT_NULL(create); + ASSERT_EQ((int)create->start_line, 8); + ASSERT_EQ((int)create->end_line, 11); + ASSERT_NOT_NULL(create->signature); + + const CBMDefinition *xxh32 = c1_def(r, "Function", "XXH32"); + ASSERT_NOT_NULL(xxh32); + ASSERT_FALSE(has_def_any(r, "XXH32_hash_t")); + ASSERT_STR_EQ(c1_call_scope(r, "XXH32_endian_align"), xxh32->qualified_name); + cbm_free_result(r); + PASS(); +} + +/* The C-LSP walks declarators on its own (c_process_function); its caller QN must + * name the recovered function too, or the exact-equality LSP join drops the + * type-aware resolution of every call inside it (the call falls back to a textual + * strategy). jemalloc-style prefix: not an export-macro candidate, so the + * preprocessed pass keeps the same misparse and cannot mask the raw one. */ +TEST(extract_cpp_macro_prefixed_lsp_caller_c1) { + CBMFileResult *r = extract("static int je_helper(int v) { return v; }\n" + "JE_ALWAYS_INLINE je_errorcode je_reset(je_state_t *st)\n" + "{\n" + " return je_helper(1);\n" + "}\n" + "JE_ALWAYS_INLINE je_hash_t je_hash (const void *in, size_t n)\n" + "{\n" + " return je_helper(2);\n" + "}\n", + CBM_LANG_CPP, "p", "je.h"); + ASSERT_NOT_NULL(r); + const CBMDefinition *reset = c1_def(r, "Function", "je_reset"); + const CBMDefinition *hash = c1_def(r, "Function", "je_hash"); + ASSERT_NOT_NULL(reset); + ASSERT_NOT_NULL(hash); + int from_reset = 0; + int from_hash = 0; + for (int i = 0; i < r->resolved_calls.count; i++) { + const char *caller = r->resolved_calls.items[i].caller_qn; + from_reset += caller && strcmp(caller, reset->qualified_name) == 0; + from_hash += caller && strcmp(caller, hash->qualified_name) == 0; + } + ASSERT_GT(from_reset, 0); + ASSERT_GT(from_hash, 0); + ASSERT_STR_EQ(c1_lsp_caller(r, "je_helper"), reset->qualified_name); + cbm_free_result(r); + PASS(); +} + +/* The same recovery in the C grammar (lua's LUA_API): named after the return type + * before. A function-like macro invocation that only looks like a declarator + * (`... TEST_BEGIN (test_alignment) {`) has an untyped "parameter" no definition can + * have, so its macro name is never adopted as the function name. */ +TEST(extract_c_macro_prefixed_function_name_c1) { + CBMFileResult *r = + extract("#if defined(LUA_CORE)\n" + "LUA_API lua_CFunction lua_atpanic (lua_State *L, lua_CFunction panicf) {\n" + " lua_CFunction old = G(L)->panic;\n" + " return old;\n" + "}\n" + "\n" + "XXH_TEST_API XXH32_hash_t TEST_BEGIN (test_alignment) {\n" + " helper(1);\n" + "}\n" + "#endif\n", + CBM_LANG_C, "p", "lapi.c"); + ASSERT_NOT_NULL(r); + ASSERT_NOT_NULL(c1_def(r, "Function", "lua_atpanic")); + ASSERT_FALSE(has_def_any(r, "lua_CFunction")); + ASSERT_FALSE(has_def_any(r, "TEST_BEGIN")); + cbm_free_result(r); + PASS(); +} + +/* `struct X` / `enum X` / `class X` without a body is a reference, not a definition: + * no Class/Enum def, so nothing competes with the real definition for its QN (a later + * or smaller-path reference used to take the node: curl 132 structs, kernel + * task_struct, redis redisServer) and no phantom type node appears in every file that + * merely mentions the type (`struct timeval`). */ +TEST(extract_c_tag_reference_is_not_a_definition_c1) { + CBMFileResult *r = extract("struct node {\n" + " int value;\n" + " struct node *next;\n" + "};\n" + "struct node *head;\n" + "struct timeval tv;\n" + "enum color { RED, GREEN };\n" + "enum color paint(enum color c);\n" + "void walk(struct node *n);\n" + "struct opaque;\n", + CBM_LANG_C, "p", "list.c"); + ASSERT_NOT_NULL(r); + ASSERT_EQ(count_defs_named(r, "Class", "node"), 1); + const CBMDefinition *node = c1_def(r, "Class", "node"); + ASSERT_EQ((int)node->start_line, 1); + ASSERT_EQ((int)node->end_line, 4); + ASSERT_EQ(count_defs_named(r, "Enum", "color"), 1); + ASSERT_EQ((int)c1_def(r, "Enum", "color")->start_line, 7); + ASSERT_FALSE(has_def_any(r, "timeval")); + ASSERT_FALSE(has_def_any(r, "opaque")); + cbm_free_result(r); + + /* .h is parsed as C++: a forward declaration before the definition. */ + r = extract("class Widget;\n" + "struct Config;\n" + "Widget *make_widget(struct Config *cfg);\n" + "class Widget {\n" + " int size;\n" + "};\n", + CBM_LANG_CPP, "p", "widget.h"); + ASSERT_NOT_NULL(r); + ASSERT_EQ(count_defs_named(r, "Class", "Widget"), 1); + ASSERT_EQ((int)c1_def(r, "Class", "Widget")->start_line, 4); + ASSERT_FALSE(has_def_any(r, "Config")); + cbm_free_result(r); + PASS(); +} + +/* tree-sitter can give up on a whole region and leave the definition head as loose + * ERROR tokens (kernel include/linux/sched.h: `struct task_struct {` at 826 is never a + * struct_specifier). Its only nodes were then the references (642 won); with + * references no longer minting defs the head itself must be recovered, spanning to + * its matching brace counted over the first branch of each #if group. */ +TEST(extract_c_struct_head_in_error_region_recovered_c1) { + CBMFileResult *r = extract("#ifndef SCHED_H\n" + "#define SCHED_H\n" + "struct task_struct;\n" + "typedef struct task_struct *(*pick_f)(int cpu);\n" + "\n" + "/* Per-task scheduler state. */\n" + "struct task_struct {\n" + "\tint state;\n" + "#ifdef CONFIG_SMP\n" + "\tstruct {\n" + "#else\n" + "\tunion {\n" + "#endif\n" + "\t\tint on_cpu;\n" + "\t};\n" + "\tint prio;\n" + "};\n" + "#endif\n", + CBM_LANG_CPP, "p", "sched.h"); + ASSERT_NOT_NULL(r); + ASSERT_EQ(count_defs_named(r, "Class", "task_struct"), 1); + const CBMDefinition *ts = c1_def(r, "Class", "task_struct"); + ASSERT_EQ((int)ts->start_line, 7); + ASSERT_EQ((int)ts->end_line, 17); + ASSERT_NOT_NULL(c1_def(r, "Type", "pick_f")); + cbm_free_result(r); + PASS(); +} + +/* A stray `;` after an if-block leaves the following `else if (...) {...}` without its + * `if`; inside a function the raw parse already lost (whole-file guard + split brace), + * error recovery reads it as a DEFINITION: type `else`, name `if`. A C keyword is never + * a function name (curl lib/vtls/openssl.c:4242, redis deps/tre/lib/tre-parse.c:1315). */ +TEST(extract_c_keyword_is_never_a_function_name_c1) { + CBMFileResult *r = extract("#ifdef USE_OPENSSL\n" + "static int check(int lib, int reason)\n" + "{\n" + " int result = 0;\n" + " if(lib == 1) {\n" + "#ifndef HAVE_BORINGSSL_LIKE\n" + " result = 1; {\n" + "#else\n" + " result = 2; {\n" + "#endif\n" + " }\n" + " };\n" + " else if(reason == 2) {\n" + " result = trace_retry(reason);\n" + " }\n" + " return result;\n" + "}\n" + "#endif\n", + CBM_LANG_C, "p", "openssl.c"); + ASSERT_NOT_NULL(r); + ASSERT_FALSE(has_def_any(r, "if")); + ASSERT_TRUE(has_call(r, "trace_retry")); + cbm_free_result(r); + PASS(); +} + +/* Preprocessor-split code the raw parse cannot read: `#ifndef X (a)) { #else (b)) { + * #endif` leaves an extra `{`, the file-level ERROR swallows the function header and + * the function gets no node. The #961 rescue re-parses the preprocessed text, but a + * whole-file `#ifdef USE_OPENSSL` with no build defines leaves it nothing to parse. + * The first-branch projection restores it (curl ossl_connect_step2, + * Curl_ossl_ctx_init, mbed_configure_ssl, ...: all 11 curl misses). */ +TEST(extract_c_function_lost_to_split_branches_restored_c1) { + CBMFileResult *r = extract("#ifdef USE_OPENSSL\n" + "static int ossl_step(int lib, int reason)\n" + "{\n" + " int result = 0;\n" + " if(lib == 1) {\n" + " result = 1;\n" + " }\n" + " else if((lib == 2) &&\n" + "#ifndef HAVE_BORINGSSL_LIKE\n" + " (reason == 3)) {\n" + "#else\n" + " (reason == 4)) {\n" + "#endif\n" + " result = trace_retry(reason);\n" + " }\n" + " return result;\n" + "}\n" + "\n" + "int ossl_after(void)\n" + "{\n" + " return ossl_step(1, 2);\n" + "}\n" + "#endif\n", + CBM_LANG_C, "p", "openssl.c"); + ASSERT_NOT_NULL(r); + const CBMDefinition *step = c1_def(r, "Function", "ossl_step"); + ASSERT_NOT_NULL(step); + ASSERT_EQ((int)step->start_line, 2); + ASSERT_EQ((int)step->end_line, 17); + ASSERT_EQ(count_defs_named(r, "Function", "ossl_after"), 1); + cbm_free_result(r); + PASS(); +} + +/* C typedefs carry no `name` field (the alias lives in the declarators), so no typedef + * ever became a node and an anonymous `typedef struct {...} Name;` lost its fields and + * enum constants too. xxhash: `@ref XXH32_state_t` could only resolve to the namespace + * rename macro `#define XXH32_state_t XXH_IPREF(XXH32_state_t)`. */ +TEST(extract_c_typedef_names_are_definitions_c1) { + CBMFileResult *r = extract("typedef struct XXH32_state_s XXH32_state_t;\n" + "typedef struct {\n" + " unsigned char digest[4];\n" + "} XXH32_canonical_t;\n" + "typedef enum { XXH_OK = 0, XXH_ERROR = 1 } XXH_errorcode;\n" + "typedef unsigned int XXH32_hash_t;\n" + "typedef int (*cmp_fn)(const void *a, const void *b);\n" + "typedef struct opaque_s opaque_s;\n" + "typedef struct node node;\n" + "struct node { int v; };\n" + "typedef struct named_s { int y; } named_t;\n" + "typedef struct same { int z; } same;\n", + CBM_LANG_C, "p", "xxhash.c"); + ASSERT_NOT_NULL(r); + char want[256]; + ASSERT_NOT_NULL(c1_def(r, "Type", "XXH32_state_t")); + const CBMDefinition *canon = c1_def(r, "Class", "XXH32_canonical_t"); + ASSERT_NOT_NULL(canon); + ASSERT_EQ((int)canon->start_line, 2); + snprintf(want, sizeof(want), "%s.XXH32_canonical_t.digest", r->module_qn); + ASSERT_TRUE(has_def_qn(r, want)); + ASSERT_NOT_NULL(c1_def(r, "Enum", "XXH_errorcode")); + snprintf(want, sizeof(want), "%s.XXH_OK", r->module_qn); + ASSERT_TRUE(has_def_qn(r, want)); + ASSERT_NOT_NULL(c1_def(r, "Type", "XXH32_hash_t")); + ASSERT_NOT_NULL(c1_def(r, "Type", "cmp_fn")); + /* identity alias of an incomplete struct: often the only declaration of an + * opaque handle type */ + ASSERT_NOT_NULL(c1_def(r, "Type", "opaque_s")); + /* identity alias of a struct defined in the same file: the definition owns it */ + ASSERT_EQ(count_defs_named(r, "Class", "node"), 1); + ASSERT_EQ(count_defs_named(r, "Type", "node"), 0); + ASSERT_NOT_NULL(c1_def(r, "Class", "named_s")); + ASSERT_NOT_NULL(c1_def(r, "Type", "named_t")); + ASSERT_EQ(count_defs_named(r, "Class", "same"), 1); + ASSERT_EQ(count_defs_named(r, "Type", "same"), 0); + cbm_free_result(r); + PASS(); +} + +static const CBMDefinition *c1_def_qn(CBMFileResult *r, const char *qn) { + for (int i = 0; i < r->defs.count; i++) { + const CBMDefinition *d = &r->defs.items[i]; + if (d->qualified_name && strcmp(d->qualified_name, qn) == 0) { + return d; + } + } + return NULL; +} + +/* The def with this label and name that starts on `line`. */ +static const CBMDefinition *c1_def_at(CBMFileResult *r, const char *label, const char *name, + int line) { + for (int i = 0; i < r->defs.count; i++) { + const CBMDefinition *d = &r->defs.items[i]; + if (d->label && d->name && strcmp(d->label, label) == 0 && strcmp(d->name, name) == 0 && + (int)d->start_line == line) { + return d; + } + } + return NULL; +} + +/* A preprocessor macro's QN is `.#macro` (CBM_MACRO_QN_SUFFIX); name and + * label stay. C keeps macros apart from ordinary identifiers, so a macro and the + * function, type or enumerator of the same name are two entities and need two QNs: the + * namespace rename `#define XXH32 XXH_NAME2(XXH_NAMESPACE, XXH32)` held the QN of the + * function XXH32 (110 of 174 audited xxhash doc links ended on such a macro), and the + * #else stand-in `#define match_proxy(a) 1` replaced the function it stands in for + * (curl, 41 functions). Every language routed through the #define path is covered, + * and in each of them a macro defined twice carries `variants`: the namespace and the + * list follow one language predicate. */ +TEST(extract_c_macro_qn_is_fenced_c1) { + static const struct { + CBMLanguage lang; + const char *file; + } cases[] = {{CBM_LANG_C, "cfg.c"}, {CBM_LANG_CPP, "cfg.h"}, + {CBM_LANG_OBJC, "cfg.m"}, {CBM_LANG_CUDA, "cfg.cu"}, + {CBM_LANG_GLSL, "cfg.glsl"}, {CBM_LANG_ISPC, "cfg.ispc"}}; + for (size_t i = 0; i < sizeof(cases) / sizeof(cases[0]); i++) { + CBMFileResult *r = extract("#define MAX_LEN 10\n" + "#define SQUARE(x) ((x) * (x))\n" + "#define MAX_LEN 20\n", + cases[i].lang, "p", cases[i].file); + ASSERT_NOT_NULL(r); + char want[256]; + const CBMDefinition *max_len = c1_def(r, "Macro", "MAX_LEN"); + ASSERT_NOT_NULL(max_len); + snprintf(want, sizeof(want), "%s.MAX_LEN" CBM_MACRO_QN_SUFFIX, r->module_qn); + ASSERT_STR_EQ(max_len->qualified_name, want); + const CBMDefinition *square = c1_def(r, "Macro", "SQUARE"); + ASSERT_NOT_NULL(square); + snprintf(want, sizeof(want), "%s.SQUARE#macro", r->module_qn); + ASSERT_STR_EQ(square->qualified_name, want); + /* the plain QN is free for a definition of that name */ + snprintf(want, sizeof(want), "%s.MAX_LEN", r->module_qn); + ASSERT_NULL(c1_def_qn(r, want)); + const CBMDefinition *redefined = c1_def_at(r, "Macro", "MAX_LEN", 3); + ASSERT_NOT_NULL(redefined); + ASSERT_NOT_NULL(redefined->variants); + cbm_free_result(r); + } + + CBMFileResult *r = extract("#ifdef USE_PROXY\n" + "int match_proxy(int a) { return a; }\n" + "#else\n" + "#define match_proxy(a) 1\n" + "#endif\n", + CBM_LANG_C, "p", "proxy.c"); + ASSERT_NOT_NULL(r); + const CBMDefinition *fn = c1_def(r, "Function", "match_proxy"); + const CBMDefinition *mac = c1_def(r, "Macro", "match_proxy"); + ASSERT_NOT_NULL(fn); + ASSERT_NOT_NULL(mac); + char want[256]; + snprintf(want, sizeof(want), "%s.match_proxy", r->module_qn); + ASSERT_STR_EQ(fn->qualified_name, want); + snprintf(want, sizeof(want), "%s.match_proxy#macro", r->module_qn); + ASSERT_STR_EQ(mac->qualified_name, want); + cbm_free_result(r); + PASS(); +} + +/* The enumerators of a C enum live in the scope that holds the enum, so their QN is + * `.`: the enum's own name is not a segment. That is how a reference + * spells them (`XXH_OK`, never `XXH_errorcode.XXH_OK`), and it gives the constants of + * an anonymous enum, which has no name to nest under, a QN at all. `parent_class` + * (the enum's QN) keeps the membership; an anonymous enum has no node to point at. */ +TEST(extract_c_enumerators_are_flat_c1) { + CBMFileResult *r = extract("enum color { RED, GREEN = 2 };\n" + "enum { ANON_A, ANON_B };\n" + "typedef enum { TD_A, TD_B } td_t;\n", + CBM_LANG_C, "p", "e.c"); + ASSERT_NOT_NULL(r); + char qn[256]; + char parent[256]; + + snprintf(qn, sizeof(qn), "%s.RED", r->module_qn); + snprintf(parent, sizeof(parent), "%s.color", r->module_qn); + const CBMDefinition *red = c1_def_qn(r, qn); + ASSERT_NOT_NULL(red); + ASSERT_STR_EQ(red->label, "Variable"); + ASSERT_NOT_NULL(red->parent_class); + ASSERT_STR_EQ(red->parent_class, parent); + snprintf(qn, sizeof(qn), "%s.GREEN", r->module_qn); + ASSERT_NOT_NULL(c1_def_qn(r, qn)); + snprintf(qn, sizeof(qn), "%s.color.RED", r->module_qn); + ASSERT_NULL(c1_def_qn(r, qn)); + ASSERT_NOT_NULL(c1_def_qn(r, parent)); /* the Enum itself keeps its QN */ + + snprintf(qn, sizeof(qn), "%s.ANON_A", r->module_qn); + const CBMDefinition *anon = c1_def_qn(r, qn); + ASSERT_NOT_NULL(anon); + ASSERT_STR_EQ(anon->label, "Variable"); + ASSERT_NULL(anon->parent_class); + snprintf(qn, sizeof(qn), "%s.ANON_B", r->module_qn); + ASSERT_NOT_NULL(c1_def_qn(r, qn)); + + snprintf(qn, sizeof(qn), "%s.TD_A", r->module_qn); + snprintf(parent, sizeof(parent), "%s.td_t", r->module_qn); + const CBMDefinition *td = c1_def_qn(r, qn); + ASSERT_NOT_NULL(td); + ASSERT_NOT_NULL(td->parent_class); + ASSERT_STR_EQ(td->parent_class, parent); + ASSERT_EQ(count_defs_named(r, "Variable", "TD_A"), 1); + cbm_free_result(r); + PASS(); +} + +/* C++: an unscoped enum's enumerators belong to the enclosing scope (module, namespace + * or class), whatever nesting holds the enum; `enum class` and `enum struct` are + * scopes of their own and keep `.`. */ +TEST(extract_cpp_enum_scoping_c1) { + CBMFileResult *r = extract("enum Color { RED, GREEN };\n" + "namespace gfx {\n" + "enum Blend { ADD, MULTIPLY };\n" + "class Brush {\n" + "public:\n" + " enum Shape { ROUND, SQUARE };\n" + "};\n" + "}\n" + "enum class Mode { FAST, SLOW };\n" + "enum struct Level { LOW, HIGH };\n", + CBM_LANG_CPP, "p", "shapes.hpp"); + ASSERT_NOT_NULL(r); + static const struct { + const char *qn_tail; + const char *parent_tail; + } want[] = {{"RED", "Color"}, + {"GREEN", "Color"}, + {"gfx.ADD", "gfx.Blend"}, + {"gfx.MULTIPLY", "gfx.Blend"}, + {"gfx.Brush.ROUND", "gfx.Brush.Shape"}, + {"gfx.Brush.SQUARE", "gfx.Brush.Shape"}, + {"Mode.FAST", "Mode"}, + {"Mode.SLOW", "Mode"}, + {"Level.LOW", "Level"}, + {"Level.HIGH", "Level"}}; + for (size_t i = 0; i < sizeof(want) / sizeof(want[0]); i++) { + char qn[256]; + char parent[256]; + snprintf(qn, sizeof(qn), "%s.%s", r->module_qn, want[i].qn_tail); + snprintf(parent, sizeof(parent), "%s.%s", r->module_qn, want[i].parent_tail); + const CBMDefinition *d = c1_def_qn(r, qn); + if (!d) { + fprintf(stderr, " [c1] no def with QN %s\n", qn); + } + ASSERT_NOT_NULL(d); + ASSERT_STR_EQ(d->label, "Variable"); + ASSERT_NOT_NULL(d->parent_class); + ASSERT_STR_EQ(d->parent_class, parent); + } + char qn[256]; + snprintf(qn, sizeof(qn), "%s.Color.RED", r->module_qn); + ASSERT_NULL(c1_def_qn(r, qn)); + snprintf(qn, sizeof(qn), "%s.FAST", r->module_qn); + ASSERT_NULL(c1_def_qn(r, qn)); + cbm_free_result(r); + PASS(); +} + +/* A named and an anonymous enum inside a struct, and an anonymous one inside a union: + * one C-looking source, read under each C-family language by the two tests below. */ +static const char *const C1_ENUM_IN_STRUCT_SRC = "struct conn {\n" + " enum state { ST_IDLE, ST_BUSY } st;\n" + " enum { KIND_A, KIND_B } kind;\n" + " int fd;\n" + "};\n" + "union box {\n" + " enum { BOX_A, BOX_B } tag;\n" + " int v;\n" + "};\n"; + +/* Is `[.].` the ONE Variable def of that name, with parent_class + * `.` (NULL: no parent_class)? */ +static bool c1_enumerator_at(CBMFileResult *r, const char *scope, const char *name, + const char *parent_tail) { + char qn[256]; + snprintf(qn, sizeof(qn), "%s%s%s.%s", r->module_qn, scope[0] ? "." : "", scope, name); + const CBMDefinition *d = c1_def_qn(r, qn); + if (!d || !d->label || strcmp(d->label, "Variable") != 0 || + count_defs_named(r, "Variable", name) != 1) { + const CBMDefinition *other = c1_def(r, "Variable", name); + fprintf(stderr, " [c1] want the one Variable %s, the Variable of that name is %s\n", qn, + other ? other->qualified_name : "(none)"); + return false; + } + if (!parent_tail) { + return d->parent_class == NULL; + } + char parent[256]; + snprintf(parent, sizeof(parent), "%s.%s", r->module_qn, parent_tail); + return d->parent_class && strcmp(d->parent_class, parent) == 0; +} + +/* C has no struct scope for ordinary identifiers: an enum declared inside a struct or + * union puts its enumerators in the scope around it, and the code names them + * unqualified (curl lib/cf-h1-proxy.c: `ts->keepon = KEEPON_CONNECT;`). In C and + * Objective-C the struct is therefore no segment of a flat enumerator's QN, named + * enum or anonymous; `parent_class` still names a named enum. */ +TEST(extract_c_enum_inside_struct_is_file_scope_c1) { + static const struct { + CBMLanguage lang; + const char *file; + } cases[] = {{CBM_LANG_C, "link.c"}, {CBM_LANG_OBJC, "link.m"}}; + for (size_t i = 0; i < sizeof(cases) / sizeof(cases[0]); i++) { + CBMFileResult *r = extract(C1_ENUM_IN_STRUCT_SRC, cases[i].lang, "p", cases[i].file); + ASSERT_NOT_NULL(r); + ASSERT_TRUE(c1_enumerator_at(r, "", "ST_IDLE", "conn.state")); + ASSERT_TRUE(c1_enumerator_at(r, "", "ST_BUSY", "conn.state")); + ASSERT_TRUE(c1_enumerator_at(r, "", "KIND_A", NULL)); + ASSERT_TRUE(c1_enumerator_at(r, "", "KIND_B", NULL)); + ASSERT_TRUE(c1_enumerator_at(r, "", "BOX_A", NULL)); + ASSERT_TRUE(c1_enumerator_at(r, "", "BOX_B", NULL)); + cbm_free_result(r); + } + PASS(); +} + +/* C++ does scope them: `conn::ST_IDLE`. The enclosing class stays the scope in C++ and + * CUDA -- and so in every .h, which is parsed as C++ -- for the same source. This pins + * the other side of the rule above: the C treatment must not reach C++. */ +TEST(extract_cpp_enum_inside_class_keeps_class_scope_c1) { + static const struct { + CBMLanguage lang; + const char *file; + } cases[] = {{CBM_LANG_CPP, "link.h"}, {CBM_LANG_CUDA, "link.cu"}}; + for (size_t i = 0; i < sizeof(cases) / sizeof(cases[0]); i++) { + CBMFileResult *r = extract(C1_ENUM_IN_STRUCT_SRC, cases[i].lang, "p", cases[i].file); + ASSERT_NOT_NULL(r); + ASSERT_TRUE(c1_enumerator_at(r, "conn", "ST_IDLE", "conn.state")); + ASSERT_TRUE(c1_enumerator_at(r, "conn", "ST_BUSY", "conn.state")); + ASSERT_TRUE(c1_enumerator_at(r, "conn", "KIND_A", NULL)); + ASSERT_TRUE(c1_enumerator_at(r, "conn", "KIND_B", NULL)); + ASSERT_TRUE(c1_enumerator_at(r, "box", "BOX_A", NULL)); + ASSERT_TRUE(c1_enumerator_at(r, "box", "BOX_B", NULL)); + cbm_free_result(r); + } + PASS(); +} + +/* A macro inside an enumerator list is not enumerator grammar; the parser keeps every + * bare identifier it can as an `enumerator` and wraps the rest in ERROR nodes. curl.h + * declares its options as `CURLOPT(CURLOPT_URL, CURLOPTTYPE_STRINGPOINT, 2),`: the + * second argument came back as a constant 115 times and, flattened, took the plain QN + * of the macro `CURLOPTTYPE_LONG`. `X CURL_DEPRECATED(7.55.0, "Use Y") = BASE + 16,` + * turned BASE and the words of the message into constants. Only the enumerator that + * opens a slot is a constant; the comment above the macro call documents it. */ +TEST(extract_c_enum_macro_wrapped_list_c1) { + CBMFileResult *r = extract("typedef enum {\n" + " /* doc of the first option */\n" + " OPT(OPT_FIRST, TYPE_LONG, 1),\n" + " OPT(OPT_SECOND, TYPE_STRING, 2),\n" + " OPT_THIRD DEPRECATED(1.0, \"Use OPT_FIRST instead\")\n" + " = BASE_LONG + 3,\n" + " OPT_LAST\n" + "} option_t;\n", + CBM_LANG_C, "p", "opts.c"); + ASSERT_NOT_NULL(r); + static const char *const real[] = {"OPT_FIRST", "OPT_SECOND", "OPT_THIRD", "OPT_LAST"}; + for (size_t i = 0; i < sizeof(real) / sizeof(real[0]); i++) { + ASSERT_EQ(count_defs_named(r, "Variable", real[i]), 1); + } + static const char *const bogus[] = {"TYPE_LONG", "TYPE_STRING", "BASE_LONG", + "Use", "instead", "OPT"}; + for (size_t i = 0; i < sizeof(bogus) / sizeof(bogus[0]); i++) { + ASSERT_EQ(count_defs_named(r, "Variable", bogus[i]), 0); + } + const CBMDefinition *first = c1_def(r, "Variable", "OPT_FIRST"); + ASSERT_NOT_NULL(first->docstring); + ASSERT_NOT_NULL(strstr(first->docstring, "doc of the first option")); + ASSERT_NULL(c1_def(r, "Variable", "OPT_SECOND")->docstring); + cbm_free_result(r); + PASS(); +} + +/* One file may define a QN several times under one label: both branches of an #if, a + * macro redefined per platform. The graph keeps one node per QN (the last by start + * line), and before this the other definitions left no trace (curl 123 groups, redis + * 87). The definition the graph keeps carries `variants`: every span of the group, + * the kept one included, sorted by start line. A name defined once carries nothing. */ +TEST(extract_c_variants_lists_same_file_duplicates_c1) { + CBMFileResult *r = extract("#if A\n" + "int pick(int a) { return a; }\n" + "#else\n" + "int pick(int a) {\n" + " return a + 1;\n" + "}\n" + "#endif\n" + "#ifdef B\n" + "#define LIMIT 1\n" + "#else\n" + "#define LIMIT 2\n" + "#endif\n" + "int once(void) { return 0; }\n", + CBM_LANG_C, "p", "v.c"); + ASSERT_NOT_NULL(r); + ASSERT_EQ(count_defs_named(r, "Function", "pick"), 2); + const CBMDefinition *kept = c1_def_at(r, "Function", "pick", 4); + const CBMDefinition *other = c1_def_at(r, "Function", "pick", 2); + ASSERT_NOT_NULL(kept); + ASSERT_NOT_NULL(other); + ASSERT_NOT_NULL(kept->variants); + ASSERT_STR_EQ(kept->variants, "[{\"start\":2,\"end\":2},{\"start\":4,\"end\":6}]"); + ASSERT_NULL(other->variants); /* one carrier per group: the node the graph keeps */ + + const CBMDefinition *lim_a = c1_def_at(r, "Macro", "LIMIT", 9); + const CBMDefinition *lim_b = c1_def_at(r, "Macro", "LIMIT", 11); + ASSERT_NOT_NULL(lim_a); + ASSERT_NOT_NULL(lim_b); + char want[128]; + snprintf(want, sizeof(want), "[{\"start\":%u,\"end\":%u},{\"start\":%u,\"end\":%u}]", + lim_a->start_line, lim_a->end_line, lim_b->start_line, lim_b->end_line); + ASSERT_NOT_NULL(lim_b->variants); + ASSERT_STR_EQ(lim_b->variants, want); + ASSERT_NULL(lim_a->variants); + + ASSERT_NULL(c1_def(r, "Function", "once")->variants); + cbm_free_result(r); + PASS(); +} + +/* `variants` describes conditional compilation, so only the C-preprocessor languages + * get it (cbm_is_c_preprocessor_lang, the six of the test above). A repeated QN + * elsewhere is another matter: a JSON file repeats one key name thousands of times + * (redis src/commands: 1,618 nodes, one of them with a 75-span list), a Python module + * may rebind a function. Both fixtures DO repeat a (QN, label) on DIFFERENT lines + * (two defs on one span count as one and would get no list anyway), which is asserted + * so the guard cannot go inert. */ +TEST(extract_variants_only_in_c_preprocessor_languages_c1) { + static const struct { + CBMLanguage lang; + const char *file; + const char *src; + } cases[] = { + {CBM_LANG_JSON, "cmd.json", "{\n \"a\": {\"name\": 1},\n \"b\": {\"name\": 2}\n}\n"}, + {CBM_LANG_PYTHON, "mod.py", + "def pick(a):\n return a\n\n\ndef pick(a):\n return a + 1\n"}, + }; + for (size_t i = 0; i < sizeof(cases) / sizeof(cases[0]); i++) { + CBMFileResult *r = extract(cases[i].src, cases[i].lang, "p", cases[i].file); + ASSERT_NOT_NULL(r); + int listed = 0; + bool distinct_lines = false; + for (int d = 0; d < r->defs.count; d++) { + const CBMDefinition *a = &r->defs.items[d]; + listed += a->variants != NULL; + for (int e = 0; a->qualified_name && a->label && e < d; e++) { + const CBMDefinition *b = &r->defs.items[e]; + distinct_lines = + distinct_lines || + (b->qualified_name && b->label && + strcmp(a->qualified_name, b->qualified_name) == 0 && + strcmp(a->label, b->label) == 0 && a->start_line != b->start_line); + } + } + if (listed != 0 || !distinct_lines) { + fprintf(stderr, + " [c1-guard] %s: defs with variants=%d, repeated on distinct lines=%d\n", + cases[i].file, listed, distinct_lines); + } + ASSERT_TRUE(distinct_lines); + ASSERT_EQ(listed, 0); + cbm_free_result(r); + } + PASS(); +} + /* #1989 review (P1): an over-long candidate-shaped identifier (>= 96 chars) * arriving AFTER a stored candidate must be rejected BEFORE the dedup * compares — the pre-fix dedup ran strncmp(out[k], src, len) with @@ -9566,6 +10301,22 @@ SUITE(extraction) { RUN_TEST(extract_cpp_export_macro_inline_method_recovery_issue1989); RUN_TEST(extract_cpp_export_macro_negative_control_ordinary_caps_issue1989); RUN_TEST(extract_c_export_macro_recovery_issue1989); + RUN_TEST(extract_cpp_macro_prefixed_function_names_c1); + RUN_TEST(extract_cpp_macro_prefixed_lsp_caller_c1); + RUN_TEST(extract_c_macro_prefixed_function_name_c1); + RUN_TEST(extract_c_tag_reference_is_not_a_definition_c1); + RUN_TEST(extract_c_struct_head_in_error_region_recovered_c1); + RUN_TEST(extract_c_keyword_is_never_a_function_name_c1); + RUN_TEST(extract_c_function_lost_to_split_branches_restored_c1); + RUN_TEST(extract_c_typedef_names_are_definitions_c1); + RUN_TEST(extract_c_macro_qn_is_fenced_c1); + RUN_TEST(extract_c_enumerators_are_flat_c1); + RUN_TEST(extract_cpp_enum_scoping_c1); + RUN_TEST(extract_c_enum_inside_struct_is_file_scope_c1); + RUN_TEST(extract_cpp_enum_inside_class_keeps_class_scope_c1); + RUN_TEST(extract_c_enum_macro_wrapped_list_c1); + RUN_TEST(extract_c_variants_lists_same_file_duplicates_c1); + RUN_TEST(extract_variants_only_in_c_preprocessor_languages_c1); RUN_TEST(extract_cpp_export_macro_overlong_candidate_safe_issue1989); RUN_TEST(extract_cpp_export_macro_comment_string_budget_issue1989); RUN_TEST(extract_cpp_export_macro_candidate_cap_issue1989); diff --git a/tests/test_graph_buffer.c b/tests/test_graph_buffer.c index 8a32d5a341..dd257f187d 100644 --- a/tests/test_graph_buffer.c +++ b/tests/test_graph_buffer.c @@ -582,6 +582,71 @@ TEST(gbuf_upsert_same_qn_updates_all_fields) { PASS(); } +/* PR C1: when a Macro and a definition share a QN, the definition owns it whichever + * line comes first; before, the later line won. C-preprocessor macros no longer get + * here (their QN ends in "#macro" and cannot equal a definition's). Chialisp does: + * `defmacro` / `defmac` defs are labelled Macro under a plain QN, so a macro and a + * `defun` of one name collide in one file or in a same-stem .clsp / .clib pair. */ +TEST(gbuf_upsert_definition_beats_same_qn_macro_c1) { + cbm_gbuf_t *gb = cbm_gbuf_new("test", "/tmp"); + /* macro AFTER the function */ + cbm_gbuf_upsert_node(gb, "Function", "assert", "p.cond.assert", "cond.clsp", 742, 779, "{}"); + cbm_gbuf_upsert_node(gb, "Macro", "assert", "p.cond.assert", "cond.clsp", 781, 782, "{}"); + /* macro BEFORE the definition */ + cbm_gbuf_upsert_node(gb, "Macro", "curry", "p.util.curry", "util.clib", 464, 465, "{}"); + cbm_gbuf_upsert_node(gb, "Function", "curry", "p.util.curry", "util.clib", 3193, 3209, "{}"); + /* a macro in the smaller path (.clib sorts before .clsp) still loses */ + cbm_gbuf_upsert_node(gb, "Constant", "limit", "p.x.limit", "x.clsp", 653, 653, "{}"); + cbm_gbuf_upsert_node(gb, "Macro", "limit", "p.x.limit", "x.clib", 9, 10, "{}"); + /* two macros: the classic contract (later line wins) is unchanged */ + cbm_gbuf_upsert_node(gb, "Macro", "M", "p.m.M", "m.clib", 1, 2, "{}"); + cbm_gbuf_upsert_node(gb, "Macro", "M", "p.m.M", "m.clib", 9, 10, "{}"); + + const cbm_gbuf_node_t *n = cbm_gbuf_find_by_qn(gb, "p.cond.assert"); + ASSERT_NOT_NULL(n); + ASSERT_STR_EQ(n->label, "Function"); + ASSERT_EQ(n->start_line, 742); + n = cbm_gbuf_find_by_qn(gb, "p.util.curry"); + ASSERT_NOT_NULL(n); + ASSERT_STR_EQ(n->label, "Function"); + ASSERT_EQ(n->start_line, 3193); + n = cbm_gbuf_find_by_qn(gb, "p.x.limit"); + ASSERT_NOT_NULL(n); + ASSERT_STR_EQ(n->label, "Constant"); + ASSERT_STR_EQ(n->file_path, "x.clsp"); + n = cbm_gbuf_find_by_qn(gb, "p.m.M"); + ASSERT_NOT_NULL(n); + ASSERT_EQ(n->start_line, 9); + /* the label index follows the survivor: only the macro-vs-macro QN keeps one */ + const cbm_gbuf_node_t **macros = NULL; + int macro_count = -1; + ASSERT_EQ(cbm_gbuf_find_by_label(gb, "Macro", ¯os, ¯o_count), 0); + ASSERT_EQ(macro_count, 1); + cbm_gbuf_free(gb); + PASS(); +} + +/* The parallel merge (worker gbuf -> main gbuf) applies the same rule in both + * directions, so the survivor never depends on worker merge order. */ +TEST(gbuf_merge_definition_beats_same_qn_macro_c1) { + for (int order = 0; order < 2; order++) { + cbm_gbuf_t *dst = cbm_gbuf_new("test", "/tmp"); + cbm_gbuf_t *src = cbm_gbuf_new("test", "/tmp"); + cbm_gbuf_t *with_fn = order == 0 ? dst : src; + cbm_gbuf_t *with_macro = order == 0 ? src : dst; + cbm_gbuf_upsert_node(with_fn, "Function", "f", "p.a.f", "a.clsp", 10, 20, "{}"); + cbm_gbuf_upsert_node(with_macro, "Macro", "f", "p.a.f", "a.clsp", 30, 31, "{}"); + ASSERT_EQ(cbm_gbuf_merge(dst, src), 0); + const cbm_gbuf_node_t *n = cbm_gbuf_find_by_qn(dst, "p.a.f"); + ASSERT_NOT_NULL(n); + ASSERT_STR_EQ(n->label, "Function"); + ASSERT_EQ(n->start_line, 10); + cbm_gbuf_free(dst); + cbm_gbuf_free(src); + } + PASS(); +} + TEST(gbuf_upsert_long_qn) { cbm_gbuf_t *gb = cbm_gbuf_new("test", "/tmp"); @@ -1250,6 +1315,8 @@ SUITE(graph_buffer) { RUN_TEST(gbuf_upsert_null_qn); RUN_TEST(gbuf_upsert_empty_qn); RUN_TEST(gbuf_upsert_same_qn_updates_all_fields); + RUN_TEST(gbuf_upsert_definition_beats_same_qn_macro_c1); + RUN_TEST(gbuf_merge_definition_beats_same_qn_macro_c1); RUN_TEST(gbuf_upsert_long_qn); RUN_TEST(gbuf_find_by_qn_missing); RUN_TEST(gbuf_find_by_id_missing); diff --git a/tests/test_mcp.c b/tests/test_mcp.c index c5705ddcfd..2dd2621dbb 100644 --- a/tests/test_mcp.c +++ b/tests/test_mcp.c @@ -6325,6 +6325,50 @@ TEST(tool_trace_call_path_ambiguous) { PASS(); } +/* Tie rule for the C-macro namespace (PR C1): a typedef and its rename macro + * (`typedef struct state_s\n state_t;` + `#define state_t NS(state_t)`) are + * two nodes with one name, `.state_t` and `.state_t#macro`, and + * here the same line span. The name resolves to the definition: without the + * macro ranking below it, the two tie and trace_path answers "ambiguous". */ +TEST(tool_trace_path_definition_beats_c_macro_c1) { + cbm_mcp_server_t *srv = cbm_mcp_server_new(NULL); + cbm_store_t *st = cbm_mcp_server_store(srv); + const char *proj = "tie-proj"; + cbm_mcp_server_set_project(srv, proj); + cbm_store_upsert_project(st, proj, "/tmp/tie"); + cbm_node_t type = {.project = proj, + .label = "Type", + .name = "state_t", + .qualified_name = "tie-proj.xxhash.state_t", + .file_path = "xxhash.h", + .start_line = 653, + .end_line = 654}; + cbm_node_t macro = {.project = proj, + .label = "Macro", + .name = "state_t", + .qualified_name = "tie-proj.xxhash.state_t#macro", + .file_path = "xxhash.h", + .start_line = 429, + .end_line = 430}; /* equal span: a tie without the rule */ + ASSERT_GT(cbm_store_upsert_node(st, &type), 0); + ASSERT_GT(cbm_store_upsert_node(st, ¯o), 0); + + char *resp = cbm_mcp_server_handle( + srv, "{\"jsonrpc\":\"2.0\",\"id\":62,\"method\":\"tools/call\"," + "\"params\":{\"name\":\"trace_call_path\"," + "\"arguments\":{\"function_name\":\"state_t\",\"project\":\"tie-proj\"}}}"); + ASSERT_NOT_NULL(resp); + char *inner = extract_text_content(resp); + ASSERT_NOT_NULL(inner); + ASSERT_NULL(strstr(inner, "ambiguous")); + ASSERT_NULL(strstr(inner, "suggestions")); + ASSERT_NULL(strstr(inner, "function not found")); + free(inner); + free(resp); + cbm_mcp_server_free(srv); + PASS(); +} + /* Multi-seed union hop semantics: bfs_union_same_name deduped visited nodes * keep-FIRST-seen, so a node reached at hop 2 from the first seed kept hop 2 * even when the second seed reaches it at hop 1. hop feeds risk_labels and @@ -16818,6 +16862,79 @@ TEST(snippet_unique_short_name) { PASS(); } +/* ── C-macro namespace (PR C1) ────────────────────────────────── */ + +/* A C macro's QN ends in "#macro", which no caller spells. get_code_snippet + * still returns the macro for its short name, for a dotted suffix and for the + * QN it had before the fence (`.`); when a definition owns that + * QN or that name, the definition is the answer. */ +TEST(snippet_c_macro_namespace_c1) { + char tmp[256]; + cbm_mcp_server_t *srv = setup_snippet_server(tmp, sizeof(tmp)); + ASSERT_NOT_NULL(srv); + cbm_store_t *st = cbm_mcp_server_store(srv); + + cbm_node_t only_macro = {.project = "test-project", + .label = "Macro", + .name = "MAX_LEN", + .qualified_name = "test-project.lib.cfg.MAX_LEN#macro", + .file_path = "main.go", + .start_line = 7, + .end_line = 8}; + cbm_node_t type = {.project = "test-project", + .label = "Type", + .name = "state_t", + .qualified_name = "test-project.lib.xx.state_t", + .file_path = "main.go", + .start_line = 3, + .end_line = 3}; + cbm_node_t shadowed_macro = {.project = "test-project", + .label = "Macro", + .name = "state_t", + .qualified_name = "test-project.lib.xx.state_t#macro", + .file_path = "main.go", + .start_line = 11, + .end_line = 12}; + ASSERT_GT(cbm_store_upsert_node(st, &only_macro), 0); + ASSERT_GT(cbm_store_upsert_node(st, &type), 0); + ASSERT_GT(cbm_store_upsert_node(st, &shadowed_macro), 0); + + static const char *const macro_inputs[] = { + "{\"qualified_name\":\"MAX_LEN\",\"project\":\"test-project\"}", + "{\"qualified_name\":\"cfg.MAX_LEN\",\"project\":\"test-project\"}", + "{\"qualified_name\":\"test-project.lib.cfg.MAX_LEN\",\"project\":\"test-project\"}", + "{\"qualified_name\":\"test-project.lib.cfg.MAX_LEN#macro\",\"project\":\"test-project\"}"}; + for (size_t i = 0; i < sizeof(macro_inputs) / sizeof(macro_inputs[0]); i++) { + char *resp = call_snippet(srv, macro_inputs[i]); + ASSERT_NOT_NULL(resp); + if (!strstr(resp, "\"qualified_name\":\"test-project.lib.cfg.MAX_LEN#macro\"")) { + fprintf(stderr, " [c1-snippet] %s -> %.300s\n", macro_inputs[i], resp); + } + ASSERT_NOT_NULL(strstr(resp, "\"qualified_name\":\"test-project.lib.cfg.MAX_LEN#macro\"")); + ASSERT_NOT_NULL(strstr(resp, "\"label\":\"Macro\"")); + free(resp); + } + + static const char *const definition_inputs[] = { + "{\"qualified_name\":\"test-project.lib.xx.state_t\",\"project\":\"test-project\"}", + "{\"qualified_name\":\"xx.state_t\",\"project\":\"test-project\"}", + "{\"qualified_name\":\"state_t\",\"project\":\"test-project\"}"}; + for (size_t i = 0; i < sizeof(definition_inputs) / sizeof(definition_inputs[0]); i++) { + char *resp = call_snippet(srv, definition_inputs[i]); + ASSERT_NOT_NULL(resp); + if (!strstr(resp, "\"label\":\"Type\"")) { + fprintf(stderr, " [c1-snippet] %s -> %.300s\n", definition_inputs[i], resp); + } + ASSERT_NOT_NULL(strstr(resp, "\"qualified_name\":\"test-project.lib.xx.state_t\"")); + ASSERT_NOT_NULL(strstr(resp, "\"label\":\"Type\"")); + free(resp); + } + + cbm_mcp_server_free(srv); + cleanup_snippet_dir(tmp); + PASS(); +} + /* ── TestSnippet_NameTier ─────────────────────────────────────── */ TEST(snippet_name_tier) { @@ -21101,6 +21218,7 @@ SUITE(mcp) { RUN_TEST(tool_call_invalid_project_name_leaves_no_corrupt_litter_issue1425); RUN_TEST(tool_trace_missing_function_name); RUN_TEST(tool_trace_call_path_ambiguous); + RUN_TEST(tool_trace_path_definition_beats_c_macro_c1); RUN_TEST(tool_trace_union_records_min_hop_across_seeds); RUN_TEST(tool_trace_pagination_exactly_once); RUN_TEST(tool_trace_paging_filters_before_window_and_hashes_effective_args); @@ -21275,6 +21393,7 @@ SUITE(mcp) { RUN_TEST(snippet_exact_qn); RUN_TEST(snippet_qn_suffix); RUN_TEST(snippet_unique_short_name); + RUN_TEST(snippet_c_macro_namespace_c1); RUN_TEST(snippet_name_tier); RUN_TEST(snippet_ambiguous_short_name); RUN_TEST(snippet_not_found); diff --git a/tests/test_pipeline.c b/tests/test_pipeline.c index 9a4c461c64..34f0ba8c2f 100644 --- a/tests/test_pipeline.c +++ b/tests/test_pipeline.c @@ -8576,6 +8576,581 @@ TEST(pipeline_python_project) { PASS(); } +/* Edges of `edge_type` from the node `.` to `.` + * (QNs, not names: a C macro and the definition it shadows share their name). */ +static int c1_edge_count_qn(cbm_store_t *s, const char *project, const char *edge_type, + const char *source_tail, const char *target_tail) { + char source_qn[512]; + char target_qn[512]; + snprintf(source_qn, sizeof(source_qn), "%s.%s", project, source_tail); + snprintf(target_qn, sizeof(target_qn), "%s.%s", project, target_tail); + cbm_node_t source = {0}; + cbm_node_t target = {0}; + int matches = -1; + if (cbm_store_find_node_by_qn(s, project, source_qn, &source) == CBM_STORE_OK && + cbm_store_find_node_by_qn(s, project, target_qn, &target) == CBM_STORE_OK) { + cbm_edge_t *edges = NULL; + int edge_count = 0; + matches = 0; + if (cbm_store_find_edges_by_source_type(s, source.id, edge_type, &edges, &edge_count) == + CBM_STORE_OK) { + for (int i = 0; i < edge_count; i++) { + matches += edges[i].target_id == target.id; + } + cbm_store_free_edges(edges, edge_count); + } + } + cbm_node_free_fields(&source); + cbm_node_free_fields(&target); + return matches; +} + +/* Label of the node `.`, copied into `out`; "" when there is none. */ +static const char *c1_label_of(cbm_store_t *s, const char *project, const char *qn_tail, char *out, + size_t out_size) { + char qn[512]; + snprintf(qn, sizeof(qn), "%s.%s", project, qn_tail); + cbm_node_t node = {0}; + out[0] = '\0'; + if (cbm_store_find_node_by_qn(s, project, qn, &node) == CBM_STORE_OK && node.label) { + snprintf(out, out_size, "%s", node.label); + } + cbm_node_free_fields(&node); + return out; +} + +/* Do the properties of the node `.` contain `needle`? */ +static bool c1_props_contain(cbm_store_t *s, const char *project, const char *qn_tail, + const char *needle) { + char qn[512]; + snprintf(qn, sizeof(qn), "%s.%s", project, qn_tail); + cbm_node_t node = {0}; + bool found = cbm_store_find_node_by_qn(s, project, qn, &node) == CBM_STORE_OK && + node.properties_json && strstr(node.properties_json, needle) != NULL; + cbm_node_free_fields(&node); + return found; +} + +/* PR C1, end to end: C keeps one graph node per QN, so the right entity must win. + * A header and its .c share the module QN (proj.s): `struct S s_global;` in s.c + * minted a Class for the REFERENCE and "smallest path wins" handed it the node (redis: + * struct redisServer pointed at server.c:85). A function and its #else macro + * stand-in shared a QN, and the later macro line won (curl url_match_proxy_use, 41 + * functions): the macro now has a QN of its own (`...#macro`), so both are nodes. + * The typedef alias had no node at all. */ +TEST(pipeline_c_definitions_own_their_qn_c1) { + const char *files[] = {"s.h", "s.c"}; + const char *contents[] = {"struct S {\n" + " int v;\n" + "};\n" + "typedef struct S S_t;\n", + + "#include \"s.h\"\n" + "struct S s_global;\n" + "#ifndef DISABLE_PROXY\n" + "static int match_proxy(int a)\n" + "{\n" + " return a + 1;\n" + "}\n" + "#else\n" + "#define match_proxy(a) 1\n" + "#endif\n" + "int use(void) { return match_proxy(s_global.v); }\n"}; + if (setup_lang_repo(files, contents, 2) != 0) + FAIL("tmpdir"); + char db[512]; + snprintf(db, sizeof(db), "%s/test.db", g_lang_tmpdir); + + cbm_pipeline_t *p = cbm_pipeline_new(g_lang_tmpdir, db, CBM_MODE_FULL); + ASSERT_NOT_NULL(p); + ASSERT_EQ(cbm_pipeline_run(p), 0); + cbm_store_t *s = cbm_store_open_path(db); + ASSERT_NOT_NULL(s); + const char *proj = cbm_pipeline_project_name(p); + + cbm_node_t *nodes = NULL; + int n = 0; + cbm_store_find_nodes_by_name(s, proj, "S", &nodes, &n); + ASSERT_EQ(n, 1); + ASSERT_STR_EQ(nodes[0].label, "Class"); + ASSERT_STR_EQ(nodes[0].file_path, "s.h"); + ASSERT_EQ(nodes[0].start_line, 1); + cbm_store_free_nodes(nodes, n); + + /* the function and its #else macro stand-in: two nodes, the plain QN is the + * function's */ + cbm_store_find_nodes_by_name(s, proj, "match_proxy", &nodes, &n); + ASSERT_EQ(n, 2); + cbm_store_free_nodes(nodes, n); + char qn[512]; + cbm_node_t fn = {0}; + snprintf(qn, sizeof(qn), "%s.s.match_proxy", proj); + ASSERT_EQ(cbm_store_find_node_by_qn(s, proj, qn, &fn), CBM_STORE_OK); + ASSERT_STR_EQ(fn.label, "Function"); + ASSERT_EQ(fn.start_line, 4); + cbm_node_free_fields(&fn); + cbm_node_t mac = {0}; + snprintf(qn, sizeof(qn), "%s.s.match_proxy#macro", proj); + ASSERT_EQ(cbm_store_find_node_by_qn(s, proj, qn, &mac), CBM_STORE_OK); + ASSERT_STR_EQ(mac.label, "Macro"); + ASSERT_STR_EQ(mac.name, "match_proxy"); + ASSERT_EQ(mac.start_line, 9); + cbm_node_free_fields(&mac); + + cbm_store_find_nodes_by_name(s, proj, "S_t", &nodes, &n); + ASSERT_EQ(n, 1); + ASSERT_STR_EQ(nodes[0].label, "Type"); + cbm_store_free_nodes(nodes, n); + + cbm_store_close(s); + cbm_pipeline_free(p); + teardown_lang_repo(); + PASS(); +} + +/* The tie rule of the macro namespace: a name visible both as a definition and as a + * macro resolves to the DEFINITION; the macro is the target only when no definition + * of that name is visible. `pick_fn` is a function in one #if branch and a macro in + * the other: before, one node held the QN and it was the macro (the later line), so + * the call landed on a Macro. `renamed_fn` is declared by a prototype whose name a + * rename macro replaces (jemalloc smallocx): no definition exists in the repo, the + * macro is the only node, and the call keeps pointing at it. */ +TEST(pipeline_c_call_targets_definition_before_macro_c1) { + const char *files[] = {"tie.c"}; + const char *contents[] = {"#define renamed_fn NS_renamed_fn\n" + "extern int renamed_fn(int v);\n" + "#ifdef USE_REAL\n" + "static int pick_fn(int v) { return v; }\n" + "#else\n" + "#define pick_fn(v) (v)\n" + "#endif\n" + "int caller(int b) {\n" + " return pick_fn(b) + renamed_fn(b);\n" + "}\n"}; + if (setup_lang_repo(files, contents, 1) != 0) + FAIL("tmpdir"); + char db[512]; + snprintf(db, sizeof(db), "%s/test.db", g_lang_tmpdir); + + cbm_pipeline_t *p = cbm_pipeline_new(g_lang_tmpdir, db, CBM_MODE_FULL); + ASSERT_NOT_NULL(p); + ASSERT_EQ(cbm_pipeline_run(p), 0); + cbm_store_t *s = cbm_store_open_path(db); + ASSERT_NOT_NULL(s); + char proj[256]; + snprintf(proj, sizeof(proj), "%s", cbm_pipeline_project_name(p)); + + char label[64]; + ASSERT_STR_EQ(c1_label_of(s, proj, "tie.pick_fn", label, sizeof(label)), "Function"); + ASSERT_STR_EQ(c1_label_of(s, proj, "tie.pick_fn#macro", label, sizeof(label)), "Macro"); + ASSERT_STR_EQ(c1_label_of(s, proj, "tie.renamed_fn#macro", label, sizeof(label)), "Macro"); + ASSERT_STR_EQ(c1_label_of(s, proj, "tie.renamed_fn", label, sizeof(label)), ""); + int to_definition = c1_edge_count_qn(s, proj, "CALLS", "tie.caller", "tie.pick_fn"); + int to_shadowed_macro = c1_edge_count_qn(s, proj, "CALLS", "tie.caller", "tie.pick_fn#macro"); + int to_only_macro = c1_edge_count_qn(s, proj, "CALLS", "tie.caller", "tie.renamed_fn#macro"); + + cbm_store_close(s); + cbm_pipeline_free(p); + teardown_lang_repo(); + + ASSERT_EQ(to_definition, 1); + ASSERT_EQ(to_shadowed_macro, 0); + ASSERT_EQ(to_only_macro, 1); + PASS(); +} + +/* A flattened enumerator is still what `E::A` names: the reference resolves by its + * leaf, so `Color::RED`, a namespace-qualified and a class-qualified spelling and the + * bare `GREEN` all reach the node `.`; a scoped enum's `Mode::FAST` + * reaches `.Mode.FAST`. */ +TEST(pipeline_cpp_enum_reference_reaches_flat_enumerator_c1) { + const char *files[] = {"shapes.hpp", "use.cpp"}; + const char *contents[] = {"enum Color { RED, GREEN };\n" + "namespace gfx {\n" + "enum Blend { ADD, MULTIPLY };\n" + "class Brush {\n" + "public:\n" + " enum Shape { ROUND, SQUARE };\n" + "};\n" + "}\n" + "enum class Mode { FAST, SLOW };\n", + + "#include \"shapes.hpp\"\n" + "int pick(int v) {\n" + " if (v == Color::RED) return 1;\n" + " if (v == gfx::Blend::ADD) return 2;\n" + " if (v == gfx::Brush::ROUND) return 3;\n" + " if (v == static_cast(Mode::FAST)) return 4;\n" + " if (v == GREEN) return 5;\n" + " return 0;\n" + "}\n"}; + if (setup_lang_repo(files, contents, 2) != 0) + FAIL("tmpdir"); + char db[512]; + snprintf(db, sizeof(db), "%s/test.db", g_lang_tmpdir); + + cbm_pipeline_t *p = cbm_pipeline_new(g_lang_tmpdir, db, CBM_MODE_FULL); + ASSERT_NOT_NULL(p); + ASSERT_EQ(cbm_pipeline_run(p), 0); + cbm_store_t *s = cbm_store_open_path(db); + ASSERT_NOT_NULL(s); + char proj[256]; + snprintf(proj, sizeof(proj), "%s", cbm_pipeline_project_name(p)); + + static const char *const targets[] = {"shapes.RED", "shapes.gfx.ADD", "shapes.gfx.Brush.ROUND", + "shapes.Mode.FAST", "shapes.GREEN"}; + int usage[5]; + for (int i = 0; i < 5; i++) { + usage[i] = c1_edge_count_qn(s, proj, "USAGE", "use.pick", targets[i]); + } + char nested[64]; + c1_label_of(s, proj, "shapes.Color.RED", nested, sizeof(nested)); + + /* "the constants of enum Color" still enumerates: parent_class is the membership */ + char parent_marker[512]; + snprintf(parent_marker, sizeof(parent_marker), "\"parent_class\":\"%s.shapes.Color\"", proj); + cbm_node_t *vars = NULL; + int var_count = 0; + int color_members = 0; + cbm_store_find_nodes_by_label(s, proj, "Variable", &vars, &var_count); + for (int i = 0; i < var_count; i++) { + color_members += + vars[i].properties_json && strstr(vars[i].properties_json, parent_marker) != NULL; + } + cbm_store_free_nodes(vars, var_count); + + cbm_store_close(s); + cbm_pipeline_free(p); + teardown_lang_repo(); + + for (int i = 0; i < 5; i++) { + if (usage[i] != 1) { + fprintf(stderr, " [c1] USAGE use.pick -> %s: %d\n", targets[i], usage[i]); + } + ASSERT_EQ(usage[i], 1); + } + ASSERT_STR_EQ(nested, ""); + ASSERT_EQ(color_members, 2); + PASS(); +} + +/* C, an enum declared inside a struct (curl lib/cf-h1-proxy.c `enum keeponval {...} + * keepon;`): its enumerators are file-scope names, so their nodes are `.` + * and the unqualified references in the code reach them. The struct is no QN segment, + * for the named enum and the anonymous one alike. */ +TEST(pipeline_c_enum_inside_struct_reference_resolves_c1) { + const char *files[] = {"link.c"}; + const char *contents[] = {"struct conn {\n" + " enum state { ST_IDLE, ST_BUSY } st;\n" + " enum { KIND_A, KIND_B } kind;\n" + " int fd;\n" + "};\n" + "int busy(struct conn *c) {\n" + " return c->st == ST_BUSY || c->kind == KIND_B;\n" + "}\n"}; + if (setup_lang_repo(files, contents, 1) != 0) + FAIL("tmpdir"); + char db[512]; + snprintf(db, sizeof(db), "%s/test.db", g_lang_tmpdir); + + cbm_pipeline_t *p = cbm_pipeline_new(g_lang_tmpdir, db, CBM_MODE_FULL); + ASSERT_NOT_NULL(p); + ASSERT_EQ(cbm_pipeline_run(p), 0); + cbm_store_t *s = cbm_store_open_path(db); + ASSERT_NOT_NULL(s); + char proj[256]; + snprintf(proj, sizeof(proj), "%s", cbm_pipeline_project_name(p)); + + int to_named = c1_edge_count_qn(s, proj, "USAGE", "link.busy", "link.ST_BUSY"); + int to_anonymous = c1_edge_count_qn(s, proj, "USAGE", "link.busy", "link.KIND_B"); + char in_struct[64]; + c1_label_of(s, proj, "link.conn.ST_BUSY", in_struct, sizeof(in_struct)); + char parent_marker[512]; + snprintf(parent_marker, sizeof(parent_marker), "\"parent_class\":\"%s.link.conn.state\"", proj); + bool named_keeps_parent = c1_props_contain(s, proj, "link.ST_BUSY", parent_marker); + bool anonymous_has_parent = c1_props_contain(s, proj, "link.KIND_B", "\"parent_class\""); + + cbm_store_close(s); + cbm_pipeline_free(p); + teardown_lang_repo(); + + ASSERT_EQ(to_named, 1); + ASSERT_EQ(to_anonymous, 1); + ASSERT_STR_EQ(in_struct, ""); + ASSERT_TRUE(named_keeps_parent); + ASSERT_FALSE(anonymous_has_parent); + PASS(); +} + +typedef struct { + int rc; + bool opened; + bool pick_lists_both; /* .v.pick carries the two spans, in order */ + bool limit_has_list; /* .v.LIMIT#macro carries a list */ + bool once_has_list; /* a name defined once must not */ + int other_lang_nodes; /* nodes of cmd.json */ + int other_lang_lists; /* ... that carry a list (must be 0) */ +} c1_variants_obs_t; + +static c1_variants_obs_t c1_observe_variants(const char *repo, const char *db_name) { + c1_variants_obs_t obs = {0}; + char db[512]; + snprintf(db, sizeof(db), "%s/%s", repo, db_name); + cbm_pipeline_t *p = cbm_pipeline_new(repo, db, CBM_MODE_FULL); + if (!p) { + obs.rc = -1; + return obs; + } + obs.rc = cbm_pipeline_run(p); + char proj[256]; + snprintf(proj, sizeof(proj), "%s", cbm_pipeline_project_name(p)); + cbm_pipeline_free(p); + cbm_store_t *s = cbm_store_open_path(db); + if (!s) { + return obs; + } + obs.opened = true; + obs.pick_lists_both = c1_props_contain( + s, proj, "v.pick", "\"variants\":[{\"start\":2,\"end\":2},{\"start\":4,\"end\":6}]"); + obs.limit_has_list = c1_props_contain(s, proj, "v.LIMIT#macro", "\"variants\":[{\"start\":9,"); + obs.once_has_list = c1_props_contain(s, proj, "v.once", "\"variants\""); + cbm_node_t *nodes = NULL; + int count = 0; + if (cbm_store_find_nodes_by_file_overlap(s, proj, "cmd.json", 1, 1000, &nodes, &count) == + CBM_STORE_OK) { + obs.other_lang_nodes = count; + for (int i = 0; i < count; i++) { + obs.other_lang_lists += nodes[i].properties_json && + strstr(nodes[i].properties_json, "\"variants\"") != NULL; + } + cbm_store_free_nodes(nodes, count); + } + cbm_store_close(s); + return obs; +} + +/* `variants` reaches the node on BOTH execution paths. The property is computed per + * file at extraction; the sequential pass (pass_definitions.c) and the parallel one + * (pass_parallel.c) each serialize it, and the properties buffer is sized for the + * list, so a path that forgot either would publish a node without it. One fixture, + * two runs: CBM_INDEX_SINGLE_THREAD forces the sequential path, the 55 fillers plus + * CBM_WORKERS select the parallel one. A JSON file with a repeated key rides along: + * no language outside the C preprocessor ones gets the property. */ +TEST(pipeline_c_variants_on_sequential_and_parallel_paths_c1) { + char tmp[256]; + snprintf(tmp, sizeof(tmp), "/tmp/cbm_c1_variants_XXXXXX"); + if (!cbm_mkdtemp(tmp)) { + FAIL("tmpdir"); + } + char path[512]; + snprintf(path, sizeof(path), "%s/v.c", tmp); + int write_rc = th_write_file(path, "#if A\n" + "int pick(int a) { return a; }\n" + "#else\n" + "int pick(int a) {\n" + " return a + 1;\n" + "}\n" + "#endif\n" + "#ifdef B\n" + "#define LIMIT 1\n" + "#else\n" + "#define LIMIT 2\n" + "#endif\n" + "int once(void) { return 0; }\n"); + snprintf(path, sizeof(path), "%s/cmd.json", tmp); + /* the repeated key on two lines: one span would never be listed, gate or no gate */ + write_rc |= th_write_file(path, "{\n \"a\": {\"name\": 1},\n \"b\": {\"name\": 2}\n}\n"); + for (int i = 0; i < 55; i++) { + char source[96]; + snprintf(path, sizeof(path), "%s/pad_%02d.c", tmp, i); + snprintf(source, sizeof(source), "int c1_pad_%02d(void) { return %d; }\n", i, i); + write_rc |= th_write_file(path, source); + } + if (write_rc != 0) { + th_rmtree(tmp); + FAIL("failed to write the variants fixture"); + } + + char *old_workers = getenv("CBM_WORKERS"); + char *saved_workers = old_workers ? strdup(old_workers) : NULL; + char *old_single = getenv("CBM_INDEX_SINGLE_THREAD"); + char *saved_single = old_single ? strdup(old_single) : NULL; + + cbm_setenv("CBM_INDEX_SINGLE_THREAD", "1", 1); + c1_variants_obs_t sequential = c1_observe_variants(tmp, "variants-sequential.db"); + + cbm_unsetenv("CBM_INDEX_SINGLE_THREAD"); + cbm_setenv("CBM_WORKERS", "4", 1); + c1_variants_obs_t parallel = c1_observe_variants(tmp, "variants-parallel.db"); + + if (saved_workers) { + cbm_setenv("CBM_WORKERS", saved_workers, 1); + free(saved_workers); + } else { + cbm_unsetenv("CBM_WORKERS"); + } + if (saved_single) { + cbm_setenv("CBM_INDEX_SINGLE_THREAD", saved_single, 1); + free(saved_single); + } else { + cbm_unsetenv("CBM_INDEX_SINGLE_THREAD"); + } + th_rmtree(tmp); + + const c1_variants_obs_t *runs[] = {&sequential, ¶llel}; + for (int i = 0; i < 2; i++) { + if (!runs[i]->pick_lists_both || !runs[i]->limit_has_list || runs[i]->once_has_list || + runs[i]->other_lang_lists != 0) { + fprintf( + stderr, + " [c1-variants] path=%s pick=%d limit=%d once=%d json_nodes=%d json_lists=%d\n", + i == 0 ? "sequential" : "parallel", runs[i]->pick_lists_both, + runs[i]->limit_has_list, runs[i]->once_has_list, runs[i]->other_lang_nodes, + runs[i]->other_lang_lists); + } + } + /* the property on each path first, the language limit second: one concern must + * not hide the other */ + for (int i = 0; i < 2; i++) { + ASSERT_EQ(runs[i]->rc, 0); + ASSERT_TRUE(runs[i]->opened); + ASSERT_TRUE(runs[i]->pick_lists_both); + ASSERT_TRUE(runs[i]->limit_has_list); + ASSERT_FALSE(runs[i]->once_has_list); + } + for (int i = 0; i < 2; i++) { + ASSERT_GT(runs[i]->other_lang_nodes, 0); + ASSERT_EQ(runs[i]->other_lang_lists, 0); + } + PASS(); +} + +/* C++ overloads share a QN (the signature is no part of it), so two of them in one + * file are the same shape as two #if branches: one node -- the last by start line -- + * and `variants` lists both spans. */ +TEST(pipeline_cpp_overloads_are_listed_as_variants_c1) { + const char *files[] = {"ov.cpp"}; + const char *contents[] = {"int scale(int v) { return v * 2; }\n" + "double scale(double v) {\n" + " return v * 2.0;\n" + "}\n" + "int once(int v) { return scale(v); }\n"}; + if (setup_lang_repo(files, contents, 1) != 0) + FAIL("tmpdir"); + char db[512]; + snprintf(db, sizeof(db), "%s/test.db", g_lang_tmpdir); + + cbm_pipeline_t *p = cbm_pipeline_new(g_lang_tmpdir, db, CBM_MODE_FULL); + ASSERT_NOT_NULL(p); + ASSERT_EQ(cbm_pipeline_run(p), 0); + cbm_store_t *s = cbm_store_open_path(db); + ASSERT_NOT_NULL(s); + char proj[256]; + snprintf(proj, sizeof(proj), "%s", cbm_pipeline_project_name(p)); + + cbm_node_t *nodes = NULL; + int n = 0; + cbm_store_find_nodes_by_name(s, proj, "scale", &nodes, &n); + int scale_nodes = n; + int kept_start = n > 0 ? nodes[0].start_line : -1; + cbm_store_free_nodes(nodes, n); + bool lists_both = c1_props_contain( + s, proj, "ov.scale", "\"variants\":[{\"start\":1,\"end\":1},{\"start\":2,\"end\":4}]"); + bool once_has_list = c1_props_contain(s, proj, "ov.once", "\"variants\""); + + cbm_store_close(s); + cbm_pipeline_free(p); + teardown_lang_repo(); + + ASSERT_EQ(scale_nodes, 1); + ASSERT_EQ(kept_start, 2); + ASSERT_TRUE(lists_both); + ASSERT_FALSE(once_has_list); + PASS(); +} + +/* CBM_SEMANTIC_INDEX_VERSION 4: the C-family node identities changed (macro QNs end + * in "#macro", unscoped enumerators are flat, typedef names are nodes). An index + * written at version 3 holds the old QNs for every file that did not change, and an + * unchanged repository is otherwise a no-op, so the version is what makes the new + * binary rebuild it. The stored index is put back to the version-3 state by hand + * (metadata and the old macro QN); the run after that must replace it in full. */ +TEST(pipeline_semantic_version_3_index_is_rebuilt_in_full_c1) { + char tmp[256]; + snprintf(tmp, sizeof(tmp), "/tmp/cbm_c1_version_XXXXXX"); + ASSERT_NOT_NULL(cbm_mkdtemp(tmp)); + write_temp_file(tmp, "cfg.c", "#define LIMIT 10\nint limit(void) { return LIMIT; }\n"); + char db_path[512]; + snprintf(db_path, sizeof(db_path), "%s/index.db", tmp); + + cbm_pipeline_incremental_test_reset_faults(); + cbm_pipeline_t *first = cbm_pipeline_new(tmp, db_path, CBM_MODE_FULL); + ASSERT_NOT_NULL(first); + ASSERT_EQ(cbm_pipeline_run(first), 0); + char project[256]; + snprintf(project, sizeof(project), "%s", cbm_pipeline_project_name(first)); + cbm_pipeline_free(first); + + /* Put the published index back to what version 3 wrote. */ + cbm_store_t *store = cbm_store_open_path(db_path); + ASSERT_NOT_NULL(store); + char label[64]; + bool fenced_before = + strcmp(c1_label_of(store, project, "cfg.LIMIT#macro", label, sizeof(label)), "Macro") == 0; + cbm_coverage_row_t *coverage_rows = NULL; + int coverage_count = 0; + ASSERT_EQ(cbm_store_coverage_get(store, project, &coverage_rows, &coverage_count), + CBM_STORE_OK); + cbm_coverage_meta_t meta = {0}; + ASSERT_EQ(cbm_store_coverage_meta_get(store, project, &meta), CBM_STORE_OK); + cbm_coverage_meta_t old_meta = meta; + old_meta.coverage_version = 3; + ASSERT_EQ( + cbm_store_coverage_replace_ex(store, project, coverage_rows, coverage_count, &old_meta), + CBM_STORE_OK); + cbm_store_free_coverage(coverage_rows, coverage_count); + cbm_store_coverage_meta_clear(&meta); + ASSERT_EQ(cbm_store_exec(store, "UPDATE nodes SET qualified_name = " + "substr(qualified_name, 1, length(qualified_name) - 6) " + "WHERE label = 'Macro' AND qualified_name LIKE '%#macro';"), + CBM_STORE_OK); + bool plain_after_downgrade = + strcmp(c1_label_of(store, project, "cfg.LIMIT", label, sizeof(label)), "Macro") == 0; + cbm_store_close(store); + + /* Nothing in the repository changed: only the version says the index is stale. */ + cbm_pipeline_incremental_test_reset_faults(); + cbm_pipeline_t *upgrade = cbm_pipeline_new(tmp, db_path, CBM_MODE_FULL); + ASSERT_NOT_NULL(upgrade); + int upgrade_rc = cbm_pipeline_run(upgrade); + cbm_incremental_route_t upgrade_route = cbm_pipeline_incremental_test_last_route(); + cbm_pipeline_free(upgrade); + + store = cbm_store_open_path(db_path); + ASSERT_NOT_NULL(store); + bool fenced_after = + strcmp(c1_label_of(store, project, "cfg.LIMIT#macro", label, sizeof(label)), "Macro") == 0; + bool plain_after = + strcmp(c1_label_of(store, project, "cfg.LIMIT", label, sizeof(label)), "Macro") == 0; + cbm_coverage_meta_t new_meta = {0}; + ASSERT_EQ(cbm_store_coverage_meta_get(store, project, &new_meta), CBM_STORE_OK); + int stored_version = new_meta.coverage_version; + cbm_store_coverage_meta_clear(&new_meta); + cbm_store_close(store); + cbm_pipeline_incremental_test_reset_faults(); + th_rmtree(tmp); + + ASSERT_TRUE(fenced_before); + ASSERT_TRUE(plain_after_downgrade); + ASSERT_EQ(upgrade_rc, 0); + ASSERT_EQ(upgrade_route, CBM_INCREMENTAL_ROUTE_FORCED_FULL); + ASSERT_EQ(stored_version, CBM_SEMANTIC_INDEX_VERSION); + ASSERT_TRUE(fenced_after); + ASSERT_FALSE(plain_after); + ASSERT_GTE(CBM_SEMANTIC_INDEX_VERSION, 4); + PASS(); +} + /* `#include ` from arch/x/bugs.c: two headers end with the * include path (include/linux/device.h, tools/virtio/linux/device.h) and each * declares `struct device`. The exact-file lookup returned the first matching @@ -16462,6 +17037,13 @@ SUITE(pipeline) { RUN_TEST(usages_kotlin_no_duplicate_calls); /* Language integration tests */ RUN_TEST(pipeline_python_project); + RUN_TEST(pipeline_c_definitions_own_their_qn_c1); + RUN_TEST(pipeline_c_call_targets_definition_before_macro_c1); + RUN_TEST(pipeline_cpp_enum_reference_reaches_flat_enumerator_c1); + RUN_TEST(pipeline_c_enum_inside_struct_reference_resolves_c1); + RUN_TEST(pipeline_c_variants_on_sequential_and_parallel_paths_c1); + RUN_TEST(pipeline_cpp_overloads_are_listed_as_variants_c1); + RUN_TEST(pipeline_semantic_version_3_index_is_rebuilt_in_full_c1); RUN_TEST(pipeline_header_include_target_is_independent_of_registration_order); RUN_TEST(pipeline_imports_multi_symbol_edges); RUN_TEST(pipeline_go_cross_package_call); diff --git a/tests/test_store_search.c b/tests/test_store_search.c index 9110c420ae..4c1063bcc9 100644 --- a/tests/test_store_search.c +++ b/tests/test_store_search.c @@ -8,6 +8,7 @@ #include "test_framework.h" #include "test_helpers.h" #include +#include "cbm.h" /* CBM_MACRO_QN_SUFFIX — the fence the FTS feed strips */ #include "sqlite3.h" /* vendored/sqlite3 — raw nodes_fts MATCH probes */ #include #include @@ -1797,6 +1798,66 @@ TEST(store_fts_rebuild_incremental_adds_only_nodes_above_watermark) { PASS(); } +/* A C macro's QN ends in "#macro" (CBM_MACRO_QN_SUFFIX) and `#` separates tokens, so + * indexing that QN as it is would give every macro node the word "macro" a second time + * (its label already says Macro). On redis that is 5,513 rows outscoring the two nodes + * that carry the word in their NAME: tre_expand_macro fell out of the 2,000-row BM25 + * candidate window and out of the result for the query `macro`. The qualified_name + * column therefore gets a Macro's QN without the fence, and only a Macro's. The + * wholesale rebuild and the delta-merge rebuild are one writer; both are run here. */ +TEST(store_fts_macro_qn_fence_adds_no_token_c1) { + cbm_store_t *s = cbm_store_open_memory(); + ASSERT_NOT_NULL(s); + cbm_store_upsert_project(s, "p", "/tmp/p"); + cbm_node_t macro = {.project = "p", + .label = "Macro", + .name = "LIMIT", + .qualified_name = "p.cfg.LIMIT" CBM_MACRO_QN_SUFFIX, + .file_path = "cfg.h"}; + cbm_node_t named = {.project = "p", + .label = "Function", + .name = "expandMacro", + .qualified_name = "p.cfg.expandMacro", + .file_path = "cfg.c"}; + /* not a Macro: a QN that merely ends like the fence keeps all its tokens */ + cbm_node_t other = {.project = "p", + .label = "Function", + .name = "odd", + .qualified_name = "p.cfg.odd#macro", + .file_path = "cfg.c"}; + ASSERT_TRUE(cbm_store_upsert_node(s, ¯o) > 0); + ASSERT_TRUE(cbm_store_upsert_node(s, &named) > 0); + ASSERT_TRUE(cbm_store_upsert_node(s, &other) > 0); + + ASSERT_EQ(cbm_store_fts_rebuild(s, NULL, 0), CBM_STORE_OK); + ASSERT_EQ(fts_match_count(s, "qualified_name:LIMIT"), 1); /* the QN stays searchable */ + ASSERT_EQ(fts_match_count(s, "label:Macro"), 1); + ASSERT_EQ(fts_match_count(s, "qualified_name:macro"), 1); /* `odd` alone */ + ASSERT_EQ(fts_match_count(s, "name:Macro"), 1); /* expandMacro, camel-split */ + + sqlite3_stmt *st = NULL; + ASSERT_EQ(sqlite3_prepare_v2(cbm_store_get_db(s), "SELECT COALESCE(MAX(id),0) FROM nodes", -1, + &st, NULL), + SQLITE_OK); + ASSERT_EQ(sqlite3_step(st), SQLITE_ROW); + int64_t watermark = sqlite3_column_int64(st, 0); + sqlite3_finalize(st); + + cbm_node_t added = {.project = "p", + .label = "Macro", + .name = "DEPTH", + .qualified_name = "p.cfg.DEPTH" CBM_MACRO_QN_SUFFIX, + .file_path = "cfg.h"}; + ASSERT_TRUE(cbm_store_upsert_node(s, &added) > 0); + ASSERT_EQ(cbm_store_fts_rebuild(s, "p", watermark), CBM_STORE_OK); + ASSERT_EQ(fts_match_count(s, "qualified_name:DEPTH"), 1); + ASSERT_EQ(fts_match_count(s, "label:Macro"), 2); + ASSERT_EQ(fts_match_count(s, "qualified_name:macro"), 1); /* the delta path strips it too */ + + cbm_store_close(s); + PASS(); +} + SUITE(store_search) { RUN_TEST(store_search_by_label); RUN_TEST(store_search_by_name_pattern); @@ -1873,4 +1934,5 @@ SUITE(store_search) { RUN_TEST(store_fts_rebuild_survives_malformed_properties_json); RUN_TEST(store_fts_rebuild_tolerates_legacy_four_column_table); RUN_TEST(store_fts_rebuild_incremental_adds_only_nodes_above_watermark); + RUN_TEST(store_fts_macro_qn_fence_adds_no_token_c1); } From 78772eb3011f6d5c39bd60f02f3d6cdef5b6bae0 Mon Sep 17 00:00:00 2001 From: Martin Vogel Date: Sat, 3 Oct 2026 18:37:19 +0200 Subject: [PATCH 2/7] test: checkpoint cross-platform harness fixes Isolate the watcher-disabled runtime, use a deterministic Python responder for the fuzz harness environment probe, and align the version metadata contract with the current Scoop manifest. Signed-off-by: Martin Vogel --- tests/test_runtime_isolation_contract.sh | 1 + tests/test_security_fuzz_harness.sh | 39 +++++++++++++++--------- tests/test_version_metadata_contract.sh | 3 +- tests/test_watcher_disabled.sh | 19 +++++++++--- 4 files changed, 41 insertions(+), 21 deletions(-) diff --git a/tests/test_runtime_isolation_contract.sh b/tests/test_runtime_isolation_contract.sh index 217230c7fa..0d6865a490 100644 --- a/tests/test_runtime_isolation_contract.sh +++ b/tests/test_runtime_isolation_contract.sh @@ -93,6 +93,7 @@ ENTRY_POINTS=( scripts/security-install.sh scripts/security-fuzz.sh scripts/security-fuzz-random.sh scripts/security-network.sh tests/test_parent_watchdog.sh tests/test_worker_watchdog.sh tests/test_worker_error_response.sh tests/test_hook_conflict_notice.sh + tests/test_watcher_disabled.sh ) for entry in "${ENTRY_POINTS[@]}"; do grep -q 'test-runtime.sh' "$ROOT/$entry" || fail "$entry does not source the helper" diff --git a/tests/test_security_fuzz_harness.sh b/tests/test_security_fuzz_harness.sh index 713107b1f8..79f43ad3e7 100644 --- a/tests/test_security_fuzz_harness.sh +++ b/tests/test_security_fuzz_harness.sh @@ -41,24 +41,35 @@ if "$ROOT/scripts/security-fuzz.sh" "$ECHO_ONLY" \ exit 1 fi +# One python3 process answers the requests; the harness needs python3 for its +# interactive case anyway. A shell `read` of the 1 MB request costs seconds +# where the shell runs emulated, and the harness's 10-second hang limit then +# decides this test instead of its assertions. +ENV_RESPONDER="$WORKDIR/environment-probe-responder.py" +cat > "$ENV_RESPONDER" <<'EOF' +import re +import sys + +# Echo a JSON-RPC result for every request with a numeric id. This keeps the +# fixture compatible with both the current fixed ids and a future per-case +# acknowledgement id without depending on the malformed payload itself: the +# last id of a line is the one that counts. +REQUEST_ID = re.compile(rb'"id"\s*:\s*([0-9]+)') + +for line in sys.stdin.buffer: + ids = REQUEST_ID.findall(line) + if not ids: + continue + result = b'{"isError":true}' if b'"name":"index_repository"' in line else b'{}' + sys.stdout.buffer.write(b'{"jsonrpc":"2.0","id":' + ids[-1] + b',"result":' + result + b'}\n') + sys.stdout.buffer.flush() +EOF + ENV_PROBE="$WORKDIR/environment-probe-mcp" cat > "$ENV_PROBE" <<'EOF' #!/usr/bin/env bash printf '%s\t%s\t%s\n' "${HOME-}" "${CBM_CACHE_DIR-}" "${CBM_RUNTIME_DIR-}" >> "$CBM_FUZZ_ENV_PROBE" - -# Echo a JSON-RPC result for every request with a numeric id. This keeps the -# fixture compatible with both the current fixed ids and a future per-case -# acknowledgement id without depending on the malformed payload itself. -while IFS= read -r line; do - id=$(printf '%s\n' "$line" | sed -n 's/.*"id"[[:space:]]*:[[:space:]]*\([0-9][0-9]*\).*/\1/p') - if [[ -n "$id" ]]; then - if [[ "$line" == *'"name":"index_repository"'* ]]; then - printf '{"jsonrpc":"2.0","id":%s,"result":{"isError":true}}\n' "$id" - else - printf '{"jsonrpc":"2.0","id":%s,"result":{}}\n' "$id" - fi - fi -done +exec python3 "${BASH_SOURCE[0]%/*}/environment-probe-responder.py" EOF chmod +x "$ENV_PROBE" diff --git a/tests/test_version_metadata_contract.sh b/tests/test_version_metadata_contract.sh index fa66caeaa2..05655a1486 100644 --- a/tests/test_version_metadata_contract.sh +++ b/tests/test_version_metadata_contract.sh @@ -107,8 +107,8 @@ SURFACES=( "pkg/go/cmd/codebase-memory-mcp/main.go|release" "pkg/chocolatey/codebase-memory-mcp.nuspec|release" "pkg/chocolatey/tools/chocolateyInstall.ps1|release" + "pkg/scoop/codebase-memory-mcp.json|release" "pkg/homebrew/Formula/codebase-memory-mcp.rb|pin:0.10.3" - "pkg/scoop/codebase-memory-mcp.json|pin:0.11.0" "pkg/aur/PKGBUILD|pin:0.8.1" "pkg/aur/.SRCINFO|pin:0.8.1" ) @@ -119,7 +119,6 @@ SURFACES=( # AND every sha256 it pins, then move the entry to "release" above. PIN_REASONS=( "pkg/homebrew/Formula/codebase-memory-mcp.rb|pins 4 per-asset sha256 (darwin/linux x arm/intel); last re-pinned for v0.10.3" - "pkg/scoop/codebase-memory-mcp.json|pins 2 per-asset sha256 (Windows amd64 + arm64 zip); last re-pinned for v0.11.0" "pkg/aur/PKGBUILD|pins sha256sums_x86_64 + sha256sums_aarch64; last re-pinned for v0.8.1" "pkg/aur/.SRCINFO|generated from PKGBUILD, so it must move with it, not before it" ) diff --git a/tests/test_watcher_disabled.sh b/tests/test_watcher_disabled.sh index 23d91ae745..bc62093504 100755 --- a/tests/test_watcher_disabled.sh +++ b/tests/test_watcher_disabled.sh @@ -39,18 +39,27 @@ esac [[ -x "${BINARY}" ]] || { echo "missing binary: ${BINARY}" >&2; exit 2; } command -v git >/dev/null 2>&1 || { echo "git required for fixture" >&2; exit 2; } -work="$(mktemp -d)" +# A private runtime directory, not only a private cache: daemon admission is +# decided per runtime directory, so in the account-wide one a released CBM that +# a developer has in use refuses this build ("a conflicting CBM process is +# active") and the test could never start. With the helper's root these daemons +# never meet a developer's real one. Its root is short enough for the socket +# path; everything this test creates lives under it. +# shellcheck source=../scripts/test-runtime.sh +source "${ROOT}/scripts/test-runtime.sh" +cbm_test_runtime_init +work="${CBM_TEST_RUNTIME_ROOT}/work" +mkdir "${work}" -# Every run gets its own CBM_CACHE_DIR, and the daemon endpoint is derived from -# it, so these daemons are private to the test and can never touch a developer's -# real one. Retire them on any exit path regardless. +# Every run gets its own CBM_CACHE_DIR (and so its own daemon log). Retire the +# daemons on any exit path, then let the helper remove the root. cleanup() { local cache for cache in "${work}"/cache-*; do [[ -d "${cache}" ]] || continue CBM_CACHE_DIR="${cache}" "${BINARY}" daemon stop >/dev/null 2>&1 || true done - rm -rf "${work}" + cbm_test_runtime_cleanup "${BINARY}" } trap cleanup EXIT From fc25f3746890ba310abd6d83412c1dcfc85d872d Mon Sep 17 00:00:00 2001 From: Martin Vogel Date: Sun, 4 Oct 2026 17:57:32 +0200 Subject: [PATCH 3/7] fix(extract): narrow the scope of the typedef-pass start index cppcheck 2.20 (the CI linter) reports variableScope for the count taken before the definitions walk, which only the C-family branch uses. The branch now takes the count, walks and drops the shadowed typedefs; the other languages walk as before. Behaviour unchanged: extraction and pipeline suites pass. Signed-off-by: Martin Vogel --- internal/cbm/extract_defs.c | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/internal/cbm/extract_defs.c b/internal/cbm/extract_defs.c index 3d50c95f24..281b38359a 100644 --- a/internal/cbm/extract_defs.c +++ b/internal/cbm/extract_defs.c @@ -10498,10 +10498,12 @@ void cbm_extract_definitions_without_module(CBMExtractCtx *ctx) { } // Walk AST for function/class definitions - int first = ctx->result->defs.count; - walk_defs(ctx, ctx->root, spec, 0); if (is_c_declarator_lang(ctx->language)) { + int first = ctx->result->defs.count; + walk_defs(ctx, ctx->root, spec, 0); drop_c_typedefs_shadowed_by_tags(ctx, first); + } else { + walk_defs(ctx, ctx->root, spec, 0); } // Extract module-level variables From d64f7e5cad9046127e941a5d1265089288e6360c Mon Sep 17 00:00:00 2001 From: Martin Vogel Date: Wed, 7 Oct 2026 18:36:35 +0200 Subject: [PATCH 4/7] fix(extract): name an anonymous C typedef aggregate once Main (#2458) named the anonymous struct/union/enum of `typedef struct { ... } Name;` in extract_class_def from the typedef's declarator. This branch's extract_c_typedef already emits that aggregate as Name from the type_definition, so on the merged tree Name was emitted twice under one label: its members twice and a false variants entry. extract_c_typedef stays the one place that names it; the extract_class_def case is removed. Both typedef test sets pass on the result: #2458's extract_c_anonymous_typedef_aggregate_is_named_by_its_typedef and member-access tests, and this branch's C1 typedef/variant tests (extraction, pipeline, registry, graph_buffer, store_search, mcp: 1,336 passed, 4 skipped). Signed-off-by: Martin Vogel --- internal/cbm/extract_defs.c | 24 +++--------------------- 1 file changed, 3 insertions(+), 21 deletions(-) diff --git a/internal/cbm/extract_defs.c b/internal/cbm/extract_defs.c index 3bd1ecc51e..d109995bd6 100644 --- a/internal/cbm/extract_defs.c +++ b/internal/cbm/extract_defs.c @@ -7379,27 +7379,9 @@ static void extract_class_def(CBMExtractCtx *ctx, TSNode node, const CBMLangSpec } break; } - case CBM_LANG_C: - case CBM_LANG_CPP: { // `typedef struct { … } Name;`: the aggregate is - // anonymous and the typedef's declarator names it. - // A pointer/array/function declarator names another - // type, not the aggregate, so only a plain - // type_identifier counts. - if (strcmp(kind, "struct_specifier") != 0 && strcmp(kind, "union_specifier") != 0 && - strcmp(kind, "enum_specifier") != 0) { - break; - } - TSNode parent = ts_node_parent(node); - if (ts_node_is_null(parent) || strcmp(ts_node_type(parent), "type_definition") != 0) { - break; - } - TSNode declarator = ts_node_child_by_field_name(parent, TS_FIELD("declarator")); - if (!ts_node_is_null(declarator) && - strcmp(ts_node_type(declarator), "type_identifier") == 0) { - name_node = declarator; - } - break; - } + /* C/C++ `typedef struct { … } Name;` is named by extract_c_typedef from + * the type_definition; naming the anonymous specifier here as well + * emitted Name twice (members twice, a false variant). */ default: break; } From c4d6453b6dd73f7ac57fb35d368b0209d921897b Mon Sep 17 00:00:00 2001 From: Martin Vogel Date: Wed, 7 Oct 2026 18:53:30 +0200 Subject: [PATCH 5/7] test(extract): an anonymous C typedef aggregate is one definition Regression test for the merge with #2458: with its extract_class_def case restored, typedef struct { int count; } Tally; yields two Class defs named Tally (count_defs_named == 2, RED at test_extraction.c:1129); with the case removed it is one Class, one count field, one Shade enum, and no variants list (extraction 432 passed). Signed-off-by: Martin Vogel --- tests/test_extraction.c | 24 ++++++++++++++++++++++++ 1 file changed, 24 insertions(+) diff --git a/tests/test_extraction.c b/tests/test_extraction.c index 1390a25bc6..c5057e7f72 100644 --- a/tests/test_extraction.c +++ b/tests/test_extraction.c @@ -1116,6 +1116,29 @@ TEST(extract_c_anonymous_typedef_aggregate_is_named_by_its_typedef) { PASS(); } +/* The anonymous aggregate of a typedef is ONE definition. extract_c_typedef names + * it from the type_definition; a second naming from the anonymous specifier gave + * Tally two Class defs under one QN, every member twice, and a variants list. */ +TEST(extract_c_anonymous_typedef_aggregate_is_one_def) { + CBMFileResult *r = extract("typedef struct {\n" + " int count;\n" + "} Tally;\n" + "typedef enum { SHADE_RED, SHADE_GREEN } Shade;\n", + CBM_LANG_C, "t", "tally.h"); + ASSERT_NOT_NULL(r); + ASSERT_EQ(count_defs_named(r, "Class", "Tally"), 1); + ASSERT_EQ(count_defs_named(r, "Field", "count"), 1); + ASSERT_EQ(count_defs_named(r, "Enum", "Shade"), 1); + for (int i = 0; i < r->defs.count; i++) { + const CBMDefinition *d = &r->defs.items[i]; + if (d->name && strcmp(d->name, "Tally") == 0) { + ASSERT_NULL(d->variants); + } + } + cbm_free_result(r); + PASS(); +} + /* A pointer-to-function member (`void (*open)(int);`, the shape of every * kernel ops table) is a field of its struct. It shares the * field_declaration + function_declarator shape with a C++ member FUNCTION @@ -10405,6 +10428,7 @@ SUITE(extraction) { RUN_TEST(c_function_return_type_plain_unchanged); RUN_TEST(c_struct); RUN_TEST(extract_c_anonymous_typedef_aggregate_is_named_by_its_typedef); + RUN_TEST(extract_c_anonymous_typedef_aggregate_is_one_def); RUN_TEST(extract_c_function_pointer_member_is_a_field); RUN_TEST(extract_c_member_declarators_name_and_type); RUN_TEST(cpp_class); From c828105d1eafe377e00a465953613eb0e32c0a31 Mon Sep 17 00:00:00 2001 From: Martin Vogel Date: Wed, 7 Oct 2026 22:16:36 +0200 Subject: [PATCH 6/7] fix(extract): write C variants in graph_buffer's schema Main's graph_buffer.c ("Definition variants") records a definition's variant spans as [{"file_path","start_line","end_line"}], and the test-impact engine reads that schema to find every TEST(...) branch of one test. This branch's extractor wrote [{"start","end"}]. On the merge with main, graph_buffer's same-QN merge found a non-empty variants list it could not read and dropped every span, the definition's own included. The engine then no longer found the second #if branch of TEST(alpha_uses_lib), so test_impact_engine_variant_tests_are_one_test reported suite alpha UNMAPPED, in CI on macOS and Windows and locally. The extractor now writes the shared schema, with the file path JSON-escaped. The C-preprocessor-only gate and the one-carrier rule are unchanged, and the variants assertions use the shared format. test_impact_engine, extraction, pipeline, graph_buffer and conditional_variants: 857 passed. Signed-off-by: Martin Vogel --- internal/cbm/cbm.c | 25 +++++++++++++++++++++---- internal/cbm/cbm.h | 4 +++- tests/test_extraction.c | 9 ++++++--- tests/test_pipeline.c | 15 ++++++++++----- 4 files changed, 40 insertions(+), 13 deletions(-) diff --git a/internal/cbm/cbm.c b/internal/cbm/cbm.c index f4ce1354fa..366a714086 100644 --- a/internal/cbm/cbm.c +++ b/internal/cbm/cbm.c @@ -26,12 +26,14 @@ #include "foundation/compat.h" #include "foundation/compat_fs.h" // cbm_fopen — crash-supervisor per-file marker write #include "foundation/hash_table.h" // CBMHashTable — crash-supervisor quarantine set +#include "foundation/str_util.h" // cbm_json_escape — variants file_path #include "tree_sitter/api.h" // TSParser, TSNode, TSTree, TSInput, TSLanguage, TSPoint, TSParseOptions, TSParseState #include "foundation/constants.h" #include "mimalloc.h" // mi_malloc/mi_calloc/mi_realloc/mi_free/mi_usable_size — bind 3rd-party allocators (#424) #if defined(CBM_BIND_TS_ALLOCATOR) && CBM_BIND_TS_ALLOCATOR #include "sqlite3.h" // sqlite3_mem_methods, sqlite3_config, SQLITE_CONFIG_MALLOC — bind sqlite to mimalloc #endif +#include // INT_MAX #include // uint32_t, uint64_t, int64_t #include #include @@ -2165,7 +2167,9 @@ static uint64_t cbm_variant_key(const char *qn, const char *label) { /* refs[0..n) is one (QN, label) group in span order: give every member the * list of the group's distinct spans. */ static void cbm_variant_assign(CBMFileResult *result, const cbm_variant_ref_t *refs, int n) { - enum { VARIANT_ENTRY_MAX = 40 }; /* ,{"start":4294967295,"end":4294967295} */ + /* ,{"file_path":"","start_line":4294967295,"end_line":4294967295} + path; + * cbm_json_escape writes at most 6 bytes per source byte (\u00XX). */ + enum { VARIANT_ENTRY_FIXED = 64, JSON_ESCAPE_GROWTH = 6 }; int distinct = 0; for (int i = 0; i < n; i++) { if (i == 0 || refs[i].start != refs[i - 1].start || refs[i].end != refs[i - 1].end) { @@ -2175,7 +2179,17 @@ static void cbm_variant_assign(CBMFileResult *result, const cbm_variant_ref_t *r if (distinct < 2) { return; } - size_t cap = (size_t)distinct * VARIANT_ENTRY_MAX + 3; + const char *path = result->defs.items[refs[0].idx].file_path; + size_t path_cap = (path ? strlen(path) : 0) * JSON_ESCAPE_GROWTH + 1; + if (path_cap > INT_MAX) { + return; + } + char *esc = (char *)cbm_arena_alloc(&result->arena, path_cap); + if (!esc) { + return; + } + (void)cbm_json_escape(esc, (int)path_cap, path ? path : ""); + size_t cap = (size_t)distinct * (VARIANT_ENTRY_FIXED + strlen(esc)) + 3; char *json = (char *)cbm_arena_alloc(&result->arena, cap); if (!json) { return; @@ -2186,8 +2200,11 @@ static void cbm_variant_assign(CBMFileResult *result, const cbm_variant_ref_t *r if (i > 0 && refs[i].start == refs[i - 1].start && refs[i].end == refs[i - 1].end) { continue; } - int w = snprintf(json + pos, cap - pos, "%s{\"start\":%u,\"end\":%u}", pos > 1 ? "," : "", - refs[i].start, refs[i].end); + /* graph_buffer.c's variants schema, so its same-QN merge and every + * reader of the node (the test-impact engine) see one format. */ + int w = snprintf(json + pos, cap - pos, + "%s{\"file_path\":\"%s\",\"start_line\":%u,\"end_line\":%u}", + pos > 1 ? "," : "", esc, refs[i].start, refs[i].end); if (w <= 0 || (size_t)w >= cap - pos) { return; /* cannot happen: cap covers the longest entry */ } diff --git a/internal/cbm/cbm.h b/internal/cbm/cbm.h index 598d10b563..1c3b0a64a1 100644 --- a/internal/cbm/cbm.h +++ b/internal/cbm/cbm.h @@ -271,7 +271,9 @@ typedef struct { * (`#if`/`#else` twins, a macro redefined per platform, overloads): the * graph keeps one node per QN, and this lists every one of those * definitions' line spans, the surviving one included, as a JSON array - * sorted by start line: [{"start":10,"end":14},{"start":20,"end":26}]. + * sorted by start line, in graph_buffer.c's variants schema: + * [{"file_path":"a.c","start_line":10,"end_line":14}, + * {"file_path":"a.c","start_line":20,"end_line":26}]. * Carried by the definition the graph keeps for the file (the last by * start line), which makes it the node's `variants` property. NULL on * every other definition, and in every language outside diff --git a/tests/test_extraction.c b/tests/test_extraction.c index c5057e7f72..5c85441ead 100644 --- a/tests/test_extraction.c +++ b/tests/test_extraction.c @@ -7743,15 +7743,18 @@ TEST(extract_c_variants_lists_same_file_duplicates_c1) { ASSERT_NOT_NULL(kept); ASSERT_NOT_NULL(other); ASSERT_NOT_NULL(kept->variants); - ASSERT_STR_EQ(kept->variants, "[{\"start\":2,\"end\":2},{\"start\":4,\"end\":6}]"); + ASSERT_STR_EQ(kept->variants, "[{\"file_path\":\"v.c\",\"start_line\":2,\"end_line\":2}," + "{\"file_path\":\"v.c\",\"start_line\":4,\"end_line\":6}]"); ASSERT_NULL(other->variants); /* one carrier per group: the node the graph keeps */ const CBMDefinition *lim_a = c1_def_at(r, "Macro", "LIMIT", 9); const CBMDefinition *lim_b = c1_def_at(r, "Macro", "LIMIT", 11); ASSERT_NOT_NULL(lim_a); ASSERT_NOT_NULL(lim_b); - char want[128]; - snprintf(want, sizeof(want), "[{\"start\":%u,\"end\":%u},{\"start\":%u,\"end\":%u}]", + char want[256]; + snprintf(want, sizeof(want), + "[{\"file_path\":\"v.c\",\"start_line\":%u,\"end_line\":%u}," + "{\"file_path\":\"v.c\",\"start_line\":%u,\"end_line\":%u}]", lim_a->start_line, lim_a->end_line, lim_b->start_line, lim_b->end_line); ASSERT_NOT_NULL(lim_b->variants); ASSERT_STR_EQ(lim_b->variants, want); diff --git a/tests/test_pipeline.c b/tests/test_pipeline.c index 135560799b..5aeb49c1ac 100644 --- a/tests/test_pipeline.c +++ b/tests/test_pipeline.c @@ -10536,9 +10536,12 @@ static c1_variants_obs_t c1_observe_variants(const char *repo, const char *db_na return obs; } obs.opened = true; - obs.pick_lists_both = c1_props_contain( - s, proj, "v.pick", "\"variants\":[{\"start\":2,\"end\":2},{\"start\":4,\"end\":6}]"); - obs.limit_has_list = c1_props_contain(s, proj, "v.LIMIT#macro", "\"variants\":[{\"start\":9,"); + obs.pick_lists_both = + c1_props_contain(s, proj, "v.pick", + "\"variants\":[{\"file_path\":\"v.c\",\"start_line\":2,\"end_line\":2}," + "{\"file_path\":\"v.c\",\"start_line\":4,\"end_line\":6}]"); + obs.limit_has_list = c1_props_contain(s, proj, "v.LIMIT#macro", + "\"variants\":[{\"file_path\":\"v.c\",\"start_line\":9,"); obs.once_has_list = c1_props_contain(s, proj, "v.once", "\"variants\""); cbm_node_t *nodes = NULL; int count = 0; @@ -10680,8 +10683,10 @@ TEST(pipeline_cpp_overloads_are_listed_as_variants_c1) { int scale_nodes = n; int kept_start = n > 0 ? nodes[0].start_line : -1; cbm_store_free_nodes(nodes, n); - bool lists_both = c1_props_contain( - s, proj, "ov.scale", "\"variants\":[{\"start\":1,\"end\":1},{\"start\":2,\"end\":4}]"); + bool lists_both = + c1_props_contain(s, proj, "ov.scale", + "\"variants\":[{\"file_path\":\"ov.cpp\",\"start_line\":1,\"end_line\":1}," + "{\"file_path\":\"ov.cpp\",\"start_line\":2,\"end_line\":4}]"); bool once_has_list = c1_props_contain(s, proj, "ov.once", "\"variants\""); cbm_store_close(s); From 8a97baabcb5d4a52f2c2439ee58535aaa548db4a Mon Sep 17 00:00:00 2001 From: Martin Vogel Date: Thu, 8 Oct 2026 01:37:03 +0200 Subject: [PATCH 7/7] fix(extract): parse the first-branch projection under the terminated-last-line rule Merging main brought #2340's virtual_newline field in CBMStringInput. cbm_rescue_defs_from_projection still built its input with the two-field initializer, a missing-field-initializers error under -Werror. The projection keeps the raw source's length and line structure, so it now goes through cbm_parse_source like the raw parse and the C++ branch views: the same virtual newline, and the same edit that removes it again. extraction pipeline graph_buffer store_search mcp registry edge_types_probe parse_coverage c_lsp test_impact_engine conditional_variants on the merge result: 2277 passed, 4 skipped. Signed-off-by: Martin Vogel --- internal/cbm/cbm.c | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/internal/cbm/cbm.c b/internal/cbm/cbm.c index 3dab239d9d..e0d5cdd612 100644 --- a/internal/cbm/cbm.c +++ b/internal/cbm/cbm.c @@ -2571,10 +2571,10 @@ static void cbm_rescue_defs_from_projection(CBMExtractCtx *raw_ctx, const TSLang TSTree *tree = NULL; if (parser) { ts_parser_reset(parser); - CBMStringInput input = {projected, (uint32_t)raw_ctx->source_len}; - TSInput ts_input = {&input, cbm_string_read, TSInputEncodingUTF8, NULL}; TSParseOptions opts = {0}; - tree = ts_parser_parse_with_options(parser, NULL, ts_input, opts); + /* The projection keeps the raw source's length and line structure, + * so it is parsed under the same terminated-last-line rule (#2078). */ + tree = cbm_parse_source(parser, projected, (uint32_t)raw_ctx->source_len, opts); } CBMHashTable *held = tree ? cbm_ht_create(CBM_SZ_256) : NULL; if (held) {