Files
agenticCode/x-docs/agent-api-usage-ac-implementation.md
Ingo Schnabel 234ea76911 Roadmap
2026-09-23 15:31:45 +02:00

211 KiB
Raw Blame History

Using AgenticCode on This Repo (Dogfooding)

This repo is ingested as project ac at http://localhost:8787. Per CLAUDE.md's dogfooding rule, prefer the AgenticCode API over grep/Explore for call graphs, callers/callees, DB access, dataflow, and module overviews when working on this repo — it's exactly the tool this project builds.

For full API semantics (params, response shapes, error codes, language applicability, curl examples), see x-docs/agent-api-system-prompt.md — that file is the canonical reference and is not duplicated here. This file only covers what's specific to using the API as Claude Code, on this checkout.

Tool priority

REST first (GET /api/projects/ac/...) → ac CLI (ac callers, ac callees, ac call-tree, ac context, ac db-accesses, ...) → grep/Explore. Only fall back past REST/CLI when the question genuinely isn't answerable by this API at all (see "Missing capability" below). If the server is unreachable, try ./manage-ac.sh deploy before falling back further.

Re-ingest before trusting results

Query results reflect the last ingest, not the current working tree. Refresh after code changes before trusting query results: ac refresh or POST /api/projects/ac/refresh (add --deep / ?deep=true for a full field-level pass). refresh is the single (re-)ingest surface (item 42) — the eager ingest-all/ingest-module/ingest-call-graph endpoints were removed. ac refresh <name> deep-ingests one module + its callees/data areas; add --neighborhood (POST /refresh/{name}?scope=neighborhood) to also pull in the module's transitive callers (whole call-graph neighbourhood).

Reconciliation on re-ingest (item 58). A refresh now purges stale nodes: for every re-parsed file it deletes the nodes the fresh parse no longer produces (renamed/removed fields, moved statements) rather than leaving them to shadow the new ones — so identifier counts and /search/identifier results stay clean after a parser change or an edited source file. Applies to every full-parse path (whole-root refresh with or without --deep, and per-module refresh/{name}); the coarse Tier-1 scan run at project creation does not reconcile (it only ever runs on an empty graph).

Auto-invalidation (item 43). The graph self-heals when files change or are deleted on disk — you rarely need a manual refresh for staleness:

  • Changed file: a field-level query on a module whose source changed re-ingests it automatically (its stored sourceHash no longer matches), so results reflect the current file. Granularity is per-module: a changed dependency is picked up when that dependency is itself queried by name.
  • Deleted file: a whole-project refresh now also removes nodes for files deleted from disk (reconciled against the filesystem), closing the earlier gap — no project re-create needed. (A per-module refresh/{name} does not sweep; it only touches its own tree.)
  • Controlled by agenticcode.auto-invalidate.enabled (default true).

Reading source in this repo

You have direct filesystem access to this checkout — never call nodes/{id}/source or modules/{name}/source. Use sourceFile/startLine/ endLine from a graph response (context, digest, search/identifier, nodes/{id}, ...) and read the file directly. This is always cheaper and gives full surrounding context; the /source endpoints exist only for API-only agents with no filesystem access.

Stale-source check (item 41). The /source endpoints compare the file on disk against the content hash (sourceHash) stored at ingest. If the file changed since the last ingest they return 409 STALE_SOURCE instead of slicing current text against old line numbers — re-ingest (refresh) the project to update the graph. Line ranges you read directly off disk are of course always current; this only guards the API's own slicing. Copycode/INCLUDE slices are raw pre-expansion file text.

Comments: reachable, but never by default (item 141)

Every default query sees code only. search/identifier, search/annotation, search/references and a plain search/value never match comment text — a Natural * ... banner, a trailing /* ..., a Java // line or a Javadoc block.

That silence used to be a wrong answer, not a missing one, wherever a convention records something in a comment. The UPMS→PUR case: a reengineered service carries its Natural origin in a Javadoc block (ServiceEndpoint: / UPMSFunction: / UpmsObject:), and the documented Natural→Java lookup searches the program name in pur, where an empty result is read as "not yet reengineered" — which it answered for services reengineered months earlier.

Three routes now reach comment text. Pick one before concluding "not present":

Question Call
"What does this module's header/change log say?" GET /modules/{name}/comments (ac comments <module>) — blocks with the declaration each documents
"Where is this query / SQL / JSON text defined?" (Java) GET /search/value?value=…&contains=true — since item 139 a static final String built from text blocks, literals, same-class constants and + carries its full text as value; a "…".formatted(...) constant carries its template (with %s) and valueKind: template. Anything built from a method call or another class's constant stays unresolved (no value)
"Does this string appear anywhere, code or comment?" GET /search/value?value=…&includeComments=true (ac search-value --include-comments) — comment hits carry kind: "COMMENT"
"…and in text the parsers do not model at all, or in a module that is not deeply ingested?" GET /search/source?regex=… — raw grep over the files on disk
GET /pur/search/value?value=WPARTX0S&contains=true                        → []          (code only)
GET /pur/search/value?value=WPARTX0S&contains=true&includeComments=true   → the Javadoc origin block
GET /upms/modules/WAGNTX0S/comments                                       → `* #01 … Bug 266`, …

Comments stay opt-in deliberately: a comment hit is not the same evidence as a literal in code, and folding them into the default result set would move every existing completeness count (item 131's lesson). The flip side is the rule to remember — an empty default search says nothing about comments.

Diagnosing a slow refresh (items 153/155)

Two instrumentation layers, both aimed at the same question — where does the time go?

  • Always on: every persist batch logs one line with its statement breakdown (Persist deep 200 files: total 8740 ms = prepare 108, merge-nodes 1630, ..., commit 20), and every enrichment step logs duration plus created/deleted rows. The residual commit is deliberate: total minus the labels is the transaction commit, so nothing hides in an unnamed remainder.
  • Opt-in per run: POST /refresh?deep=true&profile=true (ac refresh --deep --profile) runs the enrichment steps under Cypher PROFILE and logs the five heaviest operators of every step slower than 5 s. Diagnostic only — it answers "is this step matching or writing?", which the step timing alone cannot. Measured overhead on upms: none worth reporting (908 s vs 905 s).
  • Build-time diagnostic: agenticcode.ingest.comments.enabled=false drops COMMENT nodes and their DOCUMENTS edges just before persist. It exists to measure what comments cost (item 179: ~17 % of a deep refresh) and is not a supported operating mode — with it off, /comments answers empty. It is a property, not a request parameter, so it needs a rebuild and cannot be set per run.

What that measured, so nobody re-derives it. The figures below are current (2026-09-06, two clean runs, roadmap item 175); the campaign of items 153-175 took an upms deep refresh from 1 225 s to 301 s, so any older number quoted elsewhere is stale by a factor of four.

now
deep refresh upms, end to end 301 s (run-to-run spread 3.7 %)
persist ~149 s — merge-edges 47.6 s, commit 34.1 s, merge-nodes 24.0 s
finalize ~137 s — resolve-field-placeholder W/R 43.4 s, link-args-to-params 23.7 s, resolve-bare-included W/R 37.9 s
parsing ~9 s (interleaved with persist)

The original diagnosis still holds and is why those steps shrank: the five field-resolution steps were dominated by matching, not writing — resolve-bare-included once spent 273 M database hits to produce 43 753 rows. See roadmap items 156 and 175, and item 179 for what comment nodes cost.

limit/offset work on some endpoints and are silently ignored on others

Every endpoint answers completely. The difference is whether it lets you ask for less. 19 list endpoints declare no limit/offset at all, and JAX-RS drops an undeclared query parameter without a word — so ?limit=2 there is not an error, it simply has no effect:

GET /upms/modules?limit=2                        -> 3 587 rows
GET /upms/modules/WAGNTX0S/functions?limit=2     ->    22 rows
GET /upms/modules/WAGNTX0S/data-structures?limit=1 ->  16 rows

Ignore limit/offset (always the full list): /modules · /modules/{name}/functions · /modules/{name}/functions/overrides · /modules/{name}/functions/{function}/overrides · /modules/{name}/data-structures · /modules/{name}/columns · /modules/{name}/payload · /modules/{name}/dispatch-table · /modules/{name}/sql-statements · /data-structures/{name}/fields · /db-tables/{name}/columns · /variables/{name}/reads · /variables/{name}/writes · /variables/{name}/flow-forward · /variables/{name}/flow-backward · /variables/{name}/field-flow · /duplicates · /dynamic-calls/unresolved · /dynamic-calls/overrides

Honour them (14, verified against the method signatures): the search endpoints (search/identifier, search/value, search/annotation, search/references, search/source), rest-endpoints, and per module callers, callees, context, db-accesses, graph, reaches, workfile-accesses, comments.

call-tree is bounded differently again — by depth, not by row count — which is the right shape for a tree but means limit does nothing there either.

Which direction the mistake runs matters, so be precise about it: a dropped limit means you get more than you asked for, never less. It costs tokens, never correctness — the opposite of item 131's failure, where a silent 50-row cap was read as the complete set. Nothing here can under-report.

The practical consequence is budget, not trust: /modules on upms is 3 587 rows in one response. Narrow with the filters those endpoints do have (?kind=, ?sourceFile=, ?module=, ?extendsName=) rather than with a limit that will be ignored, and prefer /modules/{name}/digest or /context when you want an overview rather than an enumeration.

(Implementing limit on those 19 was considered and deliberately not done: the endpoints are honest as they stand, and a limit that ever acquired a default would reintroduce exactly the silent truncation item 131 removed.)

callers / callees say when they are cut (item 181)

callers, callees and functions/{fn}/callers now send X-AC-Total-Count and X-AC-Truncated like the search endpoints, and their body carries total and truncated next to sourceFiles / items (also for fields=name; ac callers / ac callees print the usual "truncated" warning). The default page is still 50: upms/modules/DPARTFN0/callees answers 50 of 58 with X-AC-Truncated: true — before, the 8 missing callees (among them the module that writes the partner) looked like an inconsistency with digest. Read truncated before concluding "X does not call Y"; ask with limit=1000 or narrow with scope=external.

Truncation is now visible on the search endpoints (item 131)

search/identifier, search/value, search/annotation, search/references and rest-endpoints send two headers with every answer:

Header Meaning
X-AC-Total-Count how many rows match in total, ignoring limit/offset
X-AC-Truncated true when this page leaves some out

and all five accept ?countOnly=true (CLI --count-only), returning {"count": n} instead of rows — a completeness question is a counting question, and @Column on pur is 3 630 rows ≈ 250 k tokens if you ask for them.

This closes item 103's own follow-up. The bodies stay bare arrays (no contract change), for the same reason as item 130's scope headers. Why it matters: search/annotation?name=Immutable returned 50 of 95 rows with no total, no flag and no Link/X-Total-Count — and a real UPMS→PUR audit read that page as the whole set, recording that 17 entities had lost @Immutable when zero had. Re-checked against all 114 rows of Tables_meta.csv: 94 non-writable carry it, 20 writable do not, no deviations. The finding cost a day and was pure artefact of the cut.

search/references and rest-endpoints joined this late (item 135, 2026-08-20). Item 131 was written about the annotation search and both were overlooked. search/references was the damaging one: it capped at the default 50 and said nothing at all, so a rename scoped from that page missed every site past the fiftieth and looked complete doing it. rest-endpoints defaults to an uncapped limit and so never lost rows, but it was equally silent about how many there are. Both were found by x-scripts/verify-api.sh on its first run, not by a test.

Two mechanics worth knowing: the total costs a second query only when the page comes back full (a short page is provably the end, so the total is arithmetic), and a total that divides evenly by limit makes the last full page report truncated with the next page empty — one wasted call, never a wrong answer. ac prints a note to stderr when a response is flagged truncated, so piping the body into jq stays clean.

REST surface and scope headers (item 130)

GET /api/projects/{p}/rest-endpoints?module=&countOnly=&limit=&offset= (CLI ac rest-endpoints) lists {httpMethod, path, module, moduleSimpleName, handler, sourceFile, startLine} — the composed path (class-level @Path + method-level @Path), so "which code runs for POST /partners" is one call. Previously the two halves had to be joined by hand from two /search/annotation calls, because annotations are stored by name without their arguments; the parser now persists restPath and httpMethod, including a @Path written as a constant reference. A method with no HTTP-verb annotation is not an endpoint and is excluded. A class with no @Path of its own inherits the nearest one from its extends/implements ancestry, as JAX-RS does. Rows carry outbound: true when the declaring type is a @RegisterRestClient interface — a call the application makes, not one it serves; its path is usually empty because the base URI comes from configuration.

Scope and freshness now ride on every project-scoped response as headers:

Header Meaning
X-AC-Exclude-Dirs directories the ingest skipped, or (none)
X-AC-Ingested-At when the graph was last walked (item 126)
X-AC-Ingest-Incomplete true while a whole-root pass runs or after one that never finished (item 129); unknown when no ingest was ever recorded

Read X-AC-Exclude-Dirs before trusting an empty answer: "no callers" means "none outside tests" in a project excluding test (app) and "none at all" in one that does not (pur, ac) — the bodies are identical.

They are headers, not body fields, because most endpoints answer with a bare JSON array (db-accesses, functions, search/identifier, …); adding a field there would mean restructuring array → object and breaking the web UI's generated client, the CLI printers and any agent that indexes [0]. The trade-off is that an agent reading only the JSON body will not see them — so if you consume this API programmatically, read the headers too. The project shell is cached for ~10 s to keep this off the request's critical path, and the ingest path invalidates that cache explicitly, so X-AC-Ingest-Incomplete flips as soon as a refresh starts rather than up to 10 s later.

Every reference site of a name (item 128)

GET /api/projects/{p}/search/references?name=&kind=&countOnly=&limit=&offset= (CLI ac references <name>) returns {sourceFile, lineNo, kind, inModule, target} per mention of a type — not just per call:

kind Where it comes from
CALL a call site (CALLS)
IMPORT an import of the type
TYPE a declared field / parameter / return type
ANNOTATION the type used as an annotation
EXTENDS, IMPLEMENTS inheritance
INJECTS CDI wiring
CLASS_LITERAL X.class in argument position
INCLUDE Natural copycode inclusion

Use it to scope a rename. callers sees calls alone, so a file that only imports the class, declares a field of it, or names it in an annotation was invisible — and the rename that missed it looked complete. The name may be the identity (FQN) or the short form; target echoes what it resolved to. An unknown kind is 400 INVALID_KIND, never an empty list.

Known limits, by design:

  • Local-variable types and generic type arguments are not indexed — List<Target> x records List, not Target. They multiply edge volume for much less value than the positions above.
  • Same-package references have no import, so within one package the index rests on declared-type positions alone.
  • Imports are only indexed when they look project-internal (they share the first two package segments with the importing file). Otherwise every java.util/framework import would mint a placeholder node on every ingest, just for the finalize sweep to delete it again.
  • Natural has no import or type-position concept. It contributes CALL, INCLUDE and inheritance kinds only; this is not parity with Java and should not be read as such.
  • Reference edges are written at parse time, so they only exist for files re-parsed since this landed — a project needs a refresh before the index is complete.
  • Mentions use their own MENTIONS edge type, kept out of CALLS/REFERENCES deliberately: the call-graph traversals (callers, callees, call-tree, ego-graph) follow REFERENCES as wiring, so folding imports into it made an import surface as a caller. /search/references is the only endpoint that reads MENTIONS; the call graph is unchanged.

Refreshing only what changed (item 129)

POST /api/projects/{p}/refresh (CLI ac refresh) has two ways to avoid re-walking a whole root:

Form What it does
?paths=a/B.java,c/D.java (ac refresh --paths a/B.java,c/D.java) Re-ingests exactly those relative paths, deep, plus their dependencies. Paths that match no file — or that are not ingestible source files at all, like pom.xml — come back in unresolved; a typo'd path is never silently dropped. Does not run the deleted-file sweep and does not move ingestedAt: both need a whole-root walk.
?changedOnly=true (ac refresh --changed-only) Whole-root walk, but re-parses only files whose content hash differs from the graph's (files with no stored hash count as changed).

changedOnly is opt-in on purpose. Three whole-walk behaviours are reduced, and one of them would be outright corruption if it were hidden:

  • A changed Natural copycode disables skipping for that entire run (logged). Natural projects only — a .cpy sitting in a Java project (as a test fixture, say) is not inlined by anything and no longer stands the optimisation down. Copycode text is inlined into the including module at parse time, so a module whose .cpy changed parses differently while its own hash is unchanged — skipping it would leave a stale expansion behind with nothing to indicate it.
  • Duplicate-identity detection only sees the changed files, so it can confirm duplicates among them but not discover new ones elsewhere. Existing markers are never cleared.
  • User-exit LoC annotation (item 47) is re-stamped only on re-parsed files.

Enrichment is project-wide and still runs in full, so this cuts parse+persist time only — not the finalize pass. On small projects the whole deep refresh is already ~30 s, so measure before assuming a win.

ingest.incomplete (on /projects and /projects/{p}) is true while a whole-root pass runs and stays true if one never finished — a crash, a container stop, an aborted deep refresh. Before this, an interrupted deep refresh was indistinguishable from a clean graph: the enrichment steps that already ran are committed, so queries keep answering, just from a half-updated graph. It cannot self-heal (a killed process clears nothing) and does not distinguish "running right now" from "died an hour ago" — both mean the same thing to a caller. A completed refresh clears it.

Renaming a project (item 202)

POST /api/projects/{p}/rename with {"newName": "…"} (CLI ac project rename <old> <new>) moves the project key on every node and override and in other projects' counterparts lists, then the shell; the old name answers 404 afterwards and nothing needs re-ingesting. Refusals: 400 INVALID_REQUEST (blank or unchanged), 404, 409 PROJECT_EXISTS. The node rewrite is batched; an interrupted rename is finished by running it again (97 s for the 939k-node upms). Use it to keep a reference graph next to a fresh ingest (ac project rename upms upms_alt, then create upms again and compare). A full DELETE of a project now also removes its manual overrides; recreate keeps them.

Is this project's graph any good? (item 126)

GET /api/projects and GET /api/projects/{p} (CLI ac project list / ac project show <p>) carry an ingest object describing the last whole-root ingest:

"ingest": { "ingestedAt": "2026-08-18T10:12:44Z", "mode": "full", "filesExamined": 2981,
            "filesPersisted": 2977, "filesFailed": 4, "failures": ["a/B.java", "..."],
            "failuresTruncated": false, "durationSeconds": 176, "serverVersion": "…" }

Use it before trusting a negative answer: without it, "no such module" and "that part of the project was never ingested" are the same empty response. Three rules the field obeys:

  • ingest: null means never recorded, not "ingested nothing" — a project last walked before this existed reads as null rather than as a fabricated zero.
  • Only whole-root passes write it — the create-time Tier-1 scan, refresh, refresh?deep=true. A by-name refresh/{name}, a deep ingest or a fan-out warm ingests real files but sees a fraction of the tree, so it deliberately leaves ingestedAt alone; otherwise deepening one module would advertise the whole project as freshly walked.
  • ingestedAt is not a freshness guarantee. It says when the walk ran, not that the graph still matches disk — a file edited a minute later is stale while the timestamp still looks recent. For the real check, read a file through GET /{p}/source?file=…, which answers 409 STALE_SOURCE when the content no longer matches the ingested hash. (Targeted/incremental refresh is item 129.)

failures is capped at 200 paths while filesFailed stays exact; failuresTruncated says whether the list was cut, so a short list is never mistaken for the whole story.

Tier-1 coarse scan on project create (item 36)

Creating a project (POST /api/projects/{p}) now runs a Tier-1 coarse reference scan of the root before returning, so the project is immediately queryable — no separate ingest call. The scan is lexer-level (Natural) / a coarse projection of the full parse (Java): it records module/function /data-structure shells, the identifier index (declared fields, class members), and coarse CALLS/READS/WRITES/INCLUDES references (Natural PERFORM, CALLNAT '...', dynamic CALLNAT PGM-VAR, PARAMETER/LOCAL USING copybooks; Java resolved calls/type refs) — but no deep bodies (control flow, statement-level dataflow, arg→param). Scanned modules land CALL_GRAPH/NOT_INGESTED; field-level detail is filled in by the on-demand deep ingest below. Each module shell carries a sourceHash. Disable with agenticcode.tier1.scan-on-create=false (creates an empty project shell to scan later). A missing/unscannable root is non-fatal — the project is created empty and can be re-scanned.

Unresolved references (item 40)

A reference whose target isn't (yet) ingested — a CALLNAT/PERFORM/USING to a module/copybook absent from the project, or a dynamic CALLNAT PGM-VAR whose literal can't be recovered — is stored as a deduped placeholder node (blank sourceFile). Enrichment stamps each with an unresolved boolean: true while genuinely dangling, false once a real definition of that name is ingested. /search/identifier returns it as unresolved on each IdentifierMatch, and GET /nodes/{id} carries it in the node's properties — so an agent can tell a dangling/dynamic reference apart from a resolved one. In callers/callees such targets already appear as entries with a blank sourceFile.

Querying a module endpoint: four answers, not one (items 107, 115, 114)

Every GET /api/projects/{p}/modules/{name}/… endpoint used to answer 200 with an all-zeros shell for a name that exists nowhere in the graph, byte-identical to a real-but-empty module's answer. That is not cosmetic: call-tree?depth=4 returning 200 with 0 modules reads as "analysed, nothing found" when the truth is "not analysable" — the module's source was never in the checkout. The graph knows four states and each now gets its own status:

state answer what it means
no MODULE node for the name 404 MODULE_NOT_FOUND unknown name (typo, or not in this project)
placeholder — node exists because something calls it, source never parsed 409 + {status:"NOT_INGESTED", module, detail, nextAction} knowable in principle, not analysed yet
ambiguous — several real modules share the name (Java) 409 + {code:"AMBIGUOUS_NAME", details:{candidates:[...]}} the name does not identify a module; pick one
duplicate identity — the same name in several files, skipped at ingest 409 + {code:"DUPLICATE_IDENTITY", details:{paths:[...]}} exists twice, deliberately not ingested (item 114); see /duplicates
real, ingested module 200 the body is the answer — an empty body genuinely means "nothing found"

Ambiguous names (item 115). A Java simple name is not unique: nested @Nested test classes, Builder, Config, WorkingStorage. Measured on pur, 163 names covering 385 modules (~8%) were addressable only ambiguously, and the endpoints used to answer with the union across unrelated classes — /modules/BrokerHistoryTests/functions returned 627 functions for a class that has 107. They now refuse and list the candidates. Repeat the request with ?sourceFile=<candidate>:

GET /modules/Shared/functions                          → 409 AMBIGUOUS_NAME, candidates ["a/Shared.java","b/Shared.java"]
GET /modules/Shared/functions?sourceFile=a/Shared.java → 200

A Java module's name IS its fully-qualified name (item 117) — com.example.OrderService, and com.example.Outer.Inner for a nested class. That is what makes same-simple-name classes distinguishable at all. Every module endpoint accepts either form:

GET /modules/com.example.OrderService/digest   → 200, always exact
GET /modules/OrderService/digest               → 200 when unique, else 409 AMBIGUOUS_NAME

Responses carry simpleName alongside name for display. The ?module= and ?extends= filters and the ac CLI take either form too; --source-file remains available on every module command.

Which Java types are modules (item 119). Classes, interfaces, enums, records and annotation types — moduleKind is one of CLASS | INTERFACE | ENUM | RECORD | ANNOTATION (Natural adds PROGRAM | SUBPROGRAM | …), and ?moduleKind= filters on it. Their content is modelled the way each kind carries it: a record's components and an annotation type's members are FIELDs (the latter with defaultValue where declared), an enum's constants are CONSTANTs, and an enum's or record's implements is a real edge — so a call against an interface fans out to an enum implementing it.

Before 119 these three kinds were not parsed at all: GET /modules/SomeEnum/digest answered 404, and a record referenced from elsewhere stayed an unresolved placeholder (409 NOT_INGESTED). One gap remains by design — a record's compact canonical constructor is not a function node, so calls made in its body are invisible.

Natural is unaffected throughout: its module names are file stems, and colliding identities are skipped at ingest, so they are unique by construction (upms has zero ambiguous names, pur 163).

One limit worth knowing: only the request's root module is disambiguated. Inside a traversal this no longer merges anything, because module names are unique per project after item 117 (pur and upms have zero duplicate names) — a reference the parser could not qualify is returned flagged unresolved: true rather than attached to an arbitrary candidate.

A name skipped at ingest because it exists in more than one file answers 409 DUPLICATE_IDENTITY with the conflicting details.paths (item 114); GET /duplicates lists them all. Note what that does not fix: the skipped file's own calls were never parsed, so caller lists elsewhere can still be short — they just no longer look complete.

The 409 applies to everything derived from the module's own source: digest, context, call-tree, callees, db-accesses, workfile-accesses, sql-statements, functions, functions/overrides, functions/{fn}/overrides, functions/{fn}/callers, data-structures, dispatch-table, payload, columns.

callers and graph stay 200 for a placeholder — their data comes from the calling modules' source and is genuine. When you get a 409, fall back to /callers: it is the one honest answer available for a module whose own source is missing. (digest no longer surfaces those callers, since its other fields would all be structurally zero.)

/modules/{name}/source is unchanged: it already answered 404 MODULE_NOT_FOUND for both an absent module and a placeholder, since there is no source to serve either way.

Scope limit — this guards the root module of a request only. A call-tree that traverses into placeholder targets still reports that subtree as empty without flagging it, so a dispatcher whose targets are all placeholders still returns 200 with a silently truncated tree. Cross-check the targets you care about individually (a 409 tells you it is unanalysed) — see item 103.

Data literals are not call targets (item 62). A CALLNAT <bareword> whose target is really a data value — a browse key reaching the call site through a copycode/macro argument — used to leave a permanent unresolved placeholder callee. Enrichment now reaps such a placeholder when it has no Natural sigil (#/&/+), is a real VARIABLE/CONSTANT of the project, matches no real MODULE, and is only ever reached by inferred (CALLNAT_DYNAMIC/INCLUDE_MACRO) edges. So callees, call-tree, the ego graph and the unresolved list no longer list data names as modules. Genuine dynamic dispatch (CALLNAT #PGM-VAR, sigil'd) is still reported as an unresolved target, and a static CALLNAT 'X' is always trusted even when X collides with a field name.

The ingest summary agrees with the graph (item 64). A refresh/refresh/{name} response's unresolved list is built during the file walk, independently of the graph — before item 64 it therefore reported data fields as missing modules (MODULE CO-TABLA, MODULE NAME-DESC-SP) that enrichment had already reaped, so the summary and the graph disagreed. It now applies the same item-62 gates as the graph reap, with the "is this a real field?" test asked project-wide (a data literal is often declared outside the ingested tree). Genuinely missing modules are still reported — including sigil'd dynamic dispatch (MODULE #GETSHORT-MODUL) and a missing module whose name collides with a field name but is called statically (the RPC-CNTX class). Treat unresolved as "dependencies that really are absent".

Constant-folded string-assembled targets (item 83). A dispatcher often builds the CALLNAT <var> name from a base literal plus one or more SUBSTR overlays — e.g. #GETSHORT-MODUL in YGEAGGNH, assembled by MOVE 'YGEAGKEY' TO #GETSHORT-MODUL then MOVE 'GN0' TO SUBSTR(#GETSHORT-MODUL,6,3) → YGEAGGN0. The parser records each SUBSTR write as a WRITES carrying substrPos/substrLen (1-based), and the resolve-dynamic-callnat-fold enrichment step folds the last full-var literal written before the call site with the intervening overlays (left/substring) and MERGEs a resolved CALLS edge (callKind=CALLNAT_DYNAMIC, folded=true) to the assembled module when it is a real MODULE. It runs in every ingest mode (cheap, bounded by dynamic call sites) and over-approximates like the other dynamic resolvers. This auto-recovers the Y…GNH → Y…GN0 family with no manual override, so folded sites drop out of dynamic-calls/unresolved and the assembled target appears in callees/call-tree/graph/ego graph tagged CALLNAT_DYNAMIC.

Pin what the resolvers can't: manual dynamic-CALLNAT overrides (item 82). Some CALLNAT <var> targets are beyond the auto-resolvers — a name never present as a recoverable literal in the reachable code (read from a work-file record, supplied by an unlinked caller, …). Such a site stays an unresolved placeholder. A human or agent resolves it via POST /api/projects/{p}/dynamic-calls/overrides with the call site's originFile + lineNo (from GET .../dynamic-calls/unresolved) and the target module name(s) — multiple targets for a genuine branch. The override is stored as a :DynamicCallOverride node outside the :AstNode graph, so a refresh never deletes it and an enrichment step (apply-manual-dynamic-callnat, after the auto dynamic-CALLNAT resolvers, before the placeholder cleanup) re-applies it automatically — MERGEing a CALLS edge (callKind=CALLNAT_DYNAMIC, resolvedBy='manual') to each target and flagging the placeholder manualHidden so callees/digest/graph/call-tree show the real target, not the #var. It only applies while the site is still unresolved: once an auto-resolver catches up, the override is skipped and listed obsolete — except a constant-fold (item 83), which a manual override outranks: the fold skips a site carrying a :DynamicCallOverride, and a stale folded edge there is dropped (delete-folded-overridden-dynamic-callnat) before apply-manual-dynamic-callnat runs, so the pinned target replaces it. DELETE .../dynamic-calls/overrides?originFile=&lineNo= resets one site (omit both = all), clearing the manual edges and un-hiding the placeholder inline (no refresh). A target that is not a real MODULE is rejected 400 UNKNOWN_TARGET. Bug B fix: the callees items now carry unresolved (mirroring what graph already exposed), so an unresolved dynamic target is machine-distinguishable from a resolved one without inspecting sourceFile. Since item 200 the apply and the reset also rebuild the calling modules' derived CALLS_MODULE edges in the same transaction, so reaches and field-flow honour a pinned target without a refresh, like callees and call-tree already did.

Dispatch guards: read guards — it is the only complete condition (item 72). A dispatch-table row's guardField/guardValue/guardValues describe the innermost DECIDE only. Natural nests value-DECIDEs inside each other (824 of upms's 7,729 blocks, 178 modules), and then the innermost guard is just one conjunct: in VMULTMN4, the row for YTABLMA0.TX-TABLA reports #FIELD-NAME = 'TX-TABLA', but the assignment also requires #SHORT-VIEW = 'TABL'. guards is the full chain — [{field, values}], outermost first, joined by AND, each link's values joined by OR. Reading only the legacy fields over-generalises: port that to Java and you get a branch firing where Natural never would. For an unnested DECIDE the chain has one link and says the same as the legacy fields.

  • Still incomplete for NONE/ANY branches (item 73): an assignment in a NONE branch is reported under its enclosing chain alone, but its real condition is "enclosing guard AND NOT any sibling VALUE" — a negation a chain of equalities cannot express. guards is strictly better than the legacy fields, not a total answer.

A dispatch row's lineNo belongs to sourceFile, not to the module (item 122). dispatch-table rows now carry the same provenance quartet as callees/db-accesses/workfile-accesses/functions: sourceFile, viaCopycode, includedAt, includePath. Resolve lineNo against sourceFile — when viaCopycode is non-null the assignment is written in that copycode and lineNo is a line of the .cpy, while includedAt is the INCLUDE line in the module. Before item 122 the row carried only lineNo, so following it against the module file landed somewhere arbitrary: 26 of VCOMIN50's 44 rows reported line 18 or 20, which in that module is a change-history comment; the real sites are ISICINDE.cpy:18 and ISICINDI.cpy:20. Note the guard chain may span the include boundary — the outer DECIDE in the module, the inner one in the copycode — so a row can have a multi-link guards chain whose links live in different files.

Within one guard: prefer guardValues over guardValue (item 64). dispatch-table rows carry both. guardValue is lossy and kept only for compatibility: it comma-joins the branch's VALUE literals, which drops blank alternatives entirely and, for a multi-alternative branch, yields a synthetic string the guarded field never equals ("A1, A2"). guardValues is the faithful list — every alternative in source order, blanks included — so VALUE 'GENAGREE-WOUT-SP', ' ' reports ["GENAGREE-WOUT-SP", " "], recording that a blank guard field also routes into that branch. When reasoning about routing (or porting a DECIDE to Java), read guardValues; guardValue will silently under-report the branch's conditions.

Text that is not code never yields a call (items 61 & 63). The CALLNAT/CALLNAT_DYNAMIC patterns are unanchored (a CALLNAT may legally appear mid-line), so both ingest tiers first neutralise non-code text: a full-line * comment is skipped, a trailing /* … is stripped (item 61), and a match whose keyword falls inside a quoted string literal is rejected (item 63). A real CALLNAT 'MOD' is unaffected — its keyword sits outside the quotes. This matters for trusting callers/callees/ call-tree: before item 63, prose such as WRITE(#MSG) 'NACH CALLNAT ISINGEAG:' or #ERR-TYPE := 'Callnat USIA008N' fabricated a CALLNAT_DYNAMIC edge to the real module of that name, so a mere log message appeared as a genuine call — and, because a real module existed, it was not flagged unresolved and could not be reaped by the item-62 cleanup. If you query a graph ingested before 2026-07-16, re-ingest (ac refresh) before trusting call-graph edges into modules that are also mentioned in log/error text.

Calls made through a copycode's arguments (items 120/121/123). In Natural the target of a CALLNAT is often not written at the call site at all: a copycode receives the module name as a positional INCLUDE argument and issues CALLNAT &2&. Three defects in that argument path — arguments continued on the next line, the doubled-quote escape '''X''', and the double-quote delimiter '"X"' — meant such a call produced no edge and no unresolved-dynamic-call entry, so callees, callers, call-tree and reaches agreed on an answer that was simply absent, with nothing saying "not analysed". This hit the browse/access layer hardest, because that is where the idiom lives: YCARPBN1 and YPOLIBN1 reported 0 callers each. Fixed 2026-08-07; re-ingest recovered 2672 copycode-derived call pairs (+49%) in upms with none lost. A graph ingested before 2026-08-07 under-reports Natural callers/callees, and does so silently — re-ingest before concluding a Natural module is unused. A copycode parameter that genuinely has no argument now surfaces in /dynamic-calls/unresolved as &n& rather than being dropped, so "not analysable" is visible.

A call edge no longer outlives the call it was parsed from (item 124). Until 2026-08-09 a refresh only added the corrected call and left the old one in place, because an edge is reaped only when one of its endpoints is — and a parser fix changes neither (the calling subroutine is unchanged, the old target is a never-swept placeholder). So callees/callers could report a call that no source line makes, flagged unresolved: true and indistinguishable from a genuine unresolved dynamic call. A re-parsed Natural file's call edges are now reaped before the fresh ones are merged, and a call-target placeholder left with no callers is deleted. Two consequences for a graph ingested before 2026-08-09: an unresolved: true callee may be an artefact of an already-fixed parser bug rather than a real dynamic call, and search/identifier may list module names that exist nowhere in the source. Both clear on the next deep refresh. Note the reap deliberately spares CALLNAT_DYNAMIC edges onto real modules — those are the dynamic-call resolvers' output, not the parser's.

LoC / SLoC metrics (item 46)

Every file-level node (a MODULE program/class, or a DATA_STRUCTURE for a Natural .lda/.pda data area) is stamped at ingest with two deterministic line metrics:

  • loc — physical lines of the file (language-independent; a trailing newline adds no phantom line).
  • sloc — source lines of code: non-blank, non-comment lines, computed per language. Natural drops full-line */**//* and inline /* comments; Java drops // and /* … */ blocks while keeping those tokens when they appear inside string literals.

Both the cheap Tier-1 coarse scan and the deep Tier-2 parse use the same per-language counter, so a module's loc/sloc are identical at any ingest depth — you can sum them to get exact, reproducible project totals.

Where to read them:

  • GET /modules — loc/sloc on each row.
  • GET /modules/{name}/context — loc/sloc on the module.
  • GET /nodes/{id} — loc/sloc in the node's raw properties.
  • GET /loc (ac loc) — the rollup: a per-language breakdown (fileCount, loc, sloc) plus a project-wide total, optionally narrowed by ?language= / ?sourceFile=. Each source file is counted once even when it yields several nodes (Java inner classes, Natural inline groups).

null metrics mean the node predates item 46 — re-ingest (refresh) to backfill.

Generated vs. user-exit split (item 47)

A project can be created with a source language (required at creation; an attribute only — ingest still classifies files by extension) and a generatedDir/userExitDir pair (directory names, matched as path components like excludeDirs; both or neither). Generated modules already contain their hand-written user-exit twin inline, so at ingest a module under generatedDir whose name also occurs under userExitDir is annotated with that twin's LoC/SLoC (userExitLoc/userExitSloc). User-exit files are not ingested as standalone modules (they would collide by name) — the walk skips userExitDir.

Consequence for all non-LoC analysis: the generatedDir copy is the canonical, sole module for every structural query (call graph, DB access, functions, data structures, identifiers, dataflow, dispatch table). userExitDir exists only to compute the generated-vs-manually-written LoC split below; it never contributes nodes/edges. So when verifying an API response against source for a Natural module, always read the generatedDir file (e.g. generated_src/subprogram/WGEAGB0S.nat), not the user_exit fragment.

GET /loc (ac loc) then reports, per language row and in the project total:

  • loc/sloc — the total (generated, which already includes the user exits).
  • userExitLoc/userExitSloc — the sum of the annotated user-exit twins (the hand-written part).
  • generatedExclusiveLoc/generatedExclusiveSloc — total − user-exit, clamped ≥0 per file (the purely generated part).

All three are 0 for projects without a generated/user-exit split. Create with ac project create <name> <root> -l natural -g generated_src -u user_exit, or add the split to an existing project via ac project update <name> -g generated_src -u user_exit.

Java DB_ACCESS in a project without JPA entities (item 140, 2026-08-27)

A DB_TABLE node is only ever created from a JPA @Entity or a Panache active-record class. In a Java project that has none, no DB_ACCESS candidate can resolve — and the parser's candidate heuristic is a deliberate over-approximation: its read gate admits any static receiver whose method starts with get/find/read/list/… , so UserContext.getCurrent() and TextUtils.getColumn(line, 0, 8) become candidates. On app that produced 2219 DB_ACCESS nodes in a codebase with no database access whatsoever.

db-accesses never showed them (it joins the table with a plain MATCH), but sql-statements did: it joins with OPTIONAL MATCH, so unresolved candidates came back as rows with "table": null — 76 of them on a single app module.

Since 2026-08-27 the enrichment step reap-java-db-access-without-tables deletes every Java DB_ACCESS of a project that holds no DB_TABLE. For such a project sql-statements is now empty instead of noisy. Three things to know:

  • The gate is project-level, not per node. One entity anywhere in the project switches the reaper off, and the unresolved candidates stay. pur (2014 of 3951 unresolved) and ac (221 of 335) are unaffected, and still return table: null rows. Treat a sql-statements row whose table is null as unverified, in any project that has tables.
  • Natural is untouched. A Natural DB_ACCESS comes from a literal READ/FIND/STORE and is a real access whether or not its view resolved (upms: 14302 of 14303 resolve).
  • Recovering from it needs a full refresh. If such a project later gains its first entity, the reaped nodes only come back for files that are actually re-parsed — changedOnly will not restore them, refresh without it (or recreate) will.

?depth= means module hops (item 65)

On db-accesses / sql-statements (and the ?module= scope of variables/{name}/reads|writes), depth=N means N module calls away — the same unit /modules/{name}/graph?depth= and call-tree neighbours use. depth=1 = the modules this one directly CALLNATs, regardless of how deeply the calling statement sits inside subroutines.

Before item 65 these endpoints bounded the traversal on raw CALLS edges. A CALLS edge starts at the statement making the call, not at the MODULE node, so the traversal also stepped through internal PERFORM jumps and depth measured statement nesting, not dependency distance. Concretely: WGEAGB0S reached YGEAGBNH's tables through two module calls, but the raw path is 5 edges (WGEAGB0S → GET-DATA → GET-MAIN-DATA → BGEAGFN0 → READ-FILE → YGEAGBNH), so db-accesses?depth=2 returned [] — reading as "no DB access" — and only depth=5 was truthful.

If you scripted a depth workaround (a deliberately large depth to compensate), drop it: depth is now the value you'd naturally expect, and inflated values just widen the result set.

call-tree's depth column is still raw-hop based and mixes internal subroutines into the tree: a direct dependency called from the main body shows depth=1 while one called two subroutines deep shows depth=3. Use the ego graph (/modules/{name}/graph) when you need module-level distance. Tracked as an open roadmap item.

Framework-mediated DB access via INCLUDE macros (item 44)

Natural's generic table-access framework hides a CALLNAT inside a copycode member, invoked with a statement-level macro:

INCLUDE YFRAMGC0 'YFRAML01.C-MOD-GET' '"CO-TABLA-CO-ELEMENTO-ALT"' '"YELEMGN0"' 'YELEMKEY' 'YELEMREC'

The CALLNAT to the generic accessor (YELEMGN0) lives in the copycode, not in the including module, so before item 44 both callees and db-accesses were empty for such modules. The parser now recognises the framework macro and emits a CALLS edge to the accessor named in the macro arguments (de-quoted; e.g. '"YELEMGN0"' → YELEMGN0), tagged with edgeKind = INCLUDE_MACRO on callers/callees. Because the edge is a normal CALLS, the accessor's own table access surfaces transitively: GET /modules/{name}/db-accesses?depth=N reports the table with via = the accessor module. The recognised macros and which argument names the accessor are described declaratively in FrameworkMacros (ac-parser-natural). Scope: the targeted recogniser only — general .nsc copycode expansion is still open.

db-accesses?depth=N is a superset of db-accesses (item 93). Besides READS/WRITES it also returns the mode: "DECLARES" rows — a Java entity's own MAPS_TO table and a repository's repositoryEntity table (item 32) — for every module in the closure, with via naming the declaring module. Before item 93 the transitive query carried only the READS/WRITES branch, so asking the same module with depth dropped its declared table and a Java caller's transitive db-accesses came back empty although the entity it persists through maps to a real table.

Natural view aliases are resolved to the underlying table (item 95). A Natural DML statement names a view variable (1 VDB2-VERSIS_LITERALES VIEW OF VERSVW_LITERALES), not the DDM. db-accesses reports the table — FIND VDB2-VERSIS_LITERALES, FIND NUMBER NEXT-VIEW and STORE VDB2-VERSIS_LITERALES in YLITEMN0 all come back as VERSVW_LITERALES, matching the SQL SELECT … FROM rows in the same module. Before item 95 the alias itself was the reported name, which (a) split one table across several names, (b) made the generator's boilerplate alias NEXT-VIEW a single node shared by 11 modules meaning 11 different tables, and (c) hid every VERSVW_LOGFILE write behind 11 VDB2-*-VLOG aliases. Table names are upper-cased (Natural is case-insensitive).

…including aliases declared in a USING data area (item 98). A view is often declared not in the module but in a LOCAL USING area, in the data-area export form (V 1VDB2-VERSIS_GENAGREE VERSVW_GENAGREE … — no VIEW OF text). Those resolve too: YGEAGBNH's FIND (1) VDB2-VERSIS_GENAGREE reports VERSVW_GENAGREE. Resolution is scoped to each module's own USING set, never by name — alias names are boilerplate, and NEXT-VIEW alone is declared over 100 different tables in upms. A module whose USING areas give two different tables for one alias is left unresolved rather than guessed.

Natural UPDATE(<label>.) / DELETE(<label>.) count as writes (item 96). These act on the current record of the labelled FIND/READ loop, and are reported as WRITES on that loop's table. This is what makes the Y****MN0 access layer's update/delete path visible: YLITEMN0 reports WRITES VERSVW_LITERALES at the STORE and at UPDATE(HOLD-PRIME.) / DELETE(HOLD-PRIME.), where before item 96 it reported only the STORE — reading, wrongly, as an insert-only layer. A reference that resolves to no labelled loop (an unknown label, or the numeric source-line form) records nothing rather than guessing a table.

call-tree/graph agree with callees about overridden dynamic calls (item 97). A manual dynamic-call override hides the placeholder marker rather than deleting it. All read paths now filter it, so a pinned CALLNAT <var> shows the real target and never the variable name. Everything driven by the call-tree BFS — graph, db-accesses?depth=N, sql-statements?depth=N — inherits this.

callers on a dynamically-called module is an over-approximation, and says so. A Natural web-service module is reached by CALLNAT #WIF, resolved by naming pattern: W-LST-N0.nat:362 alone resolves to 29 W****B*S/W****X*S targets, so WGEAGB0S lists W-LST-N0 and W-MNT-N0 as callers. The rows are tagged edgeKind: "CALLNAT_DYNAMIC" — treat those as may-call, not does-call, and check dynamic-calls/overrides / dynamic-calls/unresolved when the distinction matters.

XML payload / interface schema (item 45)

Natural XML wrapper subprograms build a wire payload by mapping data-area fields to XML tags via the ADD-XML-LINE idiom (#W-TAG := '<tag>' / #W-VALUE := <field> / PERFORM ADD-XML-LINE, where the subroutine COMPRESSes '<' #W-TAG '>' #W-VALUE). The deep parser extracts that contract as PAYLOAD_FIELD nodes and exposes it:

  • GET /modules/{name}/payload (ac payload <module>) → an array of {tag, field, direction, lineNo, sourceFile} triples. direction is REQUEST for an emitted (outbound) field. field is the unqualified payload field name (WXMLIN.P-COD-USUARIO → P-COD-USUARIO). sourceFile is the file lineNo refers to — the module's own file for source=IDIOM, or the interface PDA's file for source=PDA (so a caller opens the right file at the line, not the module at a stray line).

  • source=IDIOM (item 45): extracted from a static ADD-XML-LINE emit sequence with literal tags.

  • source=PDA (item 46b): the module is a generic, runtime-driven serializer (it calls the YFRAMN07 tag-builder or has an ADD-XML-LINE/ADD-XML-ACT subroutine) with no static tag list in its source — real production wrappers like WNAUTD0S are this shape. The contract is then derived from the module's PARAMETER USING interface PDA: each field is a payload field, the wire tag is the field name with the framework's EXAMINE … '#' REPLACE '_' normalisation applied (#→_), direction REQUEST. Idiom fields take precedence when both exist.

The static idiom also handles derived tags (EXAMINE #W-TAG FOR '#' REPLACE '_' → #→_) and both directions: an ADD-XML-LINE-style emit sub is REQUEST; a GET-XML-LINE-style parse sub (with the reverse field := #W-VALUE binding) is RESPONSE.

Empty for modules that neither use the idiom nor are a flagged XML wrapper (or are only coarse-ingested).

Copycode (.cpy) expansion (item 46a)

Natural INCLUDE <member> <args> is a compile-time macro: the copycode body is spliced into the including module (with positional &1&… substitution), so a copycode's CALLNAT/PERFORM, DB access and dataflow live in the copycode, not the module. The deep and coarse parsers now expand statement-level copycode includes before parsing, so those constructs surface on the including module — e.g. a READ or CALLNAT that only exists in a .cpy shows up in the host's db-accesses/callees.

  • Copycode-origin nodes/edges report the real .cpy file + line (so navigation lands in the copycode), and carry viaCopycode=<member> + includedAt=<host line>; host statements keep their own file + line (line numbers are remapped after the splice, never shifted).
  • A line number alone is not a location (item 66). Because of the above, one module's calls and field accesses come from more than one file, and host and copycode lines are freely mixed — so always read the line together with the file the endpoint gives it:
    • callers/callees/functions/{f}/callers return sites: [{lineNo, callSiteFileIndex, viaCopycode, includedAt, includePath}] (not a bare lineNos array). callSiteFileIndex indexes sourceFiles and is the file the call is written in; the entry's own sourceFileIndex is a different thing — the file the named module/function is defined in. viaCopycode/includedAt/includePath are set only for copycode sites.
    • variables/{name}/reads|writes return sourceFile = the file lineNo is in (the .cpy for a copycode access), plus viaCopycode + includedAt. Item 109: a writes row also carries assignedValue — the right-hand side as written: 'WREQUD0S' (literal), *PROGRAM (system variable), #DISPLAY(1) (indexed), #SELECTED-KEY.NUM (qualified reference) — plus assignedSubstrPos/assignedSubstrLen for a SUBSTR(...) target (item 83). It is deliberately not normalized to literals: that would drop most write sites. This is what answers "which values does this module put into field X" without opening the source, the dispatcher question item 108 keeps running into. assignedValue is always null for Java — the Java parser does not capture the right-hand side. Read it as "not captured for this language", not as "nothing is assigned". Natural carries it on ~97% of its write edges. On a reads row it is always null by definition. Before item 66 the copycode's line was paired with the host's file: #W-OPTIONS writes in WGEAGB0S were reported at WGEAGB0S.nat:18/20/22, which is its generated comment banner — the writes are really ISICINDI.cpy:18/20/22. Use includedAt when you want the spot in the host module instead.
    • includedAt is always a line in the module's own file, and includePath shows the whole chain (item 104). Natural INCLUDE nests, often through a positional argument (INCLUDE USIX050C 'YFRAMMC1' → INCLUDE &1& → INCLUDE YFRAMC01), and only the innermost member is named by viaCopycode. includedAt used to be the INCLUDE line in the enclosing .cpy — a file the response never named — so ISI173N0 → YFRAMN04 reported line 27, which is a comment in ISI173N0.nat and in truth line 27 of YFRAMMC1.cpy. Now includedAt is the host-module line (232), and includePath: [{sourceFile, lineNo}, …] lists every hop, host first, innermost last (empty for a direct statement, one entry for a one-level include).
    • db-accesses / workfile-accesses return sites: [{lineNo, sourceFile, viaCopycode, includedAt, includePath}] alongside the (kept, backward-compatible) lineNos array — one entry per statement, each tying its line to the file it truly lives in. sql-statements gains sourceFile + viaCopycode on each statement (its startLine/endLine are lines in sourceFile). Before this, a DB/work-file access written in an INCLUDEd copycode reached the API as a bare copycode-local lineNo with nothing to attribute it to — e.g. the DB2 sequence read SELECT … FROM SYSIBM-SYSDUMMY1 lives in USIX043C.cpy at lines 31/39/45/51/57, but db-accesses for the 9 including modules (YAPRFMN0, YUGRPMN0, …) reported those as bare line numbers that land on the host's own comment/DEFINE DATA lines. The sites file context is the same fix item 66 applied to variables/reads|writes and callees.
  • A copycode's nodes belong to the including module, not to the copycode (item 75-B, 2026-08-22). Until now every module that included a .cpy shared one set of nodes for its body. That is no longer so: a copycode-resident node is keyed per including module (ownerModule), so what an agent sees changes in one visible way — counts go up, and they are now per-module. A READ written in a copycode that 20 modules include is 20 access nodes, one per module, instead of one shared node; the same holds for a DEFINE SUBROUTINE in a .cpy and for its control-flow statements. Read it as "each of these modules really does perform this access", which is what the API always claimed but could not previously represent. search/identifier for a name defined in a widely-included copycode therefore returns one hit per including module — filter/group by sourceFile + the module you care about rather than expecting a single row. Modules and DB tables are deliberately not per-module: a module declared inside a copycode (ZDTSTBP6 in ZDTSTBC6.cpy) and every DB_TABLE stay shared, so module lookups are unchanged. Item 75-C (same day) takes this one step further: identity is per expansion site, not per module. A copycode included several times by the same module (JX0031N0.nat includes YFRAMBC0 16 times) now yields one set of nodes per include site, keyed by includePath. So counts rise again for those modules, and — the point of the change — a copycode that opens a block it does not close no longer collects every site's nesting into one node. includedAt alone does not identify a site (item 104 makes it the host's INCLUDE line at every nesting level, and VPARTC02.cpy includes L4NLOGIC 136 times behind a single host line); use includePath when you need to tell two expansions apart. Measured on upms after the recreate: copycode-resident nodes 27,551 -> 51,895 (project total +5.0%), spread over 19,565 distinct owners. Endpoint latency on the heaviest module (JX0030N0.nat, 91 include sites) is unaffected: digest 0.83 s, context 0.25 s, graph 0.19 s.
  • The same line number can legitimately appear twice (item 69). A host statement on line 10 and a copycode statement on line 10 are two different statements, and both are returned — as separate entries differing only in their file. Until item 69 the graph could not hold both: an edge was identified by (source, target, type, lineNo) with no file, so the second one overwrote the first and a real access was missing from every answer. Treat (file, lineNo) as the identity of a site, never lineNo.
  • call-tree's depth counts module hops (item 67). depth is how many module boundaries the shortest call path crosses, not raw CALLS edges — a CALLNAT made from two subroutines deep is still one hop. The root module's own subroutines are therefore depth 0. Measured on upms: WGEAGB0S's seven direct dependencies used to report depth 1..3 (BGEAGFN0 was 3); all seven now report 1.
    • Results at a given depth are larger than before. A subroutine of a module within depth hops is now inside the bound, because it crosses no further boundary. Previously call-tree?depth=1 could hide a DEFINE SUBROUTINE of the very module you asked about, just because it was PERFORMed from another subroutine (raw depth 2) — that is the same bug seen from the inside.
    • call-tree also returns truncated. true means the intra-module subroutine walk stopped at its raw-hop budget, so some FUNCTION items may be missing — not that your depth was exceeded (that is a normal, complete answer). It is conservative and can be true for a complete result. Tune via agenticcode.call-tree.internal-budget (default 20; the deepest internal chain observed in upms is 9).
    • Since item 94 the budget cannot hide a module. MODULE rows come from the same module-hop BFS that db-accesses/sql-statements use, so a callee one hop away is always listed even when its call site sits behind a long internal PERFORM chain (before item 94 it was dropped, and call-tree then contradicted db-accesses). This also removed the path enumeration that made call-tree?followWiring=true time out on Java projects at depth ≥ 2; followWiring is now usable at full depth.
  • field-flow's depth counts module hops (item 68). Like db-accesses/sql-statements (item 65), variables/{name}/field-flow?depth=N now means "up to N module calls apart", not N raw CALLS edges. Before item 68 a consumer called from inside a subroutine sat several raw hops away and was dropped at depth=1, so the endpoint answered "nothing downstream consumes this field" — read that answer with suspicion on any graph ingested before this change.
  • field-flow no longer fabricates flows between same-named fields (item 77). A bare field reference is resolved against the referencing module's own USING includes. It used to be resolved project-wide: an unresolved bare field is one shared node per (name, project), and the resolver aggregated over all owning modules at once, so (a) two modules including different data areas that both declare the name left both unresolved on the shared node, and (b) a module with no matching include was redirected onto another module's field. Either way the two modules ended up on one node, and field-flow — which pairs a producer with a consumer only when both touch the same node — reported a dataflow between modules that share nothing but a field name. In upms: 38 + 28 placeholders affected (199 module-field pairs). Read any pre-item-77 field-flow result for a common field name with suspicion, and note the answer only changes after a deep re-ingest, since resolution runs there. reads/writes are unaffected — they match every node with the name and report only the accessing side, so they never distinguished the targets in the first place. Residue (item 76): a bare field shared via copycode (132 of 18,539 source nodes in upms) is still one node for several modules; per-module identity is a schema change, not a query fix.
  • Copycode provenance survives field resolution (item 70). viaCopycode/includedAt are now kept for fields addressed by qualified name (MYLDA.Q-FIELD, i.e. a field of a LOCAL USING data area) as well as bare ones. Before item 70 only bare references kept it; qualified ones silently came back with viaCopycode: null and the host file, i.e. they looked exactly like host statements.
  • Excluded from expansion: framework macros (handled by the item-44 targeted recogniser), data-area USING includes, unknown members, and any copycode that declares DEFINE DATA. Recursion is cycle-guarded. Copycodes (.cpy) are not standalone modules — they enter the graph only through the including module.
  • Staleness caveat: the item-41/43 hash check hashes the host file, so auto-invalidation triggers on a change to the host — but a change to an included .cpy alone (host unchanged) is not detected; re-ingest the host (refresh/{host}) to pick it up.

Global Data Areas (.gda) (item 46c)

.gda files are now ingested as DATA_STRUCTUREs like .lda/.pda, and DEFINE DATA GLOBAL USING <gda> resolves to them (the INCLUDE/USING recogniser now accepts GLOBAL, not just PARAMETER/LOCAL).

Deep-ingest: now automatic (lazy Tier-2)

Field-level endpoints (flow-forward, flow-backward, field-flow) and cross-module dynamic CALLNAT resolution need a per-module deep ingest, not just a whole-root refresh. This deep ingest is now triggered automatically on demand: calling a field-level endpoint for a module that is only CALL_GRAPH-ingested runs a scoped deep ingest of that module (and its dependency tree) transparently, then returns the resolved result — no 409, no manual POST /refresh/{name} step. The first such call to a cold module is therefore slower (it walks the root and parses the program tree); subsequent calls hit the already-FULL graph.

The deep ingest is best-effort: if the module cannot be resolved to a source file, the endpoint still falls back to the 409 NOT_DEEPLY_INGESTED / NOT_INGESTED hint with a nextAction rather than a misleading empty result.

Flow path-ingest (auto, cross-module fixpoint). flow-forward, flow-backward, and field-flow go one step further than the single start-module deep ingest: after deep-ingesting the start module they run an ingest-and-re-traverse fixpoint. Each round deep-ingests the frontier — the modules the trace surfaced together with their direct callee modules — in one scope, then re-traverses. This is what lets a dataflow trace cross into a dynamically-dispatched callee (CALLNAT PGM-VAR): that callee is not a static dependency of the start module, so it is only pulled in and linked (caller.arg → callee.param) once a round resolves the dynamic CALLS edge and ingests the target. The loop is bounded by agenticcode.deep-ingest.flow-rounds (default 3) and the per-round agenticcode.deep-ingest.fanout-nodes budget, and stops early (fixpoint) as soon as a round pulls in nothing new — so on an already-deep graph a flow query costs one traversal plus one cheap frontier check, no re-run.

Fan-out warm (auto, on the result set). The fan-out / traversal queries callers, /search/identifier, and call-tree also auto-deep-ingest — but on the set of modules their result surfaced, not a single named module. Each runs against the graph as-is, deep-ingests the surfaced modules (blocking, bounded by the fan-out node budget agenticcode.deep-ingest.fanout-nodes, default 50), and — only if that warm actually deepened something — re-runs so the response reflects newly-resolved dynamic dispatch (e.g. a call-tree grows to include a dynamically-dispatched callee once the surfaced program is deep). When everything is already FULL (or the warm resolves nothing) the query returns its first result with no redundant re-run. Note callers warms the already-surfaced callers, so it improves downstream precision but cannot reveal a caller that was invisible at the coarse (call-graph) tier. Other module-level endpoints (context, callees, db-accesses) work regardless of ingest depth and do not trigger a deep ingest.

Bounded fan-out. A by-name deep ingest walks the transitive dependency tree breadth-first, bounded by maxDepth (hops from the named module, default 5, ceiling 20) and maxNodes (files, default 300). When a bound is hit the walk stops early and the ingest response carries a truncation object ({reason: DEPTH|NODES|NODES_AND_DEPTH, maxDepth, maxNodes, hint}) — the modules actually reached are marked FULL, the remainder stays as it was. Raise the limits on an explicit module refresh to pull in more: POST /refresh/{name}?maxDepth=&maxNodes=, or CLI ac refresh <name> --max-depth --max-nodes. Auto-triggered ingests use the server defaults; if a field-level query returns partial data because the target's deep ingest truncated, re-run the explicit refresh with higher limits. (The auto-trigger does not yet accept per-query limit overrides.)

Durable ingest status + coalescing (item 36). Each real MODULE node carries a durable ingestStatus lifecycle — NOT_INGESTED (only its call graph is in the graph) → INGESTING (a deep ingest is in flight) → INGESTED (deeply ingested, ingestDepth = FULL) — separate from ingestDepth. When two calls trigger the same module's deep ingest at once they coalesce rather than both ingesting: within one process an in-process lock serialises them; across processes/restarts a best-effort DB claim marks the module INGESTING and a loser waits for the winner to reach FULL (re-claiming if the claim is released or goes stale after agenticcode.deep-ingest.ingesting-ttl-seconds, default 1800; wait bounded by claim-wait-seconds, default 120). A crash mid-ingest leaves the module re-triggerable (it never reached FULL), and the stale INGESTING is reclaimed on the next call. GET /nodes/{id} / /nodes/{id}/source expose ingestStatus/ingestStatusAt on the module node.

Warm concurrency cap (item 37). All auto deep-ingest/warm work (by-name, fan-out, and flow-frontier) shares a global permit pool (agenticcode.deep-ingest.max-concurrent-warms, default 2), so a burst of queries cannot spawn unbounded parallel parses/Neo4j writes. A permit is acquired only around the actual ingest; if none frees up within agenticcode.deep-ingest.warm-acquire-timeout-seconds (default 10) the warm is skipped and the query returns its Tier-1 (coarse) answer immediately rather than blocking — so under sustained load a query may transiently return shallower data; retry once load subsides, or force it with an explicit POST /refresh/{name}.

OpenAPI contract & CORS (items 48/50)

The server now ships an OpenAPI 3 spec (quarkus-smallrye-openapi): the machine contract the web-UI TypeScript client is generated against. All REST endpoints carry @APIResponse/@Schema annotations, so response bodies are typed in the spec even though the JAX-RS methods return raw Response. Access it at:

  • GET /q/openapi — YAML (or Accept: application/json for JSON)
  • GET /q/swagger-ui — interactive UI (dev)

CORS is enabled (quarkus.http.cors.enabled=true) and restricted to the UI's dev origins (http://localhost:5173, http://localhost:4173) — extend the quarkus.http.cors.origins list per deployment; never ship a wildcard.

Endpoint quick reference

Endpoint Use for
GET /modules?sourceFile=&moduleKind=&extends= List/filter modules; map a source file to its module name(s). Each row carries loc/sloc (item 46) and ingestStatus/ingestDepth (item 50) for status badges without a per-module round trip
GET /loc?language=&sourceFile= Per-language LoC/SLoC rollup (typescript/css too, item 192) (fileCount/loc/sloc) + project total; each file counted once (item 46). For a generated/user_exit project also userExitLoc/userExitSloc + generatedExclusiveLoc/generatedExclusiveSloc (item 47)
GET /modules/{name}/digest Tiny triage view before deciding which modules to expand
GET /modules/{name}/context One-shot overview: functions, callers, callees, DB accesses, SQL/variable summaries (?include= for full lists)
GET /modules/{name}/callers | /callees Direct callers/callees incl. EXTENDS/IMPLEMENTS/INJECTS/REFERENCES. callers scope: external (default) = modules that call this one (CALLNAT/inheritance), rolled up to the calling MODULE: a call made from inside a subroutine/method is attributed to its owning module (never the calling FUNCTION node), and repeated call sites from one caller collapse to a single row whose sites list every line — symmetric with how callees anchors its source side. internal = the module's own subroutines' PERFORM wiring (function-level). The default is external-only, module-typed only, and never lists the module as its own caller (no MODULE→MODULE self-loop); use scope=internal or /functions/{fn}/callers for intra-module / function-level wiring. callees is unchanged (default lists both external CALLNAT and internal PERFORM targets)
GET /modules/{name}/functions/{function}/callers FUNCTION-level callers (item 52): who PERFORMs (Natural) or calls (Java, TypeScript — same-module and cross-module, item 197) a specific subroutine/method, with call-site lineNos. Cross-module callers come from the module-to-module CALLS edge's callerFn/calleeMethod, matched by name (overloads over-approximate; a call from top-level code with no enclosing function shows only in the module-level /callers). Finer-grained than the module-level /callers (which is module→module). Same CallRefResponse shape. CLI ac function-callers <module> <function>
GET /modules/{name}/call-tree?depth= Transitive call graph to scope a feature
GET /modules/{name}/reaches?target=A,B,C&direction=up|down&depth= Item 110 — "can A reach B, and how?" Returns {reachable, paths, truncated} with one witness route per reached target (module names, source→target). direction=down (default): paths from this module to each target. up: paths from each target to this module. The counterpart to call-tree, which only walks downward and returns a closure without routes — one audit hand-rolled this as ~100 /callers requests. reachable: false means "no path over known edges", not "no path": the traversal runs on resolved module calls, so a route through an unresolved dynamic CALLNAT (item 82) is invisible. Bounded by depth (item 75: the call graph has cycles). CLI ac reaches <module> --target A,B --direction up
GET /duplicates Item 114 — identities skipped at ingest because they exist in more than one file ({name, kind, paths}, paths relative to the project root). These are not in /modules; asking for one by name gives 409 DUPLICATE_IDENTITY. Their own calls are absent from the graph, so caller lists elsewhere can be short. CLI ac duplicates
GET /dynamic-calls/unresolved | /overrides · POST/DELETE /overrides Manual dynamic-CALLNAT overrides (item 82). unresolved lists open CALLNAT <var> sites {module, originFile, lineNo, variable}; POST /overrides {originFile, lineNo, targets[], variable?, note?} pins a site to real module(s) (applied at once, persisted across refreshes, 400 UNKNOWN_TARGET for a non-module); DELETE /overrides?originFile=&lineNo= resets one site (omit both = all) and restores the placeholder inline; GET /overrides lists them with an obsolete flag. CLI ac dynamic-calls unresolved|overrides|set|reset
GET /modules/{name}/graph?direction=&depth=&limit= Ego graph (item 49): bounded module-level call neighbourhood as nodes + edges (unlike call-tree). direction = out/in/both; limit caps nodes (BFS order) and sets truncated; unresolved targets carry unresolved=true + empty sourceFile. CLI ac ego-graph
GET /counterparts?module=&kind=&unmatched= (ac counterparts) Item 193: this project's web-service calls / generated DTOs / fields with their twin in the counterpart project; unmatched=true = what nothing serves or mirrors yet
GET /store?slice= (ac store) Item 194: the frontend Redux store — one row per slice (reducer key, RTK sliceName, stateType, the top-level state keys with type/optional/read/write counts, reducer count, total access sites)
GET /store/{slice}/accesses?field=&mode=reads|writes&module= (ac store-accesses) Item 194: who reads/writes a slice — reducers (functionKind=reducer, via=reducer) and the components/hooks/thunks selecting from it (via = useAppSelector, a wrapper hook, getState), with the full sub-path and line. Store fields also answer variables/<slice>.<field>/reads|writes
GET /bindings?dto=&field=&mode=reads|writes&module=&partial= (ac bindings) Item 195: which component reads/writes which DTO field through the generated Fields path objects (<SmartInput field={X.broker.ebene}>), with the field's backend counterpart — --dto Broker --field ebene --mode writes = which page edits Java Broker.ebene. data-structures/{dto}/fields carries boundReads/boundWrites
GET /theme?unused= · GET /theme/{token}/usages (ac theme, ac theme-usages) Item 196: the MUI theme's tokens (createTheme leaves + theme constants, value, uses; declared=false = read by the code but declared by no theme) and where one token is read (style block + CSS property, or plain context)
GET /styles?module=&kind=sx|style|styled|css&withLiterals= (ac styles) Item 196: the style inventory — every sx/style/styled block and CSS rule with CSS keys, hard-coded literals and the theme tokens it reads; withLiterals=true = what bypasses the theme
GET /modules/{name}/db-accesses | /sql-statements DB tables + mode, raw statement text (pass ?depth= for Natural). db-accesses/workfile-accesses return every row when no limit is given (item 103) — they used to default to 50, and since the response is a bare array with no total and no truncated flag the cut was invisible: WGEAGB0S?depth=10 returned 50 of 64 rows and hid 7 tables outright. An explicit limit is still honoured exactly. db-accesses items carry sites: [{lineNo, sourceFile, viaCopycode, includedAt}] (+ kept lineNos); sql-statements items carry sourceFile + viaCopycode — so a copycode-sourced access (e.g. SELECT … FROM SYSIBM-SYSDUMMY1 in USIX043C.cpy) reports the .cpy line, not a bare number that reads as a host-file line
GET /modules/{name}/workfile-accesses Natural work files (sequential/flat-file I/O — READ/WRITE WORK FILE n), the work-file analogue of db-accesses (item 84): [{workFile, physicalName, mode: READS|WRITES, recordBuffers, lineNos, sites}], aggregated per work-file number + mode. sites: [{lineNo, sourceFile, viaCopycode, includedAt}] gives each access its file context (copycode-aware), like db-accesses. physicalName comes from a DEFINE WORK FILE n '<name>', else null. Kept separate from db-accesses — a work file is not an ADABAS/SQL table (fixes a former bug where READ WORK FILE created a phantom DB_TABLE 'WORK'). CLI ac workfile-accesses <module>
GET /modules/{name}/data-structures Which copybooks/inline groups a module uses. A USING <member> binds by member (file) name, never by a level-1 record inside the file (item 100) — before that, WGEAGB0S USING W-WIF-A2 reported old/W-WIF-A7.pda (whose level-1 record is a copy-pasted 1W-WIF-A2), and a data area with several level-1 records and none named after the member (VLAYERLA.lda, USIX020L.lda) resolved to nothing at all (sourceFile: null, area: UNKNOWN, fieldCount: 0) although the file was ingested. One row per resolved definition, (name, sourceFile) (item 102) — never one row blending an arbitrary file with another definition's fieldCount
GET /modules/{name}/payload Natural XML wire-payload contract: {tag, field, direction, source, lineNo, sourceFile} — static ADD-XML-LINE idiom (source=IDIOM, item 45) or derived from the wrapper's interface PDA (source=PDA, item 46b). sourceFile is the file lineNo refers to (module for IDIOM, PDA for PDA)
GET /modules/{name}/comments?kind=&limit=&offset= (ac comments) Item 141: the module's comment blocks — {text, kind, sourceFile, startLine, endLine, target, targetType, truncated}, one row per contiguous block, ordered by line. target/targetType name the declaration the block documents: the declaration immediately below it, else the one enclosing it (so a file header banner documents the MODULE, a /* comment on a field's own line documents that field). kind is JAVADOC|LINE|BLOCK (Java) or NATURAL_BANNER|NATURAL_INLINE|SAG (Natural); ?kind= filters to one. SAG is excluded by default — **SAG directives are generator metadata, not human notes, and would otherwise be most of the answer for every generated Natural module. Text is cut at 4 000 chars (truncated:true); read the file for the rest. Natural copycode comments belong to the copycode's own module, not to each includer. Deep-gated: comments come from the full parse, not the Tier-1 coarse scan, so the module is deep-ingested on demand and a still-shallow module answers 409 NOT_DEEPLY_INGESTED rather than a misleading []
GET /modules/{name}/dispatch-table Natural DECIDE ON VALUE OF routing table
GET /modules/{name}/functions?kind= | /functions/{fn}/overrides | /functions/overrides Method list, modifier filter (Java), subclass overrides (single/bulk). Each item carries sourceFile + viaCopycode (item 84): a Natural subroutine pulled in via INCLUDE reports the copycode file and viaCopycode:true, so its startLine/endLine are read as offsets into that copycode — not into the including module's own file (which is shorter). viaCopycode:false = declared inline. Always false for Java
GET /data-structures/{name}/fields | /db-tables/{name}/columns | /modules/{name}/columns Field/column schemas for DTO/entity generation. Every field carries sourceFile (item 101). When a structure name resolves to several definitions (42 level-1 names recur across upms data areas), the member root — the definition whose file basename equals the name, i.e. what a USING <member> binds to — wins; ?sourceFile= pins a specific one. Before item 101 the definitions were silently unioned: W-WIF-A2 returned 15 fields, the merge of W-WIF-A2.pda (5) and W-WIF-A7.pda (10), a layout that exists nowhere
GET /variables/{name}/reads | /writes | /flow-forward | /flow-backward | /field-flow Impact analysis and dataflow tracing
GET /search/identifier | /search/value | /search/annotation Cross-project lookup by name / literal value / annotation. All three see code only by default — an empty result is not evidence that the string is absent. search/value takes includeComments=true (CLI --include-comments, item 141) to search comment blocks as well; those hits come back as kind: "COMMENT", so a comment is never read as code. It is opt-in because a comment hit is different evidence from a literal, and folding it in silently would move every existing completeness count (item 131). search/identifier and search/annotation never match comments at all — use includeComments, /modules/{name}/comments or /search/source before concluding "not present" (see "Comments: reachable, but never by default" above). search/identifier matches the exact declared name but is sigil-insensitive: a leading Natural sigil (# user, & AIV, + GDA) is ignored on both sides, so name=K-OUT-MAX finds the declared #K-OUT-MAX (and vice-versa). Item 125: a Java type declaration is matched by its short name as well as by the fully-qualified identity the graph stores (item 117) — name=PartnerUpdateLogic finds com.example.PartnerUpdateLogic; before this it answered [], which reads as "no such name". Every match carries simpleName and moduleKind (CLASS/INTERFACE/ENUM/RECORD, PROGRAM/SUBPROGRAM for Natural), both null for non-MODULE hits — so "is this name a type or a method?" needs no second call. contains=true (CLI --contains) switches to a case-insensitive substring match, as on /search/value; it was previously accepted and silently dropped. It matches the FQN too, so a package fragment also hits — filter with type=MODULE/moduleKind if that is noise. contains without a name is 400 MISSING_NAME (a substring search for nothing is a full node dump). The substring scan is unindexed: it is bounded to offset+limit rows, so keep a limit on large projects. Optional scope filters sourceFile=<relpath> and module=<name> (item 53) narrow the match to one file / one module — use them to pinpoint a module-local declaration when a name recurs across dozens of modules (the result is otherwise paginated and the local one may fall off the page). To keep the full cross-project list yet still guarantee a given module's own declaration is on the first page, pass priorityModule=<name> instead of module=: it does not filter, but pins that module's matches to the front (ahead of the otherwise sourceFile-ordered rest) so they survive the limit. This is what the web UI's click-to-identify sends for the open module. CLI ac search-identifier --module --priority-module --source-file --type --contains accept the same filters. Latency (item 105): a lookup whose hits lie in a Natural data area used to take 60-75 s — every fan-out query deep-ingested the surfaced .lda/.pda, which can never reach FULL (a data area yields no MODULE node), so it was re-warmed on every call and each warm dragged a whole-project finalize behind it. Data areas are now excluded from the fan-out warm; they have no deep tier to gain
GET /search/source?regex=&limit=&ignoreCase= (ac search-source) Regex grep over module source text (item 54): {module, sourceFile, lineNo, line} hits + truncated. Case-insensitive by default. Complements /search/identifier (declared names) — use for code patterns (statements, table names, literals). Sees everything in the file, comments included, and needs no ingest depth — so it is the fallback when a module is not deeply ingested, or when the text is something the parsers do not model. For comments specifically, prefer the graph routes added by item 141 (/modules/{name}/comments, search/value?includeComments=true), which also tell you which declaration a comment belongs to
GET /nodes/{id} Every property of one node (when a curated DTO is missing something)
GET /nodes/{id}/source | /modules/{name}/source | /source?file= Source text — only when you have no other access to the source (you always do in this repo, see "Reading source in this repo" above). /modules/{name}/source returns the whole file when the line range is omitted (M1), or a [startLine,endLine] slice when both are given. /source?file=<relpath> (CLI ac file-source) serves a file by relative path rather than module name — for files that aren't standalone modules, e.g. a Natural data area (PDA/LDA) USING'd by a module, whose field line numbers refer to that file. Same whole-file/range + stale-source semantics; the client-supplied path is rejected (400 INVALID_SOURCE_FILE) if it escapes the project root

Full endpoint list, request params, and response field details: x-docs/agent-api-system-prompt.md.

Errors are structured JSON — always (item 136)

Every failure now answers { "error": ..., "code": ..., "details": {} }, including the ones nobody planned for: an unhandled exception is mapped to 500 INTERNAL_ERROR with an errorId in details that matches the stack trace in the server log (the trace itself is never in the response). Before this, an unexpected fault escaped as a plain-text Quarkus error page with no code to branch on — which is exactly the moment a client most needs a machine-readable answer. Deliberate statuses (PROJECT_NOT_FOUND, MISSING_NAME, STALE_SOURCE, the runtime's own routing 404s) pass through unchanged.

The fault that exposed this: search/identifier coerced startLine/endLine unconditionally, and item 114's duplicate markers were the one kind of node created without them, so any page long enough to reach a marker (row 487 on ac) died. Both halves are fixed — the markers now carry lines, and the row mapper no longer trusts that they will.

Verifying the API after a deploy (item 134)

After ./manage-ac.sh deploy (or ./rebuild-and-refresh.sh), run:

./x-scripts/verify-api.sh              # defaults to project 'ac'
./x-scripts/verify-api.sh -p upms      # any ingested project
AC_SERVER_URL=http://host:8787 ./x-scripts/verify-api.sh

It answers one question in ~10 s: does the server that is running right now still return plausible data over the real graph? Exit 0 = all green, 1 = at least one check failed. Every line is PASS, FAIL or SKIP; SKIP means the endpoint family does not apply to that project (a pure Natural project has no rest-endpoints, a leaf module has no callees).

What it covers: /api/version and the project list; the item-130 scope headers (X-AC-Exclude-Dirs, X-AC-Ingested-At, X-AC-Ingest-Incomplete — a true there means a refresh was aborted and every later answer is drawn from a half-updated graph); the item-131 paging contract on the search endpoints; per-family data plausibility; and the structured-error negative cases.

What it is not: a substitute for mvn test. The integration tests pin semantics; this pins "the deployed thing is not obviously broken". A green run is not a quality gate. All assertions are invariants, never fixed counts — counts move with every refresh.

TypeScript / React projects (item 192)

A project may declare language: typescript (ac project create purfe --root … --language typescript). Its .ts/.tsx files (not .d.ts) and plain .css files are ingested; node_modules and dist are excluded by default. Only a typescript project ingests TypeScript — a Java project with a bundled web UI (ac has ac-ui/) never parses it. Java and Natural files stay language-agnostic.

Identities are paths. A TypeScript MODULE is named by its root-relative path without the script extension — pur-r-vstamm/src/store/slices/agstammSlice — with simpleName = the file stem (agstammSlice), workspace = the first path segment, moduleKind = ts/tsx/css, and generated=true + generator (typescript-generator for the Java-side EndpointGenerator output under generated/, hey-api for @hey-api/openapi-ts). A CSS module keeps its extension (pur-ui/src/index.css). Every /modules/{name}/… endpoint accepts either form (item 117), so ac context agstammSlice works — until two workspaces have a file with the same stem, then use the path.

Two tiers, like Java/Natural. Project creation and refresh without --deep run the Tier-1 regex outline in Java: module shell (sourceHash, loc/sloc), one FUNCTION per top-level function / arrow / class (kind = function | component | hook | thunk | styled | class, exported), one DATA_STRUCTURE per interface/type/enum (dataType says which), and a REFERENCES edge per import to the module it resolves to (value = the import clause, specifier = as written). npm packages are not placeholders; they are listed on the module as externalImports. A deep pass (refresh --deep, ingest by name) runs the Node sidecar (ac-parser-typescript/sidecar/extract.mjs, TypeScript compiler API, one whole-program run per npm workspace, 3–5 s and ~0.5 GB each on the pur frontend) and replaces the outline with the checker's view: exact positions, imports resolved against the file system, and calls:

  • a callee owned by a top-level declaration of the same file → FUNCTION -CALLS-> FUNCTION;
  • a callee in another module → MODULE -CALLS-> MODULE carrying callKind (METHOD_CALL, or CONSTRUCTOR for new), callerFn (the calling function, absent at module level), calleeMethod (the owning top-level declaration in the target), callSyntax (call/new/tagged/jsx — a JSX element <HistorieDrawer/> is a call), and on a member call receiver (type of the innermost object, e.g. AgstammControllerEndpoint) and member (saveBroker.post);
  • calls into npm packages and the language library are not edges.

These are the exact properties the Java parser writes, so callers/callees/call-tree, functions/{fn}/callers (item 52) and the placeholder rewiring work unchanged. A module that got its Tier-2 pass carries ingestTier=2.

Sidecar failure is visible, not silent. If node, the script or its node_modules/typescript are missing, or a workspace run fails or times out, the refresh still completes at Tier-1 for those files and the response lists a failure with the pseudo-path sidecar or sidecar:<workspace>. Config: agenticcode.typescript.node, .sidecar-script, .max-heap-mb (1024), .timeout-seconds (600); the image carries node and the sidecar (Dockerfile.jvm), dev mode expects npm ci run once in ac-parser-typescript/sidecar/. The project root is read-only in the container; the sidecar reads the project's own node_modules for library typings and writes nothing.

Scope of the pur frontend project. The registered project covers the pur-ui and pur-ui-common workspaces only; pur-r-vstamm and pur-r-vbuch are excluded via excludeDirs, and the sidecar does not load an excluded workspace. Both workspaces call the pur backend through the legacy generated client (generated/endpoints.ts, backend pur); the hey-api client shape is recognised too but is not in scope.

Resolution notes (verified on purfe, 2026-09-22). An import of a workspace consumed through its package.json exports resolves into its build output (pur-ui-common/dist/x.d.ts); the sidecar maps that to the source twin (pur-ui-common/src/x.ts) so the edge lands on a real module. A bare specifier the checker resolves to neither a file nor a package (immer, redux — transitive dependencies the project does not list) is recorded as an external import, not a placeholder.

Transitive packages are known from root/node_modules (directory names), so Tier-1 treats them as external too.

Stale parsed edges are reaped on a deep refresh (item 198). Every edge the parser emits is stamped with the run's ingestGen at merge time; after a deep re-parse of a file, the edges from that file's nodes whose stamp is older than the run's (the fresh parse did not re-emit them) are deleted before the node sweep, for every language and edge type. An import or call the new parse names differently (a renamed class, a dist→src mapping fix) therefore no longer keeps its old placeholder alive next to the fresh edge, and a placeholder left edgeless falls to the usual placeholder sweep. Tier-1 (changedOnly or non-deep) refreshes do not reap, because a Tier-1 pass emits fewer edges than a deep one. Edges persisted before the stamp was introduced carry no generation and are never reaped; the first deep refresh after upgrading stamps them, the next one reaps — so a project never needs recreating after a parser change any more, two deep refreshes do.

A module node never lives in a copycode (item 201). A program whose body is a single INCLUDE used to get a second MODULE node named after it with the .cpy as sourceFile (ZDTSTBP6 in upms), so modules and search/identifier listed the name twice. Fixed in the parser; a graph ingested before the fix keeps the stray node until the project is recreated, or you remove it by hand: MATCH (m:MODULE {project: $p}) WHERE m.sourceFile ENDS WITH '.cpy' AND m.ownerModule = '' DETACH DELETE m (Natural copycodes are never modules of their own, so the match is exact).

Known limits of 192 (the later items fill them): no field bindings (195), no styles (196); the store is item 194 below. A changedOnly refresh re-runs the sidecar over the whole workspace but re-persists only the changed files, so an unchanged file's facts can lag one refresh (same class of caveat as 46a). The by-name deep ingest resolves dependencies by file stem, so a dependency whose stem exists in several workspaces (index) is reported as a duplicate and skipped — use refresh --deep for the whole frontend.

rest-endpoints lists the frontend's calls. Every member of a generated Endpoint class (AgstammControllerEndpoint.saveBroker) and every hey-api sdk function is a FUNCTION of kind=endpoint carrying the same restPath/httpMethod the Java parser writes for a handler, plus outbound=true — so GET /projects/purfe/rest-endpoints answers with outbound: true rows (handler = AgstammControllerEndpoint.saveBroker, path = /agstamm/ui). Who calls it: modules/{generated module}/callers names the calling slices/components (module level), and the member call api.saveBroker.post(...) is retargeted from the generic PostMethod.post signature to the endpoint function, so the module-to-module CALLS edge carries calleeMethod = AgstammControllerEndpoint.saveBroker and callerFn = <thunk>. Since item 203 the synthetic class-hierarchy edges (resolvedVia: INHERITANCE, caller → each implementation of the called interface/base) exist once per originating call site with its real lineNo, originFile, calleeMethod and callerFn — before, one edge per pair took whichever call line the merge met first. So callees lists every real line for an implementation, and functions/{impl-method}/callers also names callers that go through the interface (RepoImpl.save ← Service.store via Repo.save). Since item 197 functions/{fn}/callers joins these module-to-module edges back to the calling function, so purfe/modules/pur-ui/src/generated/endpoints/functions/GeneralAgreementUiControllerEndpoint.createNew/callers names the thunk in generalAgreementSlice, and on the Java side pur/modules/…AgstammLogic/functions/handleMerge/callers names AgstammController.mergeBroker (a REST controller method itself has no Java callers — it is the HTTP entry point). Extra properties on the node ( GET /nodes/{id}): restUrl (as composed, with placeholders and query string), restBase (an application base such as /pur-r-vbuch/v1 split off so paths compare with the backend's base-less @Path), backend (pur, pur-r-vstamm, dynamic — from the URL builder), queryParams, requestType, responseType, paramsType, generator, owner, member. A generated interface's properties are FIELDs under its DATA_STRUCTURE, named <Interface>.<member> (Broker.ebene; props field = the bare member, owner = the interface — since item 195: a node's identity is type + name + file, and one generated file declares hundreds of interfaces, so a bare vid used to be a single node under six interfaces), so GET /data-structures/AgstammUseCase/fields answers for the frontend too (bare member names; since item 195 — before, the query filtered FIELD out and returned [] for an interface), and counterparts?kind=field rows are named Broker.ebene.

COUNTERPART_OF: the same thing in another project. A project setting counterparts: ["pur"] (ac project create purfe … --counterpart pur, ac project update purfe --counterpart pur, GET /projects/purfe shows it) makes enrichment link, after every refresh of either side:

this project → counterpart matched on
outbound endpoint FUNCTION backend handler FUNCTION httpMethod + path shape (every {param} segment compares as {}, class + method @Path composed as rest-endpoints does); several matches → the handler whose source lives under the frontend's backend name
DATA_STRUCTURE in a generated=true module Java MODULE with the same simpleName unique name only; an ambiguous name stays unlinked
its FIELDs the class's FIELDs name

The edges are rebuilt from scratch on each run (never accumulated) and re-run for the frontend when the backend refreshes, because a refreshed handler node is deleted together with the edges pointing at it. Roadmap item 143 (Natural ↔ Java counterparts) will use the same edge.

GET /projects/{p}/counterparts?module=&kind=rest|dto|field&unmatched=&countOnly=&limit=&offset= (CLI ac counterparts [--module] [--kind] [--unmatched] [--count-only]) lists {kind, name, module, sourceFile, startLine, httpMethod, path, counterpartProject, counterpartName, counterpartModule, counterpartSourceFile, counterpartStartLine}; the counterpart* fields are null for an unlinked row, and unmatched=true is the planning question: which calls does nothing serve, which generated DTOs / fields have no backend twin. Paged with X-AC-Total-Count / X-AC-Truncated like the search endpoints; unknown kind → 400 KIND_UNSUPPORTED; a project naming itself as counterpart → 400 COUNTERPART_SELF.

The Redux store (item 194)

What is modelled. Every createSlice / createAppSlice in a typescript project is a STORE_SLICE node named by the reducer key it is mounted under in the project's configureStore (state.<key>; gruppenprovision for generalAgreementSlice, whose RTK name is generalAgreement) — the key is what every selector path starts with, so it is the identity; the RTK name is kept as sliceName. The sidecar traces each reducer: { key: xReducer } entry through xReducer = xSlice.reducer (also export default xSlice.reducer) back to the slice, across workspaces; a slice no store mounts is named by its own sliceName. Under the slice, one store FIELD per top-level state key, named <key>.<field> (schluesseltabelle.sucheStatus, from the checker's type of initialState, so keys only the state type declares are present too) with dataType = the TS type, optional, store=true, slice, field. Every reducer is a FUNCTION of kind=reducer in the slice's module, named by the action type it handles: schluesseltabelle/updateX for a case reducer (reducerKind=reducer), schluesseltabelle/suche/fulfilled for builder.addCase(sucheByServer.fulfilled, …) (reducerKind=case, trigger = the expression), schluesseltabelle/matcher:isSlicePending(sliceName) for addMatcher (reducerKind=matcher). A thunk whose lifecycle action a case handles CALLS that case (callSyntax=extraReducer), when both live in the same file.

Reads and writes. Inside a reducer every state.a.b chain is a READS/WRITES edge from the reducer FUNCTION to the store FIELD of its first key (state alone → the STORE_SLICE), carrying path (the full sub-path as written, keyTableUseCaseSvcResult.result.tableId), lineNo, via=reducer. A write is an assignment target (compound assignments also read), ++/--, delete, a mutating method on the chain (push, splice, set, delete, …), Object.assign(state.x, …), or a return { … } (each key written) / return other (the whole slice). Outside reducers every store read is a READS edge from the reading function (component, hook, thunk; the module when at top level) to a placeholder <key>.<field> that the finalize step resolve-store-placeholder redirects onto the real field by name and then drops — so a read in KeyTablePage.tsx lands on the field declared in keytableSlice.ts without either file knowing the other. Recognised read forms: useSelector/useAppSelector((state) => state.a.b) (every chain rooted at the arrow's parameter; const { x, y } = useAppSelector((s) => s.a) reads a.x and a.y), wrapper hooks such as useSchluesseltabelleSelector((useCase) => useCase?.result?.purMode) — the sidecar finds the wrapper's inner useAppSelector((state) => selector(state.a.b)), so the read is a.b.result.purMode with via=useSchluesseltabelleSelector (the wrapper's own inner read is recorded too) — and store.getState().a.b / const s = thunkAPI.getState(); s.a.b chains (via=getState; 81 sites on pur-ui). A read of a key the store does not declare stays a placeholder and is listed by search/identifier with an empty sourceFile.

Dispatch. A call to a slice action creator (dispatch(updateX(…)), a binding of xSlice.actions) carries actionType = <sliceName>/updateX and is a cross-module CALLS edge to the slice module with calleeMethod = <sliceName>/updateX — the reducer FUNCTION — so module callees of a component list the slices it dispatches into; a thunk call keeps the thunk FUNCTION as target and carries the thunk's type prefix as actionType. (Function-level exposure of these cross-module edges is item 197.)

Endpoints. GET /projects/{p}/store?slice= (CLI ac store [--slice]) → [{slice, sliceName, module, sourceFile, startLine, endLine, stateType, fields: [{name, type, optional, reads, writes}], reducers, reads, writes}], all slices or one by reducer key / RTK name. GET /projects/{p}/store/{slice}/accesses?field=&mode=reads|writes&module=&countOnly=&limit=&offset= (CLI ac store-accesses <slice> [--field] [--mode] [--module] [--count-only]) → [{mode, slice, field, path, function, functionType, functionKind, module, sourceFile, lineNo, via}], one row per access site, field null for a whole-slice access; paged like counterparts; unknown mode → 400 MODE_UNSUPPORTED, unknown slice → 404 SLICE_NOT_FOUND. Because store fields are FIELD nodes with a unique name, the generic GET /variables/<key>.<field>/reads|writes and search/identifier?name=<key>.<field> answer too.

Limits. Reads through getState() aliases are followed only inside the file that created the alias; a thunk→case CALLS edge is emitted only when thunk and slice share a file (the pur slices do); a case whose trigger cannot be folded to an action type (a predicate matcher) is named by its expression text; the sidecar reads the store of every workspace in scope — a key mounted only by a workspace outside the project (excludeDirs) falls back to the slice's own name (matches on pur). No USES_TYPE from a store field to the DTO it holds yet — the field's dataType says SvcResult<KeyTableUseCase>, the link is item 195. Only a top-level createSlice is a slice: a slice built inside a factory function (pur-ui-common's filetransferSlice, created per instance and not mounted in the pur-ui store) is not modelled.

Verified on purfe (2026-09-22, server 318, recreated + refresh --deep, 285 files, 0 failures, 0 placeholders left): 9 slices — error, global, healthTables, metadata (pur-ui-common) and gruppenprovisionSuche, gruppenprovision (RTK name generalAgreement), schluesseltabelle, multilinguism, translationdata (pur-ui) — 43 reducer functions, 216 read and 82 write sites. schluesseltabelle alone: 62 sites, 25 through useSchluesseltabelleSelector, 19 through getState(), 5 through useAppSelector, 13 in reducers.

DTO field bindings (item 195)

What a binding is. The generator emits, next to every DTO interface, a Fields class tree (AgstammUseCaseField = new AgstammUseCaseFields<AgstammUseCase, never>(), members broker = new BrokerFields<TRoot, Broker>(this, "broker"), list members as keyTableList = (index?) => new KeyTableDOFields(...)). A path expression on it — AgstammUseCaseField.broker.ebene on a <SmartInput field={…}>, <SmartOutput field={…}>, a table's fieldTermForRowData, a column's field:, or a Fields-typed prop such as useCaseFieldPrefix — is a binding. The sidecar types every hop through the checker (XFields<TRoot, TSelf>): the root DTO is TRoot, the owner of the leaf is the TSelf of the hop before it, the leaf name is the field. Each binding becomes a READS edge (plus a WRITES edge when the component tag matches Input$|Dropzone$|Editor$) from the binding function (component/hook; the module at top level) to the FIELD of the generated interface that item 193 already creates (Broker → ebene), carrying path (dotted hops from the root, brokerList[] for a list hop), rootDto, kind, partial, component, attribute, via=binding, lineNo.

  • kind=field: the leaf is a scalar (ebene); kind=prefix: a whole sub-object is handed on (useCaseFieldPrefix={X.tab.translationData}, fieldTermForRowData={X.keyTableList()}) — recorded as a read of the container field so nothing is silently dropped.
  • partial=true: the expression is rooted at a prop or local (props.useCaseFieldPrefix.dataName, tabPrefix.x(idx).gausVal), so the leaf and its owner are exact but the prefix of the path is unknown (only the tail is given). A carrier prop's own name is not part of the path.

Cross-file targets are placeholders <Dto>.<field> (binding=true, owner, field, targetModule) resolved by the finalize step resolve-binding-placeholder exactly — module → DATA_STRUCTURE → FIELD — and dropped afterwards; a leaf the interface does not declare stays a placeholder (listed by search/identifier with empty sourceFile). The item-74 stale-edge sweep covers binding targets on a deep re-ingest.

Endpoints. GET /projects/{p}/bindings?dto=&field=&mode=reads|writes&module=&partial=&countOnly=&limit=&offset= (CLI ac bindings [--dto] [--field] [--mode] [--module] [--partial] [--count-only]) → one row per binding site: {mode, dto, field, path, rootDto, kind, partial, component, attribute, function, functionType, functionKind, module, sourceFile, lineNo, counterpartProject, counterpartModule, counterpartField}. The counterpart* columns are the field's COUNTERPART_OF twin (item 193), so "which page edits Java Broker.ebene" is ac bindings --dto Broker --field ebene --mode writes -p purfe and needs no query on the backend project. dto is the declaring interface (Broker), not the root (AgstammUseCase) — filter on rootDto client-side when you need the latter. Paged like counterparts; unknown mode → 400 MODE_UNSUPPORTED. GET /data-structures/{dto}/fields now returns TypeScript interface fields (type=FIELD) and carries boundReads/boundWrites per field (0 for Natural/Java). The generic variables/<Interface>.<field>/reads|writes sees binding edges too. The leaf's owner is the interface that declares the member — datStart bound through GeneralAgreementDO lands on AbstractHistorizedDO.datStart.

Limits. String-form field="…" bindings and the lodash-path bindings of pur-r-vbuch (excluded workspace) are not modelled; a partial binding cannot say which list element or tab; no USES_TYPE from the component to the root DTO (the rootDto edge property answers that). A binding placeholder and a store placeholder share the FIELD type, so a DTO named exactly like a store key would merge their placeholders (Dto.field vs key.field) — not the case on pur.

Verified on purfe (2026-09-22, server 322, recreated + refresh --deep, 285 files, 0 failures): 257 binding sites (196 reads, 61 writes; 226 scalar leaves, 31 prefixes; 84 partial) over 23 declaring DTOs, every one linked to its pur counterpart field; SmartInput 122, SmartOutput 64, tables 32, column definitions 36. counterparts?kind=field: 744 fields, exactly one twin each (six-fold fan-out before the qualified names), 112 unmatched.

Styling: theme tokens and the style inventory (item 196)

The theme. The file with createTheme({...}) (pur-ui-common/src/theme.ts) gets a DATA_STRUCTURE theme (kind=theme) with one FIELD per token: theme.<path> for every leaf of the literal (palette.primary.dark, typography.h1.fontWeight, shape.borderRadius, sizes.*; props token, tokenKind=path, value folded through constants — #0054A2 — and constant when the leaf names one, PRIMARY_DARK) and theme.<NAME> for each exported string/number constant of that file (tokenKind=constant, PRIMARY). MUI components.styleOverrides are recorded as tokens too, not interpreted.

Style blocks. Every sx={…}, style={…} and styled(X)(…) block is a STYLE node under its component FUNCTION (a styled block under the kind=styled function), named <function>.<sx|style|styled>@<line>:<col>, with styleKind, element (the JSX tag or styled base: Box, 'div'), properties (the CSS keys, nested selectors flattened: &:hover.color, & .MuiPaper-root.background), literals (hard-coded colours/lengths: #005CA9, 17px, -2%, calc(100% - 16px) — mt: 2 is theme-relative and no literal), dynamic (a value the sidecar could not classify, or a whole sx={props.sx}), spread. Plain .css files get one STYLE per rule from the Tier-1 scanner (body@7, @font-face@2; styleKind=css, selector, properties, literals).

Token reads. Inside a style block every chain on a theme value is a REFERENCES edge STYLE → theme.<token> with property (the CSS key it feeds) and via=theme; a theme value is anything typed Theme (useTheme(), a ({ theme }) => styled parameter, the theme object imported under any name) or named theme. theme.spacing(2) ends at spacing, theme.palette.grey['200'] is palette.grey.200; a theme constant (color: PRIMARY) references theme.PRIMARY. A token read outside a style block (borderColor={theme.palette.grey['200']}, code) is the same edge from the enclosing FUNCTION with context = the JSX attribute or code. Cross-file targets are placeholders theme.<token> (theme=true) resolved by exact name at finalize when exactly one theme declares the token; a token no theme declares keeps its placeholder on purpose — GET /theme lists it with declared=false (MUI defaults such as palette.grey.200, palette.common.white, the spacing function, or a typo).

Endpoints. GET /projects/{p}/theme?unused= (CLI ac theme [--unused]) → [{token, kind, value, constant, declared, module, lineNo, uses}], declared tokens first. uses counts project references (style blocks and code); MUI's own consumption of a token is invisible, so unused=true means "no project code references it", never "safe to delete" (palette.primary.main colours every Button whether or not a component names it). GET /projects/{p}/theme/{token}/usages (ac theme-usages palette.primary.dark, token = dotted path or constant name) → [{function, functionKind, module, sourceFile, lineNo, styleKind, element, property, context}]; unknown token → 404 TOKEN_NOT_FOUND. GET /projects/{p}/styles?module=&kind=sx|style|styled|css&withLiterals=&countOnly=&limit=&offset= (ac styles [--module] [--kind] [--with-literals]) → [{name, styleKind, element, selector, function, module, sourceFile, lineNo, properties, literals, dynamic, tokens}], paged like counterparts; withLiterals=true is the review question: which blocks hard-code colours and lengths instead of using the theme. Unknown kind → 400 KIND_UNSUPPORTED.

Several themes (item 199). Every createTheme({...}) in a file is read; a token two themes of one file declare is one row with the first theme's value and variants: 2. A token declared by two theme files (light/dark) is one row per file, and a read of it resolves onto both, so each row counts the use and ?unused=true stays honest; theme/{token}/usages and styles[].tokens report such a read once. Two reads of one token on one line of a style block (color: PRIMARY, borderColor: PRIMARY) are one usage whose property is the comma list of the keys they feed. CSS rules: a block-less @import/@charset line is not part of the next selector, braces inside string values do not open blocks, a nested rule head inside an at-rule body is not a declaration, and a selector repeated on one line (minified CSS) gets a :col suffix in its name.

Limits. className strings are not matched to CSS rules; Emotion css templates and MUI styleOverrides are not modelled; a token read from a component-level function (not a style block) that survives a re-parse keeps its resolved edge until the file's nodes are re-created (same class as item 198).

Verified on purfe (2026-09-22, server 326, recreated + refresh --deep, 285 files, 0 failures): 129 declared tokens (111 paths, 18 constants), 93 of them with no project reference; 12 undeclared tokens the code reads (palette.common.white 4, palette.grey.200 4, spacing 3, palette.divider, palette.text.secondary, applyStyles, transitions.create, …); palette.primary.dark is the most-used token (26 reads: 17 style, 3 sx, 1 styled, 5 as a plain prop such as confirmColor). Style inventory: 304 blocks (196 sx, 78 style, 27 styled, 3 CSS rules), 109 with hard-coded literals (100% 29, 1px 21, 12px 14, 17px 14, …), 75 reading theme tokens, 41 dynamic. The only placeholders left in the project are the 12 undeclared theme tokens — by design.

Missing capability?

If the API/CLI genuinely cannot answer a question (not just unreachable — the capability doesn't exist), finish the task via grep/Explore as a fallback, then use AskUserQuestion to flag the gap and ask whether it should become a roadmap item in x-docs/roadmap.md. Don't silently fall back and move on.