211 KiB
Using AgenticCode on This Repo (Dogfooding)
This repo is ingested as project ac at http://localhost:8787. Per
CLAUDE.md's dogfooding rule, prefer the AgenticCode API over grep/Explore for
call graphs, callers/callees, DB access, dataflow, and module overviews when
working on this repo — it's exactly the tool this project builds.
For full API semantics (params, response shapes, error codes, language
applicability, curl examples), see
x-docs/agent-api-system-prompt.md — that file is the canonical reference and
is not duplicated here. This file only covers what's specific to using the
API as Claude Code, on this checkout.
Tool priority
REST first (GET /api/projects/ac/...) → ac
CLI (ac callers, ac callees, ac call-tree, ac context, ac db-accesses, ...) → grep/Explore. Only fall back past REST/CLI when the
question genuinely isn't answerable by this API at all (see "Missing
capability" below). If the server is unreachable, try
./manage-ac.sh deploy before falling back further.
Re-ingest before trusting results
Query results reflect the last ingest, not the current working tree.
Refresh after code changes before trusting query results:
ac refresh or POST /api/projects/ac/refresh (add --deep / ?deep=true for a
full field-level pass). refresh is the single (re-)ingest surface (item 42) — the
eager ingest-all/ingest-module/ingest-call-graph endpoints were removed.
ac refresh <name> deep-ingests one module + its callees/data areas; add
--neighborhood (POST /refresh/{name}?scope=neighborhood) to also pull in the
module's transitive callers (whole call-graph neighbourhood).
Reconciliation on re-ingest (item 58). A refresh now purges stale nodes:
for every re-parsed file it deletes the nodes the fresh parse no longer produces
(renamed/removed fields, moved statements) rather than leaving them to shadow the
new ones — so identifier counts and /search/identifier results stay clean after a
parser change or an edited source file. Applies to every full-parse path (whole-root
refresh with or without --deep, and per-module refresh/{name}); the coarse
Tier-1 scan run at project creation does not reconcile (it only ever runs on an empty
graph).
Auto-invalidation (item 43). The graph self-heals when files change or are deleted on
disk — you rarely need a manual refresh for staleness:
- Changed file: a field-level query on a module whose source changed re-ingests it
automatically (its stored
sourceHashno longer matches), so results reflect the current file. Granularity is per-module: a changed dependency is picked up when that dependency is itself queried by name. - Deleted file: a whole-project
refreshnow also removes nodes for files deleted from disk (reconciled against the filesystem), closing the earlier gap — no project re-create needed. (A per-modulerefresh/{name}does not sweep; it only touches its own tree.) - Controlled by
agenticcode.auto-invalidate.enabled(defaulttrue).
Reading source in this repo
You have direct filesystem access to this checkout — never call
nodes/{id}/source or modules/{name}/source. Use sourceFile/startLine/
endLine from a graph response (context, digest, search/identifier,
nodes/{id}, ...) and read the file directly. This is always cheaper and gives
full surrounding context; the /source endpoints exist only for API-only
agents with no filesystem access.
Stale-source check (item 41). The /source endpoints compare the file on
disk against the content hash (sourceHash) stored at ingest. If the file
changed since the last ingest they return 409 STALE_SOURCE instead of slicing current text against old line
numbers — re-ingest (refresh) the project to update the graph. Line ranges you
read directly off disk are of course always current; this only guards the API's
own slicing. Copycode/INCLUDE slices are raw pre-expansion file text.
Comments: reachable, but never by default (item 141)
Every default query sees code only. search/identifier, search/annotation,
search/references and a plain search/value never match comment text — a Natural * ... banner,
a trailing /* ..., a Java // line or a Javadoc block.
That silence used to be a wrong answer, not a missing one, wherever a convention records
something in a comment. The UPMS→PUR case: a reengineered service carries its Natural origin in a
Javadoc block (ServiceEndpoint: / UPMSFunction: / UpmsObject:), and the documented
Natural→Java lookup searches the program name in pur, where an empty result is read as "not yet
reengineered" — which it answered for services reengineered months earlier.
Three routes now reach comment text. Pick one before concluding "not present":
| Question | Call |
|---|---|
| "What does this module's header/change log say?" | GET /modules/{name}/comments (ac comments <module>) — blocks with the declaration each documents |
| "Where is this query / SQL / JSON text defined?" (Java) | GET /search/value?value=…&contains=true — since item 139 a static final String built from text blocks, literals, same-class constants and + carries its full text as value; a "…".formatted(...) constant carries its template (with %s) and valueKind: template. Anything built from a method call or another class's constant stays unresolved (no value) |
| "Does this string appear anywhere, code or comment?" | GET /search/value?value=…&includeComments=true (ac search-value --include-comments) — comment hits carry kind: "COMMENT" |
| "…and in text the parsers do not model at all, or in a module that is not deeply ingested?" | GET /search/source?regex=… — raw grep over the files on disk |
GET /pur/search/value?value=WPARTX0S&contains=true → [] (code only)
GET /pur/search/value?value=WPARTX0S&contains=true&includeComments=true → the Javadoc origin block
GET /upms/modules/WAGNTX0S/comments → `* #01 … Bug 266`, …
Comments stay opt-in deliberately: a comment hit is not the same evidence as a literal in code, and folding them into the default result set would move every existing completeness count (item 131's lesson). The flip side is the rule to remember — an empty default search says nothing about comments.
Diagnosing a slow refresh (items 153/155)
Two instrumentation layers, both aimed at the same question — where does the time go?
- Always on: every persist batch logs one line with its statement breakdown
(
Persist deep 200 files: total 8740 ms = prepare 108, merge-nodes 1630, ..., commit 20), and every enrichment step logs duration plus created/deleted rows. The residualcommitis deliberate: total minus the labels is the transaction commit, so nothing hides in an unnamed remainder. - Opt-in per run:
POST /refresh?deep=true&profile=true(ac refresh --deep --profile) runs the enrichment steps under CypherPROFILEand logs the five heaviest operators of every step slower than 5 s. Diagnostic only — it answers "is this step matching or writing?", which the step timing alone cannot. Measured overhead onupms: none worth reporting (908 s vs 905 s). - Build-time diagnostic:
agenticcode.ingest.comments.enabled=falsedropsCOMMENTnodes and theirDOCUMENTSedges just before persist. It exists to measure what comments cost (item 179: ~17 % of a deep refresh) and is not a supported operating mode — with it off,/commentsanswers empty. It is a property, not a request parameter, so it needs a rebuild and cannot be set per run.
What that measured, so nobody re-derives it. The figures below are current (2026-09-06, two clean
runs, roadmap item 175); the campaign of items 153-175 took an upms deep refresh from 1 225 s to
301 s, so any older number quoted elsewhere is stale by a factor of four.
| now | |
|---|---|
deep refresh upms, end to end |
301 s (run-to-run spread 3.7 %) |
| persist | ~149 s — merge-edges 47.6 s, commit 34.1 s, merge-nodes 24.0 s |
| finalize | ~137 s — resolve-field-placeholder W/R 43.4 s, link-args-to-params 23.7 s, resolve-bare-included W/R 37.9 s |
| parsing | ~9 s (interleaved with persist) |
The original diagnosis still holds and is why those steps shrank: the five field-resolution steps
were dominated by matching, not writing — resolve-bare-included once spent 273 M database hits to
produce 43 753 rows. See roadmap items 156 and 175, and item 179 for what comment nodes cost.
limit/offset work on some endpoints and are silently ignored on others
Every endpoint answers completely. The difference is whether it lets you ask for less. 19 list
endpoints declare no limit/offset at all, and JAX-RS drops an undeclared query parameter without
a word — so ?limit=2 there is not an error, it simply has no effect:
GET /upms/modules?limit=2 -> 3 587 rows
GET /upms/modules/WAGNTX0S/functions?limit=2 -> 22 rows
GET /upms/modules/WAGNTX0S/data-structures?limit=1 -> 16 rows
Ignore limit/offset (always the full list):
/modules · /modules/{name}/functions · /modules/{name}/functions/overrides ·
/modules/{name}/functions/{function}/overrides · /modules/{name}/data-structures ·
/modules/{name}/columns · /modules/{name}/payload · /modules/{name}/dispatch-table ·
/modules/{name}/sql-statements · /data-structures/{name}/fields · /db-tables/{name}/columns ·
/variables/{name}/reads · /variables/{name}/writes · /variables/{name}/flow-forward ·
/variables/{name}/flow-backward · /variables/{name}/field-flow · /duplicates ·
/dynamic-calls/unresolved · /dynamic-calls/overrides
Honour them (14, verified against the method signatures): the search endpoints
(search/identifier, search/value, search/annotation, search/references, search/source),
rest-endpoints, and per module callers, callees, context, db-accesses, graph, reaches,
workfile-accesses, comments.
call-tree is bounded differently again — by depth, not by row count — which is the right shape
for a tree but means limit does nothing there either.
Which direction the mistake runs matters, so be precise about it: a dropped limit means you get
more than you asked for, never less. It costs tokens, never correctness — the opposite of item
131's failure, where a silent 50-row cap was read as the complete set. Nothing here can under-report.
The practical consequence is budget, not trust: /modules on upms is 3 587 rows in one response.
Narrow with the filters those endpoints do have (?kind=, ?sourceFile=, ?module=,
?extendsName=) rather than with a limit that will be ignored, and prefer /modules/{name}/digest
or /context when you want an overview rather than an enumeration.
(Implementing limit on those 19 was considered and deliberately not done: the endpoints are honest
as they stand, and a limit that ever acquired a default would reintroduce exactly the silent
truncation item 131 removed.)
callers / callees say when they are cut (item 181)
callers, callees and functions/{fn}/callers now send X-AC-Total-Count and X-AC-Truncated
like the search endpoints, and their body carries total and truncated next to sourceFiles /
items (also for fields=name; ac callers / ac callees print the usual "truncated" warning).
The default page is still 50: upms/modules/DPARTFN0/callees answers 50 of 58 with
X-AC-Truncated: true — before, the 8 missing callees (among them the module that writes the
partner) looked like an inconsistency with digest. Read truncated before concluding "X does not
call Y"; ask with limit=1000 or narrow with scope=external.
Truncation is now visible on the search endpoints (item 131)
search/identifier, search/value, search/annotation, search/references and
rest-endpoints send two headers with every answer:
| Header | Meaning |
|---|---|
X-AC-Total-Count |
how many rows match in total, ignoring limit/offset |
X-AC-Truncated |
true when this page leaves some out |
and all five accept ?countOnly=true (CLI --count-only), returning {"count": n} instead of
rows — a completeness question is a counting question, and @Column on pur is 3 630 rows ≈ 250 k
tokens if you ask for them.
This closes item 103's own follow-up. The bodies stay bare arrays (no contract change), for the same
reason as item 130's scope headers. Why it matters: search/annotation?name=Immutable returned 50 of
95 rows with no total, no flag and no Link/X-Total-Count — and a real UPMS→PUR audit read that
page as the whole set, recording that 17 entities had lost @Immutable when zero had. Re-checked
against all 114 rows of Tables_meta.csv: 94 non-writable carry it, 20 writable do not, no
deviations. The finding cost a day and was pure artefact of the cut.
search/references and rest-endpoints joined this late (item 135, 2026-08-20). Item 131 was
written about the annotation search and both were overlooked. search/references was the damaging
one: it capped at the default 50 and said nothing at all, so a rename scoped from that page missed
every site past the fiftieth and looked complete doing it. rest-endpoints defaults to an uncapped
limit and so never lost rows, but it was equally silent about how many there are. Both were found by
x-scripts/verify-api.sh on its first run, not by a test.
Two mechanics worth knowing: the total costs a second query only when the page comes back full
(a short page is provably the end, so the total is arithmetic), and a total that divides evenly by
limit makes the last full page report truncated with the next page empty — one wasted call, never
a wrong answer. ac prints a note to stderr when a response is flagged truncated, so piping the
body into jq stays clean.
REST surface and scope headers (item 130)
GET /api/projects/{p}/rest-endpoints?module=&countOnly=&limit=&offset= (CLI ac rest-endpoints) lists
{httpMethod, path, module, moduleSimpleName, handler, sourceFile, startLine} — the composed
path (class-level @Path + method-level @Path), so "which code runs for POST /partners" is one
call. Previously the two halves had to be joined by hand from two /search/annotation calls, because
annotations are stored by name without their arguments; the parser now persists restPath and
httpMethod, including a @Path written as a constant reference. A method with no HTTP-verb
annotation is not an endpoint and is excluded. A class with no @Path of its own inherits the
nearest one from its extends/implements ancestry, as JAX-RS does. Rows carry outbound: true
when the declaring type is a @RegisterRestClient interface — a call the application makes, not one
it serves; its path is usually empty because the base URI comes from configuration.
Scope and freshness now ride on every project-scoped response as headers:
| Header | Meaning |
|---|---|
X-AC-Exclude-Dirs |
directories the ingest skipped, or (none) |
X-AC-Ingested-At |
when the graph was last walked (item 126) |
X-AC-Ingest-Incomplete |
true while a whole-root pass runs or after one that never finished (item 129); unknown when no ingest was ever recorded |
Read X-AC-Exclude-Dirs before trusting an empty answer: "no callers" means "none outside tests"
in a project excluding test (app) and "none at all" in one that does not (pur, ac) — the
bodies are identical.
They are headers, not body fields, because most endpoints answer with a bare JSON array
(db-accesses, functions, search/identifier, …); adding a field there would mean restructuring
array → object and breaking the web UI's generated client, the CLI printers and any agent that
indexes [0]. The trade-off is that an agent reading only the JSON body will not see them — so if
you consume this API programmatically, read the headers too. The project shell is cached for ~10 s to
keep this off the request's critical path, and the ingest path invalidates that cache explicitly, so
X-AC-Ingest-Incomplete flips as soon as a refresh starts rather than up to 10 s later.
Every reference site of a name (item 128)
GET /api/projects/{p}/search/references?name=&kind=&countOnly=&limit=&offset= (CLI ac references <name>)
returns {sourceFile, lineNo, kind, inModule, target} per mention of a type — not just per call:
kind |
Where it comes from |
|---|---|
CALL |
a call site (CALLS) |
IMPORT |
an import of the type |
TYPE |
a declared field / parameter / return type |
ANNOTATION |
the type used as an annotation |
EXTENDS, IMPLEMENTS |
inheritance |
INJECTS |
CDI wiring |
CLASS_LITERAL |
X.class in argument position |
INCLUDE |
Natural copycode inclusion |
Use it to scope a rename. callers sees calls alone, so a file that only imports the class,
declares a field of it, or names it in an annotation was invisible — and the rename that missed it
looked complete. The name may be the identity (FQN) or the short form; target echoes what it
resolved to. An unknown kind is 400 INVALID_KIND, never an empty list.
Known limits, by design:
- Local-variable types and generic type arguments are not indexed —
List<Target> xrecordsList, notTarget. They multiply edge volume for much less value than the positions above. - Same-package references have no import, so within one package the index rests on declared-type positions alone.
- Imports are only indexed when they look project-internal (they share the first two package
segments with the importing file). Otherwise every
java.util/framework import would mint a placeholder node on every ingest, just for the finalize sweep to delete it again. - Natural has no import or type-position concept. It contributes
CALL,INCLUDEand inheritance kinds only; this is not parity with Java and should not be read as such. - Reference edges are written at parse time, so they only exist for files re-parsed since this
landed — a project needs a
refreshbefore the index is complete. - Mentions use their own
MENTIONSedge type, kept out ofCALLS/REFERENCESdeliberately: the call-graph traversals (callers,callees,call-tree,ego-graph) followREFERENCESas wiring, so folding imports into it made animportsurface as a caller./search/referencesis the only endpoint that readsMENTIONS; the call graph is unchanged.
Refreshing only what changed (item 129)
POST /api/projects/{p}/refresh (CLI ac refresh) has two ways to avoid re-walking a whole root:
| Form | What it does |
|---|---|
?paths=a/B.java,c/D.java (ac refresh --paths a/B.java,c/D.java) |
Re-ingests exactly those relative paths, deep, plus their dependencies. Paths that match no file — or that are not ingestible source files at all, like pom.xml — come back in unresolved; a typo'd path is never silently dropped. Does not run the deleted-file sweep and does not move ingestedAt: both need a whole-root walk. |
?changedOnly=true (ac refresh --changed-only) |
Whole-root walk, but re-parses only files whose content hash differs from the graph's (files with no stored hash count as changed). |
changedOnly is opt-in on purpose. Three whole-walk behaviours are reduced, and one of them would
be outright corruption if it were hidden:
- A changed Natural copycode disables skipping for that entire run (logged). Natural projects
only — a
.cpysitting in a Java project (as a test fixture, say) is not inlined by anything and no longer stands the optimisation down. Copycode text is inlined into the including module at parse time, so a module whose.cpychanged parses differently while its own hash is unchanged — skipping it would leave a stale expansion behind with nothing to indicate it. - Duplicate-identity detection only sees the changed files, so it can confirm duplicates among them but not discover new ones elsewhere. Existing markers are never cleared.
- User-exit LoC annotation (item 47) is re-stamped only on re-parsed files.
Enrichment is project-wide and still runs in full, so this cuts parse+persist time only — not the finalize pass. On small projects the whole deep refresh is already ~30 s, so measure before assuming a win.
ingest.incomplete (on /projects and /projects/{p}) is true while a whole-root pass runs and
stays true if one never finished — a crash, a container stop, an aborted deep refresh. Before
this, an interrupted deep refresh was indistinguishable from a clean graph: the enrichment steps that
already ran are committed, so queries keep answering, just from a half-updated graph. It cannot
self-heal (a killed process clears nothing) and does not distinguish "running right now" from "died an
hour ago" — both mean the same thing to a caller. A completed refresh clears it.
Renaming a project (item 202)
POST /api/projects/{p}/rename with {"newName": "…"} (CLI ac project rename <old> <new>)
moves the project key on every node and override and in other projects' counterparts lists, then
the shell; the old name answers 404 afterwards and nothing needs re-ingesting. Refusals:
400 INVALID_REQUEST (blank or unchanged), 404, 409 PROJECT_EXISTS. The node rewrite is batched;
an interrupted rename is finished by running it again (97 s for the 939k-node upms). Use it to
keep a reference graph next to a fresh ingest (ac project rename upms upms_alt, then create upms
again and compare). A full DELETE of a project now also removes its manual overrides; recreate
keeps them.
Is this project's graph any good? (item 126)
GET /api/projects and GET /api/projects/{p} (CLI ac project list / ac project show <p>)
carry an ingest object describing the last whole-root ingest:
"ingest": { "ingestedAt": "2026-08-18T10:12:44Z", "mode": "full", "filesExamined": 2981,
"filesPersisted": 2977, "filesFailed": 4, "failures": ["a/B.java", "..."],
"failuresTruncated": false, "durationSeconds": 176, "serverVersion": "…" }
Use it before trusting a negative answer: without it, "no such module" and "that part of the project was never ingested" are the same empty response. Three rules the field obeys:
ingest: nullmeans never recorded, not "ingested nothing" — a project last walked before this existed reads as null rather than as a fabricated zero.- Only whole-root passes write it — the create-time Tier-1 scan,
refresh,refresh?deep=true. A by-namerefresh/{name}, a deep ingest or a fan-out warm ingests real files but sees a fraction of the tree, so it deliberately leavesingestedAtalone; otherwise deepening one module would advertise the whole project as freshly walked. ingestedAtis not a freshness guarantee. It says when the walk ran, not that the graph still matches disk — a file edited a minute later is stale while the timestamp still looks recent. For the real check, read a file throughGET /{p}/source?file=…, which answers409 STALE_SOURCEwhen the content no longer matches the ingested hash. (Targeted/incremental refresh is item 129.)
failures is capped at 200 paths while filesFailed stays exact; failuresTruncated says whether
the list was cut, so a short list is never mistaken for the whole story.
Tier-1 coarse scan on project create (item 36)
Creating a project (POST /api/projects/{p}) now runs a Tier-1 coarse reference scan of the
root before returning, so the project is immediately queryable — no separate ingest call. The scan
is lexer-level (Natural) / a coarse projection of the full parse (Java): it records module/function
/data-structure shells, the identifier index (declared fields, class members), and coarse
CALLS/READS/WRITES/INCLUDES references (Natural PERFORM, CALLNAT '...', dynamic CALLNAT PGM-VAR, PARAMETER/LOCAL USING copybooks; Java resolved calls/type refs) — but no deep
bodies (control flow, statement-level dataflow, arg→param). Scanned modules land
CALL_GRAPH/NOT_INGESTED; field-level detail is filled in by the on-demand deep ingest below. Each
module shell carries a sourceHash. Disable with agenticcode.tier1.scan-on-create=false (creates
an empty project shell to scan later). A missing/unscannable root is non-fatal — the project is
created empty and can be re-scanned.
Unresolved references (item 40)
A reference whose target isn't (yet) ingested — a CALLNAT/PERFORM/USING to a module/copybook
absent from the project, or a dynamic CALLNAT PGM-VAR whose literal can't be recovered — is stored
as a deduped placeholder node (blank sourceFile). Enrichment stamps each with an unresolved
boolean: true while genuinely dangling, false once a real definition of that name is ingested.
/search/identifier returns it as unresolved on each IdentifierMatch, and GET /nodes/{id} carries it in the
node's properties — so an agent can tell a dangling/dynamic
reference apart from a resolved one. In callers/callees such targets already appear as entries
with a blank sourceFile.
Querying a module endpoint: four answers, not one (items 107, 115, 114)
Every GET /api/projects/{p}/modules/{name}/… endpoint used to answer 200 with an all-zeros shell
for a name that exists nowhere in the graph, byte-identical to a real-but-empty module's answer. That
is not cosmetic: call-tree?depth=4 returning 200 with 0 modules reads as "analysed, nothing
found" when the truth is "not analysable" — the module's source was never in the checkout. The
graph knows four states and each now gets its own status:
| state | answer | what it means |
|---|---|---|
no MODULE node for the name |
404 MODULE_NOT_FOUND |
unknown name (typo, or not in this project) |
| placeholder — node exists because something calls it, source never parsed | 409 + {status:"NOT_INGESTED", module, detail, nextAction} |
knowable in principle, not analysed yet |
| ambiguous — several real modules share the name (Java) | 409 + {code:"AMBIGUOUS_NAME", details:{candidates:[...]}} |
the name does not identify a module; pick one |
| duplicate identity — the same name in several files, skipped at ingest | 409 + {code:"DUPLICATE_IDENTITY", details:{paths:[...]}} |
exists twice, deliberately not ingested (item 114); see /duplicates |
| real, ingested module | 200 |
the body is the answer — an empty body genuinely means "nothing found" |
Ambiguous names (item 115). A Java simple name is not unique: nested @Nested test classes,
Builder, Config, WorkingStorage. Measured on pur, 163 names covering 385 modules (~8%) were
addressable only ambiguously, and the endpoints used to answer with the union across unrelated
classes — /modules/BrokerHistoryTests/functions returned 627 functions for a class that has 107.
They now refuse and list the candidates. Repeat the request with ?sourceFile=<candidate>:
GET /modules/Shared/functions → 409 AMBIGUOUS_NAME, candidates ["a/Shared.java","b/Shared.java"]
GET /modules/Shared/functions?sourceFile=a/Shared.java → 200
A Java module's name IS its fully-qualified name (item 117) — com.example.OrderService, and
com.example.Outer.Inner for a nested class. That is what makes same-simple-name classes
distinguishable at all. Every module endpoint accepts either form:
GET /modules/com.example.OrderService/digest → 200, always exact
GET /modules/OrderService/digest → 200 when unique, else 409 AMBIGUOUS_NAME
Responses carry simpleName alongside name for display. The ?module= and ?extends= filters and
the ac CLI take either form too; --source-file remains available on every module command.
Which Java types are modules (item 119). Classes, interfaces, enums, records and annotation
types — moduleKind is one of CLASS | INTERFACE | ENUM | RECORD | ANNOTATION (Natural adds
PROGRAM | SUBPROGRAM | …), and ?moduleKind= filters on it. Their content is modelled the way each
kind carries it: a record's components and an annotation type's members are FIELDs (the latter with
defaultValue where declared), an enum's constants are CONSTANTs, and an enum's or record's
implements is a real edge — so a call against an interface fans out to an enum implementing it.
Before 119 these three kinds were not parsed at all: GET /modules/SomeEnum/digest answered 404,
and a record referenced from elsewhere stayed an unresolved placeholder (409 NOT_INGESTED). One
gap remains by design — a record's compact canonical constructor is not a function node, so calls
made in its body are invisible.
Natural is unaffected throughout: its module names are file stems, and colliding identities are
skipped at ingest, so they are unique by construction (upms has zero ambiguous names, pur 163).
One limit worth knowing: only the request's root module is disambiguated. Inside a traversal this
no longer merges anything, because module names are unique per project after item 117 (pur and upms
have zero duplicate names) — a reference the parser could not qualify is returned flagged
unresolved: true rather than attached to an arbitrary candidate.
A name skipped at ingest because it exists in more than one file answers 409 DUPLICATE_IDENTITY
with the conflicting details.paths (item 114); GET /duplicates lists them all. Note what that does
not fix: the skipped file's own calls were never parsed, so caller lists elsewhere can still be
short — they just no longer look complete.
The 409 applies to everything derived from the module's own source: digest, context,
call-tree, callees, db-accesses, workfile-accesses, sql-statements, functions,
functions/overrides, functions/{fn}/overrides, functions/{fn}/callers, data-structures,
dispatch-table, payload, columns.
callers and graph stay 200 for a placeholder — their data comes from the calling modules'
source and is genuine. When you get a 409, fall back to /callers: it is the one honest answer
available for a module whose own source is missing. (digest no longer surfaces those callers, since
its other fields would all be structurally zero.)
/modules/{name}/source is unchanged: it already answered 404 MODULE_NOT_FOUND for both an absent
module and a placeholder, since there is no source to serve either way.
Scope limit — this guards the root module of a request only. A call-tree that traverses into
placeholder targets still reports that subtree as empty without flagging it, so a dispatcher whose
targets are all placeholders still returns 200 with a silently truncated tree. Cross-check the
targets you care about individually (a 409 tells you it is unanalysed) — see item 103.
Data literals are not call targets (item 62). A CALLNAT <bareword> whose target is really a data
value — a browse key reaching the call site through a copycode/macro argument — used to leave a
permanent unresolved placeholder callee. Enrichment now reaps such a placeholder when it has no Natural
sigil (#/&/+), is a real VARIABLE/CONSTANT of the project, matches no real MODULE, and is
only ever reached by inferred (CALLNAT_DYNAMIC/INCLUDE_MACRO) edges. So callees, call-tree, the
ego graph and the unresolved list no longer list data names as modules. Genuine dynamic dispatch
(CALLNAT #PGM-VAR, sigil'd) is still reported as an unresolved target, and a static CALLNAT 'X' is
always trusted even when X collides with a field name.
The ingest summary agrees with the graph (item 64). A refresh/refresh/{name} response's
unresolved list is built during the file walk, independently of the graph — before item 64 it therefore
reported data fields as missing modules (MODULE CO-TABLA, MODULE NAME-DESC-SP) that enrichment had
already reaped, so the summary and the graph disagreed. It now applies the same item-62 gates as the graph
reap, with the "is this a real field?" test asked project-wide (a data literal is often declared outside
the ingested tree). Genuinely missing modules are still reported — including sigil'd dynamic dispatch
(MODULE #GETSHORT-MODUL) and a missing module whose name collides with a field name but is called
statically (the RPC-CNTX class). Treat unresolved as "dependencies that really are absent".
Constant-folded string-assembled targets (item 83). A dispatcher often builds the CALLNAT <var>
name from a base literal plus one or more SUBSTR overlays — e.g. #GETSHORT-MODUL in YGEAGGNH,
assembled by MOVE 'YGEAGKEY' TO #GETSHORT-MODUL then MOVE 'GN0' TO SUBSTR(#GETSHORT-MODUL,6,3) →
YGEAGGN0. The parser records each SUBSTR write as a WRITES carrying substrPos/substrLen
(1-based), and the resolve-dynamic-callnat-fold enrichment step folds the last full-var literal
written before the call site with the intervening overlays (left/substring) and MERGEs a resolved
CALLS edge (callKind=CALLNAT_DYNAMIC, folded=true) to the assembled module when it is a real
MODULE. It runs in every ingest mode (cheap, bounded by dynamic call sites) and over-approximates like
the other dynamic resolvers. This auto-recovers the Y…GNH → Y…GN0 family with no manual override, so
folded sites drop out of dynamic-calls/unresolved and the assembled target appears in
callees/call-tree/graph/ego graph tagged CALLNAT_DYNAMIC.
Pin what the resolvers can't: manual dynamic-CALLNAT overrides (item 82). Some CALLNAT <var>
targets are beyond the auto-resolvers — a name never present as a recoverable literal in the reachable
code (read from a work-file record, supplied by an unlinked caller, …). Such a site stays an unresolved
placeholder. A human or agent resolves it via
POST /api/projects/{p}/dynamic-calls/overrides with the call site's originFile + lineNo (from
GET .../dynamic-calls/unresolved) and the target module name(s) — multiple targets for a genuine
branch. The override is stored as a :DynamicCallOverride node outside the :AstNode graph, so a
refresh never deletes it and an enrichment step (apply-manual-dynamic-callnat, after the auto
dynamic-CALLNAT resolvers, before the placeholder cleanup) re-applies it automatically — MERGEing a
CALLS edge (callKind=CALLNAT_DYNAMIC, resolvedBy='manual') to each target and flagging the
placeholder manualHidden so callees/digest/graph/call-tree show the real target, not the
#var. It only applies while the site is still unresolved: once an auto-resolver catches up, the
override is skipped and listed obsolete — except a constant-fold (item 83), which a manual override
outranks: the fold skips a site carrying a :DynamicCallOverride, and a stale folded edge there is
dropped (delete-folded-overridden-dynamic-callnat) before apply-manual-dynamic-callnat runs, so the
pinned target replaces it. DELETE .../dynamic-calls/overrides?originFile=&lineNo=
resets one site (omit both = all), clearing the manual edges and un-hiding the placeholder inline (no
refresh). A target that is not a real MODULE is rejected 400 UNKNOWN_TARGET. Bug B fix: the
callees items now carry unresolved (mirroring what graph already exposed), so an unresolved
dynamic target is machine-distinguishable from a resolved one without inspecting sourceFile. Since item 200 the apply
and the reset also rebuild the calling
modules' derived CALLS_MODULE edges in the same transaction, so reaches and field-flow honour
a pinned target without a refresh, like callees and call-tree already did.
Dispatch guards: read guards — it is the only complete condition (item 72). A dispatch-table row's
guardField/guardValue/guardValues describe the innermost DECIDE only. Natural nests
value-DECIDEs inside each other (824 of upms's 7,729 blocks, 178 modules), and then the innermost guard
is just one conjunct: in VMULTMN4, the row for YTABLMA0.TX-TABLA reports #FIELD-NAME = 'TX-TABLA',
but the assignment also requires #SHORT-VIEW = 'TABL'. guards is the full chain — [{field, values}],
outermost first, joined by AND, each link's values joined by OR. Reading only the legacy fields
over-generalises: port that to Java and you get a branch firing where Natural never would. For an
unnested DECIDE the chain has one link and says the same as the legacy fields.
- Still incomplete for
NONE/ANYbranches (item 73): an assignment in aNONEbranch is reported under its enclosing chain alone, but its real condition is "enclosing guard AND NOT any siblingVALUE" — a negation a chain of equalities cannot express.guardsis strictly better than the legacy fields, not a total answer.
A dispatch row's lineNo belongs to sourceFile, not to the module (item 122). dispatch-table
rows now carry the same provenance quartet as callees/db-accesses/workfile-accesses/functions:
sourceFile, viaCopycode, includedAt, includePath. Resolve lineNo against sourceFile —
when viaCopycode is non-null the assignment is written in that copycode and lineNo is a line of the
.cpy, while includedAt is the INCLUDE line in the module. Before item 122 the row carried only
lineNo, so following it against the module file landed somewhere arbitrary: 26 of VCOMIN50's 44 rows
reported line 18 or 20, which in that module is a change-history comment; the real sites are
ISICINDE.cpy:18 and ISICINDI.cpy:20. Note the guard chain may span the include boundary — the
outer DECIDE in the module, the inner one in the copycode — so a row can have a multi-link guards
chain whose links live in different files.
Within one guard: prefer guardValues over guardValue (item 64). dispatch-table rows carry both.
guardValue is lossy and kept only for compatibility: it comma-joins the branch's VALUE literals,
which drops blank alternatives entirely and, for a multi-alternative branch, yields a synthetic string the
guarded field never equals ("A1, A2"). guardValues is the faithful list — every alternative in source
order, blanks included — so VALUE 'GENAGREE-WOUT-SP', ' ' reports ["GENAGREE-WOUT-SP", " "], recording
that a blank guard field also routes into that branch. When reasoning about routing (or porting a
DECIDE to Java), read guardValues; guardValue will silently under-report the branch's conditions.
Text that is not code never yields a call (items 61 & 63). The CALLNAT/CALLNAT_DYNAMIC patterns
are unanchored (a CALLNAT may legally appear mid-line), so both ingest tiers first neutralise non-code
text: a full-line * comment is skipped, a trailing /* … is stripped (item 61), and a match whose
keyword falls inside a quoted string literal is rejected (item 63). A real CALLNAT 'MOD' is
unaffected — its keyword sits outside the quotes. This matters for trusting callers/callees/
call-tree: before item 63, prose such as WRITE(#MSG) 'NACH CALLNAT ISINGEAG:' or #ERR-TYPE := 'Callnat USIA008N' fabricated a CALLNAT_DYNAMIC edge to the real module of that name, so a mere log
message appeared as a genuine call — and, because a real module existed, it was not flagged
unresolved and could not be reaped by the item-62 cleanup. If you query a graph ingested before
2026-07-16, re-ingest (ac refresh) before trusting call-graph edges into modules that are also
mentioned in log/error text.
Calls made through a copycode's arguments (items 120/121/123). In Natural the target of a
CALLNAT is often not written at the call site at all: a copycode receives the module name as a
positional INCLUDE argument and issues CALLNAT &2&. Three defects in that argument path — arguments
continued on the next line, the doubled-quote escape '''X''', and the double-quote delimiter
'"X"' — meant such a call produced no edge and no unresolved-dynamic-call entry, so callees,
callers, call-tree and reaches agreed on an answer that was simply absent, with nothing saying
"not analysed". This hit the browse/access layer hardest, because that is where the idiom lives:
YCARPBN1 and YPOLIBN1 reported 0 callers each. Fixed 2026-08-07; re-ingest recovered 2672
copycode-derived call pairs (+49%) in upms with none lost. A graph ingested before 2026-08-07
under-reports Natural callers/callees, and does so silently — re-ingest before concluding a Natural
module is unused. A copycode parameter that genuinely has no argument now surfaces in
/dynamic-calls/unresolved as &n& rather than being dropped, so "not analysable" is visible.
A call edge no longer outlives the call it was parsed from (item 124). Until 2026-08-09 a refresh
only added the corrected call and left the old one in place, because an edge is reaped only when one
of its endpoints is — and a parser fix changes neither (the calling subroutine is unchanged, the old
target is a never-swept placeholder). So callees/callers could report a call that no source line
makes, flagged unresolved: true and indistinguishable from a genuine unresolved dynamic call. A
re-parsed Natural file's call edges are now reaped before the fresh ones are merged, and a call-target
placeholder left with no callers is deleted. Two consequences for a graph ingested before
2026-08-09: an unresolved: true callee may be an artefact of an already-fixed parser bug rather
than a real dynamic call, and search/identifier may list module names that exist nowhere in the
source. Both clear on the next deep refresh. Note the reap deliberately spares CALLNAT_DYNAMIC edges
onto real modules — those are the dynamic-call resolvers' output, not the parser's.
LoC / SLoC metrics (item 46)
Every file-level node (a MODULE program/class, or a DATA_STRUCTURE for a Natural .lda/.pda
data area) is stamped at ingest with two deterministic line metrics:
loc— physical lines of the file (language-independent; a trailing newline adds no phantom line).sloc— source lines of code: non-blank, non-comment lines, computed per language. Natural drops full-line*/**//*and inline/*comments; Java drops//and/* … */blocks while keeping those tokens when they appear inside string literals.
Both the cheap Tier-1 coarse scan and the deep Tier-2 parse use the same per-language counter, so a
module's loc/sloc are identical at any ingest depth — you can sum them to get exact,
reproducible project totals.
Where to read them:
GET /modules—loc/slocon each row.GET /modules/{name}/context—loc/slocon the module.GET /nodes/{id}—loc/slocin the node's raw properties.GET /loc(ac loc) — the rollup: a per-language breakdown (fileCount,loc,sloc) plus a project-wide total, optionally narrowed by?language=/?sourceFile=. Each source file is counted once even when it yields several nodes (Java inner classes, Natural inline groups).
null metrics mean the node predates item 46 — re-ingest (refresh) to backfill.
Generated vs. user-exit split (item 47)
A project can be created with a source language (required at creation; an attribute only — ingest
still classifies files by extension) and a generatedDir/userExitDir pair (directory names,
matched as path components like excludeDirs; both or neither). Generated modules already contain their
hand-written user-exit twin inline, so at ingest a module under generatedDir whose name also occurs
under userExitDir is annotated with that twin's LoC/SLoC (userExitLoc/userExitSloc). User-exit
files are not ingested as standalone modules (they would collide by name) — the walk skips
userExitDir.
Consequence for all non-LoC analysis: the generatedDir copy is the canonical, sole module for
every structural query (call graph, DB access, functions, data structures, identifiers, dataflow,
dispatch table). userExitDir exists only to compute the generated-vs-manually-written LoC split
below; it never contributes nodes/edges. So when verifying an API response against source for a Natural
module, always read the generatedDir file (e.g. generated_src/subprogram/WGEAGB0S.nat), not the
user_exit fragment.
GET /loc (ac loc) then reports, per language row and in the project total:
loc/sloc— the total (generated, which already includes the user exits).userExitLoc/userExitSloc— the sum of the annotated user-exit twins (the hand-written part).generatedExclusiveLoc/generatedExclusiveSloc— total − user-exit, clamped ≥0 per file (the purely generated part).
All three are 0 for projects without a generated/user-exit split. Create with
ac project create <name> <root> -l natural -g generated_src -u user_exit, or add the split to an
existing project via ac project update <name> -g generated_src -u user_exit.
Java DB_ACCESS in a project without JPA entities (item 140, 2026-08-27)
A DB_TABLE node is only ever created from a JPA @Entity or a Panache active-record class. In a
Java project that has none, no DB_ACCESS candidate can resolve — and the parser's candidate
heuristic is a deliberate over-approximation: its read gate admits any static receiver whose method
starts with get/find/read/list/… , so UserContext.getCurrent() and
TextUtils.getColumn(line, 0, 8) become candidates. On app that produced 2219 DB_ACCESS
nodes in a codebase with no database access whatsoever.
db-accesses never showed them (it joins the table with a plain MATCH), but sql-statements
did: it joins with OPTIONAL MATCH, so unresolved candidates came back as rows with
"table": null — 76 of them on a single app module.
Since 2026-08-27 the enrichment step reap-java-db-access-without-tables deletes every Java
DB_ACCESS of a project that holds no DB_TABLE. For such a project sql-statements is now empty
instead of noisy. Three things to know:
- The gate is project-level, not per node. One entity anywhere in the project switches the reaper
off, and the unresolved candidates stay.
pur(2014 of 3951 unresolved) andac(221 of 335) are unaffected, and still returntable: nullrows. Treat asql-statementsrow whosetableisnullas unverified, in any project that has tables. - Natural is untouched. A Natural
DB_ACCESScomes from a literalREAD/FIND/STOREand is a real access whether or not its view resolved (upms: 14302 of 14303 resolve). - Recovering from it needs a full refresh. If such a project later gains its first entity, the
reaped nodes only come back for files that are actually re-parsed —
changedOnlywill not restore them,refreshwithout it (orrecreate) will.
?depth= means module hops (item 65)
On db-accesses / sql-statements (and the ?module= scope of variables/{name}/reads|writes),
depth=N means N module calls away — the same unit /modules/{name}/graph?depth= and call-tree neighbours
use. depth=1 = the modules this one directly CALLNATs, regardless of how deeply the calling
statement sits inside subroutines.
Before item 65 these endpoints bounded the traversal on raw CALLS edges. A CALLS edge starts at the
statement making the call, not at the MODULE node, so the traversal also stepped through internal
PERFORM jumps and depth measured statement nesting, not dependency distance. Concretely:
WGEAGB0S reached YGEAGBNH's tables through two module calls, but the raw path is 5 edges
(WGEAGB0S → GET-DATA → GET-MAIN-DATA → BGEAGFN0 → READ-FILE → YGEAGBNH), so db-accesses?depth=2
returned [] — reading as "no DB access" — and only depth=5 was truthful.
If you scripted a depth workaround (a deliberately large depth to compensate), drop it: depth
is now the value you'd naturally expect, and inflated values just widen the result set.
call-tree'sdepthcolumn is still raw-hop based and mixes internal subroutines into the tree: a direct dependency called from the main body showsdepth=1while one called two subroutines deep showsdepth=3. Use the ego graph (/modules/{name}/graph) when you need module-level distance. Tracked as an open roadmap item.
Framework-mediated DB access via INCLUDE macros (item 44)
Natural's generic table-access framework hides a CALLNAT inside a copycode member, invoked with a
statement-level macro:
INCLUDE YFRAMGC0 'YFRAML01.C-MOD-GET' '"CO-TABLA-CO-ELEMENTO-ALT"' '"YELEMGN0"' 'YELEMKEY' 'YELEMREC'
The CALLNAT to the generic accessor (YELEMGN0) lives in the copycode, not in the including module,
so before item 44 both callees and db-accesses were empty for such modules. The parser now
recognises the framework macro and emits a CALLS edge to the accessor named in the macro arguments
(de-quoted; e.g. '"YELEMGN0"' → YELEMGN0), tagged with edgeKind = INCLUDE_MACRO on
callers/callees. Because the edge is a normal CALLS, the accessor's own table access surfaces
transitively: GET /modules/{name}/db-accesses?depth=N reports the table with via = the accessor
module. The recognised macros and which argument names the accessor are described declaratively in
FrameworkMacros (ac-parser-natural). Scope: the targeted recogniser only — general .nsc copycode
expansion is still open.
db-accesses?depth=N is a superset of db-accesses (item 93). Besides READS/WRITES it also
returns the mode: "DECLARES" rows — a Java entity's own MAPS_TO table and a repository's
repositoryEntity table (item 32) — for every module in the closure, with via naming the declaring
module. Before item 93 the transitive query carried only the READS/WRITES branch, so asking the
same module with depth dropped its declared table and a Java caller's transitive db-accesses came
back empty although the entity it persists through maps to a real table.
Natural view aliases are resolved to the underlying table (item 95). A Natural DML statement names a
view variable (1 VDB2-VERSIS_LITERALES VIEW OF VERSVW_LITERALES), not the DDM. db-accesses reports
the table — FIND VDB2-VERSIS_LITERALES, FIND NUMBER NEXT-VIEW and STORE VDB2-VERSIS_LITERALES
in YLITEMN0 all come back as VERSVW_LITERALES, matching the SQL SELECT … FROM rows in the same
module. Before item 95 the alias itself was the reported name, which (a) split one table across several
names, (b) made the generator's boilerplate alias NEXT-VIEW a single node shared by 11 modules meaning
11 different tables, and (c) hid every VERSVW_LOGFILE write behind 11 VDB2-*-VLOG aliases. Table
names are upper-cased (Natural is case-insensitive).
…including aliases declared in a USING data area (item 98). A view is often declared not in the
module but in a LOCAL USING area, in the data-area export form (V 1VDB2-VERSIS_GENAGREE VERSVW_GENAGREE … — no VIEW OF text). Those resolve too: YGEAGBNH's
FIND (1) VDB2-VERSIS_GENAGREE reports VERSVW_GENAGREE. Resolution is scoped to each module's own
USING set, never by name — alias names are boilerplate, and NEXT-VIEW alone is declared over 100
different tables in upms. A module whose USING areas give two different tables for one alias is left
unresolved rather than guessed.
Natural UPDATE(<label>.) / DELETE(<label>.) count as writes (item 96). These act on the current
record of the labelled FIND/READ loop, and are reported as WRITES on that loop's table. This is
what makes the Y****MN0 access layer's update/delete path visible: YLITEMN0 reports WRITES VERSVW_LITERALES at the STORE and at UPDATE(HOLD-PRIME.) / DELETE(HOLD-PRIME.), where before
item 96 it reported only the STORE — reading, wrongly, as an insert-only layer. A reference that
resolves to no labelled loop (an unknown label, or the numeric source-line form) records nothing rather
than guessing a table.
call-tree/graph agree with callees about overridden dynamic calls (item 97). A manual
dynamic-call override hides the placeholder marker rather than deleting it. All read paths now filter it,
so a pinned CALLNAT <var> shows the real target and never the variable name. Everything driven by the
call-tree BFS — graph, db-accesses?depth=N, sql-statements?depth=N — inherits this.
callers on a dynamically-called module is an over-approximation, and says so. A Natural web-service
module is reached by CALLNAT #WIF, resolved by naming pattern: W-LST-N0.nat:362 alone resolves to 29
W****B*S/W****X*S targets, so WGEAGB0S lists W-LST-N0 and W-MNT-N0 as callers. The rows are
tagged edgeKind: "CALLNAT_DYNAMIC" — treat those as may-call, not does-call, and check
dynamic-calls/overrides / dynamic-calls/unresolved when the distinction matters.
XML payload / interface schema (item 45)
Natural XML wrapper subprograms build a wire payload by mapping data-area fields to XML tags via the
ADD-XML-LINE idiom (#W-TAG := '<tag>' / #W-VALUE := <field> / PERFORM ADD-XML-LINE, where the
subroutine COMPRESSes '<' #W-TAG '>' #W-VALUE). The deep parser extracts that contract as
PAYLOAD_FIELD nodes and exposes it:
-
GET /modules/{name}/payload(ac payload <module>) → an array of{tag, field, direction, lineNo, sourceFile}triples.directionisREQUESTfor an emitted (outbound) field.fieldis the unqualified payload field name (WXMLIN.P-COD-USUARIO→P-COD-USUARIO).sourceFileis the filelineNorefers to — the module's own file forsource=IDIOM, or the interface PDA's file forsource=PDA(so a caller opens the right file at the line, not the module at a stray line). -
source=IDIOM(item 45): extracted from a staticADD-XML-LINEemit sequence with literal tags. -
source=PDA(item 46b): the module is a generic, runtime-driven serializer (it calls theYFRAMN07tag-builder or has anADD-XML-LINE/ADD-XML-ACTsubroutine) with no static tag list in its source — real production wrappers likeWNAUTD0Sare this shape. The contract is then derived from the module'sPARAMETER USINGinterface PDA: each field is a payload field, the wire tag is the field name with the framework'sEXAMINE … '#' REPLACE '_'normalisation applied (#→_), directionREQUEST. Idiom fields take precedence when both exist.
The static idiom also handles derived tags (EXAMINE #W-TAG FOR '#' REPLACE '_' → #→_) and
both directions: an ADD-XML-LINE-style emit sub is REQUEST; a GET-XML-LINE-style parse sub
(with the reverse field := #W-VALUE binding) is RESPONSE.
Empty for modules that neither use the idiom nor are a flagged XML wrapper (or are only coarse-ingested).
Copycode (.cpy) expansion (item 46a)
Natural INCLUDE <member> <args> is a compile-time macro: the copycode body is spliced into the
including module (with positional &1&… substitution), so a copycode's CALLNAT/PERFORM, DB access
and dataflow live in the copycode, not the module. The deep and coarse parsers now expand statement-level
copycode includes before parsing, so those constructs surface on the including module — e.g. a READ
or CALLNAT that only exists in a .cpy shows up in the host's db-accesses/callees.
- Copycode-origin nodes/edges report the real
.cpyfile + line (so navigation lands in the copycode), and carryviaCopycode=<member>+includedAt=<host line>; host statements keep their own file + line (line numbers are remapped after the splice, never shifted). - A line number alone is not a location (item 66). Because of the above, one module's calls and field
accesses come from more than one file, and host and copycode lines are freely mixed — so always read the
line together with the file the endpoint gives it:
callers/callees/functions/{f}/callersreturnsites: [{lineNo, callSiteFileIndex, viaCopycode, includedAt, includePath}](not a barelineNosarray).callSiteFileIndexindexessourceFilesand is the file the call is written in; the entry's ownsourceFileIndexis a different thing — the file the named module/function is defined in.viaCopycode/includedAt/includePathare set only for copycode sites.variables/{name}/reads|writesreturnsourceFile= the filelineNois in (the.cpyfor a copycode access), plusviaCopycode+includedAt. Item 109: awritesrow also carriesassignedValue— the right-hand side as written:'WREQUD0S'(literal),*PROGRAM(system variable),#DISPLAY(1)(indexed),#SELECTED-KEY.NUM(qualified reference) — plusassignedSubstrPos/assignedSubstrLenfor aSUBSTR(...)target (item 83). It is deliberately not normalized to literals: that would drop most write sites. This is what answers "which values does this module put into field X" without opening the source, the dispatcher question item 108 keeps running into.assignedValueis alwaysnullfor Java — the Java parser does not capture the right-hand side. Read it as "not captured for this language", not as "nothing is assigned". Natural carries it on ~97% of its write edges. On areadsrow it is alwaysnullby definition. Before item 66 the copycode's line was paired with the host's file:#W-OPTIONSwrites inWGEAGB0Swere reported atWGEAGB0S.nat:18/20/22, which is its generated comment banner — the writes are reallyISICINDI.cpy:18/20/22. UseincludedAtwhen you want the spot in the host module instead.includedAtis always a line in the module's own file, andincludePathshows the whole chain (item 104). NaturalINCLUDEnests, often through a positional argument (INCLUDE USIX050C 'YFRAMMC1'→INCLUDE &1&→INCLUDE YFRAMC01), and only the innermost member is named byviaCopycode.includedAtused to be theINCLUDEline in the enclosing.cpy— a file the response never named — soISI173N0 → YFRAMN04reported line 27, which is a comment inISI173N0.natand in truth line 27 ofYFRAMMC1.cpy. NowincludedAtis the host-module line (232), andincludePath: [{sourceFile, lineNo}, …]lists every hop, host first, innermost last (empty for a direct statement, one entry for a one-level include).db-accesses/workfile-accessesreturnsites: [{lineNo, sourceFile, viaCopycode, includedAt, includePath}]alongside the (kept, backward-compatible)lineNosarray — one entry per statement, each tying its line to the file it truly lives in.sql-statementsgainssourceFile+viaCopycodeon each statement (itsstartLine/endLineare lines insourceFile). Before this, a DB/work-file access written in anINCLUDEd copycode reached the API as a bare copycode-locallineNowith nothing to attribute it to — e.g. the DB2 sequence readSELECT … FROM SYSIBM-SYSDUMMY1lives inUSIX043C.cpyat lines 31/39/45/51/57, butdb-accessesfor the 9 including modules (YAPRFMN0, YUGRPMN0, …) reported those as bare line numbers that land on the host's own comment/DEFINE DATAlines. Thesitesfile context is the same fix item 66 applied tovariables/reads|writesandcallees.
- A copycode's nodes belong to the including module, not to the copycode (item 75-B, 2026-08-22).
Until now every module that included a
.cpyshared one set of nodes for its body. That is no longer so: a copycode-resident node is keyed per including module (ownerModule), so what an agent sees changes in one visible way — counts go up, and they are now per-module. AREADwritten in a copycode that 20 modules include is 20 access nodes, one per module, instead of one shared node; the same holds for aDEFINE SUBROUTINEin a.cpyand for its control-flow statements. Read it as "each of these modules really does perform this access", which is what the API always claimed but could not previously represent.search/identifierfor a name defined in a widely-included copycode therefore returns one hit per including module — filter/group bysourceFile+ the module you care about rather than expecting a single row. Modules and DB tables are deliberately not per-module: a module declared inside a copycode (ZDTSTBP6inZDTSTBC6.cpy) and everyDB_TABLEstay shared, so module lookups are unchanged. Item 75-C (same day) takes this one step further: identity is per expansion site, not per module. A copycode included several times by the same module (JX0031N0.natincludesYFRAMBC016 times) now yields one set of nodes per include site, keyed byincludePath. So counts rise again for those modules, and — the point of the change — a copycode that opens a block it does not close no longer collects every site's nesting into one node.includedAtalone does not identify a site (item 104 makes it the host's INCLUDE line at every nesting level, andVPARTC02.cpyincludesL4NLOGIC136 times behind a single host line); useincludePathwhen you need to tell two expansions apart. Measured onupmsafter the recreate: copycode-resident nodes 27,551 -> 51,895 (project total +5.0%), spread over 19,565 distinct owners. Endpoint latency on the heaviest module (JX0030N0.nat, 91 include sites) is unaffected:digest0.83 s,context0.25 s,graph0.19 s. - The same line number can legitimately appear twice (item 69). A host statement on line 10 and a
copycode statement on line 10 are two different statements, and both are returned — as separate entries
differing only in their file. Until item 69 the graph could not hold both: an edge was identified by
(source, target, type, lineNo)with no file, so the second one overwrote the first and a real access was missing from every answer. Treat(file, lineNo)as the identity of a site, neverlineNo. call-tree'sdepthcounts module hops (item 67).depthis how many module boundaries the shortest call path crosses, not rawCALLSedges — aCALLNATmade from two subroutines deep is still one hop. The root module's own subroutines are therefore depth 0. Measured onupms:WGEAGB0S's seven direct dependencies used to report depth 1..3 (BGEAGFN0was 3); all seven now report 1.- Results at a given
depthare larger than before. A subroutine of a module withindepthhops is now inside the bound, because it crosses no further boundary. Previouslycall-tree?depth=1could hide aDEFINE SUBROUTINEof the very module you asked about, just because it wasPERFORMed from another subroutine (raw depth 2) — that is the same bug seen from the inside. call-treealso returnstruncated.truemeans the intra-module subroutine walk stopped at its raw-hop budget, so someFUNCTIONitems may be missing — not that yourdepthwas exceeded (that is a normal, complete answer). It is conservative and can betruefor a complete result. Tune viaagenticcode.call-tree.internal-budget(default 20; the deepest internal chain observed inupmsis 9).- Since item 94 the budget cannot hide a module.
MODULErows come from the same module-hop BFS thatdb-accesses/sql-statementsuse, so a callee one hop away is always listed even when its call site sits behind a long internalPERFORMchain (before item 94 it was dropped, andcall-treethen contradicteddb-accesses). This also removed the path enumeration that madecall-tree?followWiring=truetime out on Java projects atdepth ≥ 2;followWiringis now usable at full depth.
- Results at a given
field-flow'sdepthcounts module hops (item 68). Likedb-accesses/sql-statements(item 65),variables/{name}/field-flow?depth=Nnow means "up to N module calls apart", not N rawCALLSedges. Before item 68 a consumer called from inside a subroutine sat several raw hops away and was dropped atdepth=1, so the endpoint answered "nothing downstream consumes this field" — read that answer with suspicion on any graph ingested before this change.field-flowno longer fabricates flows between same-named fields (item 77). A bare field reference is resolved against the referencing module's ownUSINGincludes. It used to be resolved project-wide: an unresolved bare field is one shared node per(name, project), and the resolver aggregated over all owning modules at once, so (a) two modules including different data areas that both declare the name left both unresolved on the shared node, and (b) a module with no matching include was redirected onto another module's field. Either way the two modules ended up on one node, andfield-flow— which pairs a producer with a consumer only when both touch the same node — reported a dataflow between modules that share nothing but a field name. Inupms: 38 + 28 placeholders affected (199 module-field pairs). Read any pre-item-77field-flowresult for a common field name with suspicion, and note the answer only changes after a deep re-ingest, since resolution runs there.reads/writesare unaffected — they match every node with the name and report only the accessing side, so they never distinguished the targets in the first place. Residue (item 76): a bare field shared via copycode (132 of 18,539 source nodes inupms) is still one node for several modules; per-module identity is a schema change, not a query fix.- Copycode provenance survives field resolution (item 70).
viaCopycode/includedAtare now kept for fields addressed by qualified name (MYLDA.Q-FIELD, i.e. a field of aLOCAL USINGdata area) as well as bare ones. Before item 70 only bare references kept it; qualified ones silently came back withviaCopycode: nulland the host file, i.e. they looked exactly like host statements. - Excluded from expansion: framework macros (handled by the item-44 targeted recogniser), data-area
USINGincludes, unknown members, and any copycode that declaresDEFINE DATA. Recursion is cycle-guarded. Copycodes (.cpy) are not standalone modules — they enter the graph only through the including module. - Staleness caveat: the item-41/43 hash check hashes the host file, so auto-invalidation triggers
on a change to the host — but a change to an included
.cpyalone (host unchanged) is not detected; re-ingest the host (refresh/{host}) to pick it up.
Global Data Areas (.gda) (item 46c)
.gda files are now ingested as DATA_STRUCTUREs like .lda/.pda, and DEFINE DATA GLOBAL USING <gda> resolves to them (the INCLUDE/USING recogniser now accepts GLOBAL, not just
PARAMETER/LOCAL).
Deep-ingest: now automatic (lazy Tier-2)
Field-level endpoints (flow-forward, flow-backward, field-flow) and
cross-module dynamic CALLNAT resolution need a per-module deep ingest, not
just a whole-root refresh. This deep ingest is now triggered
automatically on demand: calling a field-level endpoint for a module that is
only CALL_GRAPH-ingested runs a scoped deep ingest of that module (and its
dependency tree) transparently, then returns the resolved result — no 409,
no manual POST /refresh/{name} step. The first such call to a cold module is
therefore slower (it walks the root and parses the program tree); subsequent
calls hit the already-FULL graph.
The deep ingest is best-effort: if the module cannot be resolved to a source
file, the endpoint still falls back to the 409 NOT_DEEPLY_INGESTED /
NOT_INGESTED hint with a nextAction rather than a misleading empty result.
Flow path-ingest (auto, cross-module fixpoint). flow-forward,
flow-backward, and field-flow go one step further than the single start-module
deep ingest: after deep-ingesting the start module they run an ingest-and-re-traverse
fixpoint. Each round deep-ingests the frontier — the modules the trace surfaced
together with their direct callee modules — in one scope, then re-traverses. This is
what lets a dataflow trace cross into a dynamically-dispatched callee (CALLNAT PGM-VAR): that callee is not a static dependency of the start module, so it is only
pulled in and linked (caller.arg → callee.param) once a round resolves the dynamic
CALLS edge and ingests the target. The loop is bounded by
agenticcode.deep-ingest.flow-rounds (default 3) and the per-round
agenticcode.deep-ingest.fanout-nodes budget, and stops early (fixpoint) as soon as a
round pulls in nothing new — so on an already-deep graph a flow query costs one
traversal plus one cheap frontier check, no re-run.
Fan-out warm (auto, on the result set). The fan-out / traversal queries
callers, /search/identifier, and call-tree also auto-deep-ingest — but on
the set of modules their result surfaced, not a single named module. Each
runs against the graph as-is, deep-ingests the surfaced modules (blocking,
bounded by the fan-out node budget agenticcode.deep-ingest.fanout-nodes,
default 50), and — only if that warm actually deepened something — re-runs so
the response reflects newly-resolved dynamic dispatch (e.g. a call-tree grows
to include a dynamically-dispatched callee once the surfaced program is deep).
When everything is already FULL (or the warm resolves nothing) the query
returns its first result with no redundant re-run. Note callers warms the
already-surfaced callers, so it improves downstream precision but cannot
reveal a caller that was invisible at the coarse (call-graph) tier. Other
module-level endpoints (context, callees, db-accesses) work regardless of
ingest depth and do not trigger a deep ingest.
Bounded fan-out. A by-name deep ingest walks the transitive dependency tree
breadth-first, bounded by maxDepth (hops from the named module, default 5,
ceiling 20) and maxNodes (files, default 300). When a bound is hit the walk
stops early and the ingest response carries a truncation object
({reason: DEPTH|NODES|NODES_AND_DEPTH, maxDepth, maxNodes, hint}) — the modules
actually reached are marked FULL, the remainder stays as it was. Raise the
limits on an explicit module refresh to pull in more:
POST /refresh/{name}?maxDepth=&maxNodes=, or CLI ac refresh <name> --max-depth --max-nodes. Auto-triggered ingests
use the
server defaults; if a field-level query returns partial data because the target's
deep ingest truncated, re-run the explicit refresh with higher limits. (The
auto-trigger does not yet accept per-query limit overrides.)
Durable ingest status + coalescing (item 36). Each real MODULE node carries a
durable ingestStatus lifecycle — NOT_INGESTED (only its call graph is in the
graph) → INGESTING (a deep ingest is in flight) → INGESTED (deeply ingested,
ingestDepth = FULL) — separate from ingestDepth. When two calls trigger the same
module's deep ingest at once they coalesce rather than both ingesting: within one
process an in-process lock serialises them; across processes/restarts a best-effort DB
claim marks the module INGESTING and a loser waits for the winner to reach FULL
(re-claiming if the claim is released or goes stale after
agenticcode.deep-ingest.ingesting-ttl-seconds, default 1800; wait bounded by
claim-wait-seconds, default 120). A crash mid-ingest leaves the module re-triggerable
(it never reached FULL), and the stale INGESTING is reclaimed on the next call.
GET /nodes/{id} / /nodes/{id}/source expose ingestStatus/ingestStatusAt on the module node.
Warm concurrency cap (item 37). All auto deep-ingest/warm work (by-name, fan-out,
and flow-frontier) shares a global permit pool
(agenticcode.deep-ingest.max-concurrent-warms, default 2), so a burst of queries
cannot spawn unbounded parallel parses/Neo4j writes. A permit is acquired only around
the actual ingest; if none frees up within
agenticcode.deep-ingest.warm-acquire-timeout-seconds (default 10) the warm is skipped
and the query returns its Tier-1 (coarse) answer immediately rather than blocking —
so under sustained load a query may transiently return shallower data; retry once load
subsides, or force it with an explicit POST /refresh/{name}.
OpenAPI contract & CORS (items 48/50)
The server now ships an OpenAPI 3 spec (quarkus-smallrye-openapi): the machine
contract the web-UI TypeScript client is generated against. All REST endpoints
carry @APIResponse/@Schema annotations, so response bodies are typed in the
spec even though the JAX-RS methods return raw Response. Access it at:
GET /q/openapi— YAML (orAccept: application/jsonfor JSON)GET /q/swagger-ui— interactive UI (dev)
CORS is enabled (quarkus.http.cors.enabled=true) and restricted to the UI's dev
origins (http://localhost:5173, http://localhost:4173) — extend the
quarkus.http.cors.origins list per deployment; never ship a wildcard.
Endpoint quick reference
| Endpoint | Use for |
|---|---|
GET /modules?sourceFile=&moduleKind=&extends= |
List/filter modules; map a source file to its module name(s). Each row carries loc/sloc (item 46) and ingestStatus/ingestDepth (item 50) for status badges without a per-module round trip |
GET /loc?language=&sourceFile= |
Per-language LoC/SLoC rollup (typescript/css too, item 192) (fileCount/loc/sloc) + project total; each file counted once (item 46). For a generated/user_exit project also userExitLoc/userExitSloc + generatedExclusiveLoc/generatedExclusiveSloc (item 47) |
GET /modules/{name}/digest |
Tiny triage view before deciding which modules to expand |
GET /modules/{name}/context |
One-shot overview: functions, callers, callees, DB accesses, SQL/variable summaries (?include= for full lists) |
GET /modules/{name}/callers | /callees |
Direct callers/callees incl. EXTENDS/IMPLEMENTS/INJECTS/REFERENCES. callers scope: external (default) = modules that call this one (CALLNAT/inheritance), rolled up to the calling MODULE: a call made from inside a subroutine/method is attributed to its owning module (never the calling FUNCTION node), and repeated call sites from one caller collapse to a single row whose sites list every line — symmetric with how callees anchors its source side. internal = the module's own subroutines' PERFORM wiring (function-level). The default is external-only, module-typed only, and never lists the module as its own caller (no MODULE→MODULE self-loop); use scope=internal or /functions/{fn}/callers for intra-module / function-level wiring. callees is unchanged (default lists both external CALLNAT and internal PERFORM targets) |
GET /modules/{name}/functions/{function}/callers |
FUNCTION-level callers (item 52): who PERFORMs (Natural) or calls (Java, TypeScript — same-module and cross-module, item 197) a specific subroutine/method, with call-site lineNos. Cross-module callers come from the module-to-module CALLS edge's callerFn/calleeMethod, matched by name (overloads over-approximate; a call from top-level code with no enclosing function shows only in the module-level /callers). Finer-grained than the module-level /callers (which is module→module). Same CallRefResponse shape. CLI ac function-callers <module> <function> |
GET /modules/{name}/call-tree?depth= |
Transitive call graph to scope a feature |
GET /modules/{name}/reaches?target=A,B,C&direction=up|down&depth= |
Item 110 — "can A reach B, and how?" Returns {reachable, paths, truncated} with one witness route per reached target (module names, source→target). direction=down (default): paths from this module to each target. up: paths from each target to this module. The counterpart to call-tree, which only walks downward and returns a closure without routes — one audit hand-rolled this as ~100 /callers requests. reachable: false means "no path over known edges", not "no path": the traversal runs on resolved module calls, so a route through an unresolved dynamic CALLNAT (item 82) is invisible. Bounded by depth (item 75: the call graph has cycles). CLI ac reaches <module> --target A,B --direction up |
GET /duplicates |
Item 114 — identities skipped at ingest because they exist in more than one file ({name, kind, paths}, paths relative to the project root). These are not in /modules; asking for one by name gives 409 DUPLICATE_IDENTITY. Their own calls are absent from the graph, so caller lists elsewhere can be short. CLI ac duplicates |
GET /dynamic-calls/unresolved | /overrides · POST/DELETE /overrides |
Manual dynamic-CALLNAT overrides (item 82). unresolved lists open CALLNAT <var> sites {module, originFile, lineNo, variable}; POST /overrides {originFile, lineNo, targets[], variable?, note?} pins a site to real module(s) (applied at once, persisted across refreshes, 400 UNKNOWN_TARGET for a non-module); DELETE /overrides?originFile=&lineNo= resets one site (omit both = all) and restores the placeholder inline; GET /overrides lists them with an obsolete flag. CLI ac dynamic-calls unresolved|overrides|set|reset |
GET /modules/{name}/graph?direction=&depth=&limit= |
Ego graph (item 49): bounded module-level call neighbourhood as nodes + edges (unlike call-tree). direction = out/in/both; limit caps nodes (BFS order) and sets truncated; unresolved targets carry unresolved=true + empty sourceFile. CLI ac ego-graph |
GET /counterparts?module=&kind=&unmatched= (ac counterparts) |
Item 193: this project's web-service calls / generated DTOs / fields with their twin in the counterpart project; unmatched=true = what nothing serves or mirrors yet |
GET /store?slice= (ac store) |
Item 194: the frontend Redux store — one row per slice (reducer key, RTK sliceName, stateType, the top-level state keys with type/optional/read/write counts, reducer count, total access sites) |
GET /store/{slice}/accesses?field=&mode=reads|writes&module= (ac store-accesses) |
Item 194: who reads/writes a slice — reducers (functionKind=reducer, via=reducer) and the components/hooks/thunks selecting from it (via = useAppSelector, a wrapper hook, getState), with the full sub-path and line. Store fields also answer variables/<slice>.<field>/reads|writes |
GET /bindings?dto=&field=&mode=reads|writes&module=&partial= (ac bindings) |
Item 195: which component reads/writes which DTO field through the generated Fields path objects (<SmartInput field={X.broker.ebene}>), with the field's backend counterpart — --dto Broker --field ebene --mode writes = which page edits Java Broker.ebene. data-structures/{dto}/fields carries boundReads/boundWrites |
GET /theme?unused= · GET /theme/{token}/usages (ac theme, ac theme-usages) |
Item 196: the MUI theme's tokens (createTheme leaves + theme constants, value, uses; declared=false = read by the code but declared by no theme) and where one token is read (style block + CSS property, or plain context) |
GET /styles?module=&kind=sx|style|styled|css&withLiterals= (ac styles) |
Item 196: the style inventory — every sx/style/styled block and CSS rule with CSS keys, hard-coded literals and the theme tokens it reads; withLiterals=true = what bypasses the theme |
GET /modules/{name}/db-accesses | /sql-statements |
DB tables + mode, raw statement text (pass ?depth= for Natural). db-accesses/workfile-accesses return every row when no limit is given (item 103) — they used to default to 50, and since the response is a bare array with no total and no truncated flag the cut was invisible: WGEAGB0S?depth=10 returned 50 of 64 rows and hid 7 tables outright. An explicit limit is still honoured exactly. db-accesses items carry sites: [{lineNo, sourceFile, viaCopycode, includedAt}] (+ kept lineNos); sql-statements items carry sourceFile + viaCopycode — so a copycode-sourced access (e.g. SELECT … FROM SYSIBM-SYSDUMMY1 in USIX043C.cpy) reports the .cpy line, not a bare number that reads as a host-file line |
GET /modules/{name}/workfile-accesses |
Natural work files (sequential/flat-file I/O — READ/WRITE WORK FILE n), the work-file analogue of db-accesses (item 84): [{workFile, physicalName, mode: READS|WRITES, recordBuffers, lineNos, sites}], aggregated per work-file number + mode. sites: [{lineNo, sourceFile, viaCopycode, includedAt}] gives each access its file context (copycode-aware), like db-accesses. physicalName comes from a DEFINE WORK FILE n '<name>', else null. Kept separate from db-accesses — a work file is not an ADABAS/SQL table (fixes a former bug where READ WORK FILE created a phantom DB_TABLE 'WORK'). CLI ac workfile-accesses <module> |
GET /modules/{name}/data-structures |
Which copybooks/inline groups a module uses. A USING <member> binds by member (file) name, never by a level-1 record inside the file (item 100) — before that, WGEAGB0S USING W-WIF-A2 reported old/W-WIF-A7.pda (whose level-1 record is a copy-pasted 1W-WIF-A2), and a data area with several level-1 records and none named after the member (VLAYERLA.lda, USIX020L.lda) resolved to nothing at all (sourceFile: null, area: UNKNOWN, fieldCount: 0) although the file was ingested. One row per resolved definition, (name, sourceFile) (item 102) — never one row blending an arbitrary file with another definition's fieldCount |
GET /modules/{name}/payload |
Natural XML wire-payload contract: {tag, field, direction, source, lineNo, sourceFile} — static ADD-XML-LINE idiom (source=IDIOM, item 45) or derived from the wrapper's interface PDA (source=PDA, item 46b). sourceFile is the file lineNo refers to (module for IDIOM, PDA for PDA) |
GET /modules/{name}/comments?kind=&limit=&offset= (ac comments) |
Item 141: the module's comment blocks — {text, kind, sourceFile, startLine, endLine, target, targetType, truncated}, one row per contiguous block, ordered by line. target/targetType name the declaration the block documents: the declaration immediately below it, else the one enclosing it (so a file header banner documents the MODULE, a /* comment on a field's own line documents that field). kind is JAVADOC|LINE|BLOCK (Java) or NATURAL_BANNER|NATURAL_INLINE|SAG (Natural); ?kind= filters to one. SAG is excluded by default — **SAG directives are generator metadata, not human notes, and would otherwise be most of the answer for every generated Natural module. Text is cut at 4 000 chars (truncated:true); read the file for the rest. Natural copycode comments belong to the copycode's own module, not to each includer. Deep-gated: comments come from the full parse, not the Tier-1 coarse scan, so the module is deep-ingested on demand and a still-shallow module answers 409 NOT_DEEPLY_INGESTED rather than a misleading [] |
GET /modules/{name}/dispatch-table |
Natural DECIDE ON VALUE OF routing table |
GET /modules/{name}/functions?kind= | /functions/{fn}/overrides | /functions/overrides |
Method list, modifier filter (Java), subclass overrides (single/bulk). Each item carries sourceFile + viaCopycode (item 84): a Natural subroutine pulled in via INCLUDE reports the copycode file and viaCopycode:true, so its startLine/endLine are read as offsets into that copycode — not into the including module's own file (which is shorter). viaCopycode:false = declared inline. Always false for Java |
GET /data-structures/{name}/fields | /db-tables/{name}/columns | /modules/{name}/columns |
Field/column schemas for DTO/entity generation. Every field carries sourceFile (item 101). When a structure name resolves to several definitions (42 level-1 names recur across upms data areas), the member root — the definition whose file basename equals the name, i.e. what a USING <member> binds to — wins; ?sourceFile= pins a specific one. Before item 101 the definitions were silently unioned: W-WIF-A2 returned 15 fields, the merge of W-WIF-A2.pda (5) and W-WIF-A7.pda (10), a layout that exists nowhere |
GET /variables/{name}/reads | /writes | /flow-forward | /flow-backward | /field-flow |
Impact analysis and dataflow tracing |
GET /search/identifier | /search/value | /search/annotation |
Cross-project lookup by name / literal value / annotation. All three see code only by default — an empty result is not evidence that the string is absent. search/value takes includeComments=true (CLI --include-comments, item 141) to search comment blocks as well; those hits come back as kind: "COMMENT", so a comment is never read as code. It is opt-in because a comment hit is different evidence from a literal, and folding it in silently would move every existing completeness count (item 131). search/identifier and search/annotation never match comments at all — use includeComments, /modules/{name}/comments or /search/source before concluding "not present" (see "Comments: reachable, but never by default" above). search/identifier matches the exact declared name but is sigil-insensitive: a leading Natural sigil (# user, & AIV, + GDA) is ignored on both sides, so name=K-OUT-MAX finds the declared #K-OUT-MAX (and vice-versa). Item 125: a Java type declaration is matched by its short name as well as by the fully-qualified identity the graph stores (item 117) — name=PartnerUpdateLogic finds com.example.PartnerUpdateLogic; before this it answered [], which reads as "no such name". Every match carries simpleName and moduleKind (CLASS/INTERFACE/ENUM/RECORD, PROGRAM/SUBPROGRAM for Natural), both null for non-MODULE hits — so "is this name a type or a method?" needs no second call. contains=true (CLI --contains) switches to a case-insensitive substring match, as on /search/value; it was previously accepted and silently dropped. It matches the FQN too, so a package fragment also hits — filter with type=MODULE/moduleKind if that is noise. contains without a name is 400 MISSING_NAME (a substring search for nothing is a full node dump). The substring scan is unindexed: it is bounded to offset+limit rows, so keep a limit on large projects. Optional scope filters sourceFile=<relpath> and module=<name> (item 53) narrow the match to one file / one module — use them to pinpoint a module-local declaration when a name recurs across dozens of modules (the result is otherwise paginated and the local one may fall off the page). To keep the full cross-project list yet still guarantee a given module's own declaration is on the first page, pass priorityModule=<name> instead of module=: it does not filter, but pins that module's matches to the front (ahead of the otherwise sourceFile-ordered rest) so they survive the limit. This is what the web UI's click-to-identify sends for the open module. CLI ac search-identifier --module --priority-module --source-file --type --contains accept the same filters. Latency (item 105): a lookup whose hits lie in a Natural data area used to take 60-75 s — every fan-out query deep-ingested the surfaced .lda/.pda, which can never reach FULL (a data area yields no MODULE node), so it was re-warmed on every call and each warm dragged a whole-project finalize behind it. Data areas are now excluded from the fan-out warm; they have no deep tier to gain |
GET /search/source?regex=&limit=&ignoreCase= (ac search-source) |
Regex grep over module source text (item 54): {module, sourceFile, lineNo, line} hits + truncated. Case-insensitive by default. Complements /search/identifier (declared names) — use for code patterns (statements, table names, literals). Sees everything in the file, comments included, and needs no ingest depth — so it is the fallback when a module is not deeply ingested, or when the text is something the parsers do not model. For comments specifically, prefer the graph routes added by item 141 (/modules/{name}/comments, search/value?includeComments=true), which also tell you which declaration a comment belongs to |
GET /nodes/{id} |
Every property of one node (when a curated DTO is missing something) |
GET /nodes/{id}/source | /modules/{name}/source | /source?file= |
Source text — only when you have no other access to the source (you always do in this repo, see "Reading source in this repo" above). /modules/{name}/source returns the whole file when the line range is omitted (M1), or a [startLine,endLine] slice when both are given. /source?file=<relpath> (CLI ac file-source) serves a file by relative path rather than module name — for files that aren't standalone modules, e.g. a Natural data area (PDA/LDA) USING'd by a module, whose field line numbers refer to that file. Same whole-file/range + stale-source semantics; the client-supplied path is rejected (400 INVALID_SOURCE_FILE) if it escapes the project root |
Full endpoint list, request params, and response field details:
x-docs/agent-api-system-prompt.md.
Errors are structured JSON — always (item 136)
Every failure now answers { "error": ..., "code": ..., "details": {} }, including the ones nobody
planned for: an unhandled exception is mapped to 500 INTERNAL_ERROR with an errorId in details
that matches the stack trace in the server log (the trace itself is never in the response). Before
this, an unexpected fault escaped as a plain-text Quarkus error page with no code to branch on —
which is exactly the moment a client most needs a machine-readable answer. Deliberate statuses
(PROJECT_NOT_FOUND, MISSING_NAME, STALE_SOURCE, the runtime's own routing 404s) pass through
unchanged.
The fault that exposed this: search/identifier coerced startLine/endLine unconditionally, and
item 114's duplicate markers were the one kind of node created without them, so any page long enough
to reach a marker (row 487 on ac) died. Both halves are fixed — the markers now carry lines, and the
row mapper no longer trusts that they will.
Verifying the API after a deploy (item 134)
After ./manage-ac.sh deploy (or ./rebuild-and-refresh.sh), run:
./x-scripts/verify-api.sh # defaults to project 'ac'
./x-scripts/verify-api.sh -p upms # any ingested project
AC_SERVER_URL=http://host:8787 ./x-scripts/verify-api.sh
It answers one question in ~10 s: does the server that is running right now still return
plausible data over the real graph? Exit 0 = all green, 1 = at least one check failed. Every line is
PASS, FAIL or SKIP; SKIP means the endpoint family does not apply to that project (a pure
Natural project has no rest-endpoints, a leaf module has no callees).
What it covers: /api/version and the project list; the item-130 scope headers
(X-AC-Exclude-Dirs, X-AC-Ingested-At, X-AC-Ingest-Incomplete — a true there means a refresh
was aborted and every later answer is drawn from a half-updated graph); the item-131 paging contract
on the search endpoints; per-family data plausibility; and the structured-error negative cases.
What it is not: a substitute for mvn test. The integration tests pin semantics; this pins
"the deployed thing is not obviously broken". A green run is not a quality gate. All assertions are
invariants, never fixed counts — counts move with every refresh.
TypeScript / React projects (item 192)
A project may declare language: typescript (ac project create purfe --root … --language typescript).
Its .ts/.tsx files (not .d.ts) and plain .css files are ingested; node_modules and dist
are excluded by default. Only a typescript project ingests TypeScript — a Java project with a
bundled web UI (ac has ac-ui/) never parses it. Java and Natural files stay language-agnostic.
Identities are paths. A TypeScript MODULE is named by its root-relative path without the script
extension — pur-r-vstamm/src/store/slices/agstammSlice — with simpleName = the file stem
(agstammSlice), workspace = the first path segment, moduleKind = ts/tsx/css, and
generated=true + generator (typescript-generator for the Java-side EndpointGenerator output
under generated/, hey-api for @hey-api/openapi-ts). A CSS module keeps its extension
(pur-ui/src/index.css). Every /modules/{name}/… endpoint accepts either form (item 117), so
ac context agstammSlice works — until two workspaces have a file with the same stem, then use the
path.
Two tiers, like Java/Natural. Project creation and refresh without --deep run the Tier-1
regex outline in Java: module shell (sourceHash, loc/sloc), one FUNCTION per top-level
function / arrow / class (kind = function | component | hook | thunk | styled | class,
exported), one DATA_STRUCTURE per interface/type/enum (dataType says which), and a
REFERENCES edge per import to the module it resolves to (value = the import clause, specifier
= as written). npm packages are not placeholders; they are listed on the module as
externalImports. A deep pass (refresh --deep, ingest by name) runs the Node sidecar
(ac-parser-typescript/sidecar/extract.mjs, TypeScript compiler API, one whole-program run per npm
workspace, 3–5 s and ~0.5 GB each on the pur frontend) and replaces the outline with the checker's
view: exact positions, imports resolved against the file system, and calls:
- a callee owned by a top-level declaration of the same file →
FUNCTION -CALLS-> FUNCTION; - a callee in another module →
MODULE -CALLS-> MODULEcarryingcallKind(METHOD_CALL, orCONSTRUCTORfornew),callerFn(the calling function, absent at module level),calleeMethod(the owning top-level declaration in the target),callSyntax(call/new/tagged/jsx— a JSX element<HistorieDrawer/>is a call), and on a member callreceiver(type of the innermost object, e.g.AgstammControllerEndpoint) andmember(saveBroker.post); - calls into npm packages and the language library are not edges.
These are the exact properties the Java parser writes, so callers/callees/call-tree,
functions/{fn}/callers (item 52) and the placeholder rewiring work unchanged. A module that got its
Tier-2 pass carries ingestTier=2.
Sidecar failure is visible, not silent. If node, the script or its node_modules/typescript
are missing, or a workspace run fails or times out, the refresh still completes at Tier-1 for those
files and the response lists a failure with the pseudo-path sidecar or sidecar:<workspace>.
Config: agenticcode.typescript.node, .sidecar-script, .max-heap-mb (1024), .timeout-seconds
(600); the image carries node and the sidecar (Dockerfile.jvm), dev mode expects
npm ci run once in ac-parser-typescript/sidecar/. The project root is read-only in the
container; the sidecar reads the project's own node_modules for library typings and writes nothing.
Scope of the pur frontend project. The registered project covers the pur-ui and
pur-ui-common workspaces only; pur-r-vstamm and pur-r-vbuch are excluded via excludeDirs,
and the sidecar does not load an excluded workspace. Both workspaces call the pur backend through
the legacy generated client (generated/endpoints.ts, backend pur); the hey-api client shape is
recognised too but is not in scope.
Resolution notes (verified on purfe, 2026-09-22). An import of a workspace consumed through
its package.json exports resolves into its build output (pur-ui-common/dist/x.d.ts); the sidecar
maps that to the source twin (pur-ui-common/src/x.ts) so the edge lands on a real module. A bare
specifier the checker resolves to neither a file nor a package (immer, redux — transitive
dependencies the project does not list) is recorded as an external import, not a placeholder.
Transitive packages are known from root/node_modules (directory names), so Tier-1 treats them
as external too.
Stale parsed edges are reaped on a deep refresh (item 198). Every edge the parser emits is
stamped with the run's ingestGen at merge time; after a deep re-parse of a file, the edges from
that file's nodes whose stamp is older than the run's (the fresh parse did not re-emit them) are
deleted before the node sweep, for every language and edge type. An import or call the new parse
names differently (a renamed class, a dist→src mapping fix) therefore no longer keeps its old
placeholder alive next to the fresh edge, and a placeholder left edgeless falls to the usual
placeholder sweep. Tier-1 (changedOnly or non-deep) refreshes do not reap, because a Tier-1 pass
emits fewer edges than a deep one. Edges persisted before the stamp was introduced carry no
generation and are never reaped; the first deep refresh after upgrading stamps them, the next one
reaps — so a project never needs recreating after a parser change any more, two deep refreshes do.
A module node never lives in a copycode (item 201). A program whose body is a single INCLUDE
used to get a second MODULE node named after it with the .cpy as sourceFile (ZDTSTBP6 in
upms), so modules and search/identifier listed the name twice. Fixed in the parser; a graph
ingested before the fix keeps the stray node until the project is recreated, or you remove it by hand:
MATCH (m:MODULE {project: $p}) WHERE m.sourceFile ENDS WITH '.cpy' AND m.ownerModule = '' DETACH DELETE m
(Natural copycodes are never modules of their own, so the match is exact).
Known limits of 192 (the later items fill them): no field bindings (195), no styles (196);
the store is item 194 below. A changedOnly refresh re-runs the sidecar over the
whole workspace but re-persists only the changed files, so an unchanged file's facts can lag one
refresh (same class of caveat as 46a). The by-name deep ingest resolves dependencies by file stem, so
a dependency whose stem exists in several workspaces (index) is reported as a duplicate and skipped
— use refresh --deep for the whole frontend.
Web-service calls and the counterpart link (item 193)
rest-endpoints lists the frontend's calls. Every member of a generated Endpoint class
(AgstammControllerEndpoint.saveBroker) and every hey-api sdk function is a FUNCTION of
kind=endpoint carrying the same restPath/httpMethod the Java parser writes for a handler, plus
outbound=true — so GET /projects/purfe/rest-endpoints answers with outbound: true rows
(handler = AgstammControllerEndpoint.saveBroker, path = /agstamm/ui). Who calls it:
modules/{generated module}/callers names the calling slices/components (module level), and the
member call api.saveBroker.post(...) is retargeted from the generic PostMethod.post signature to
the endpoint function, so the module-to-module CALLS edge carries calleeMethod = AgstammControllerEndpoint.saveBroker and callerFn = <thunk>. Since item 203 the synthetic
class-hierarchy edges (resolvedVia: INHERITANCE, caller → each implementation of the called
interface/base) exist once per originating call site with its real lineNo, originFile,
calleeMethod and callerFn — before, one edge per pair took whichever call line the merge met first.
So callees lists every real line for an implementation, and functions/{impl-method}/callers also
names callers that go through the interface (RepoImpl.save ← Service.store via Repo.save).
Since item 197
functions/{fn}/callers joins these module-to-module edges back to the calling function, so
purfe/modules/pur-ui/src/generated/endpoints/functions/GeneralAgreementUiControllerEndpoint.createNew/callers
names the thunk in generalAgreementSlice, and on the Java side
pur/modules/…AgstammLogic/functions/handleMerge/callers names AgstammController.mergeBroker
(a REST controller method itself has no Java callers — it is the HTTP entry point). Extra properties on the node (
GET /nodes/{id}): restUrl (as composed,
with placeholders and query string), restBase (an application base such as /pur-r-vbuch/v1
split off so paths compare with the backend's base-less @Path), backend (pur, pur-r-vstamm,
dynamic — from the URL builder), queryParams, requestType, responseType, paramsType,
generator, owner, member. A generated interface's properties are FIELDs under its
DATA_STRUCTURE, named <Interface>.<member> (Broker.ebene; props field = the bare member,
owner = the interface — since item 195: a node's identity is type + name + file, and one generated
file declares hundreds of interfaces, so a bare vid used to be a single node under six interfaces),
so GET /data-structures/AgstammUseCase/fields answers for the frontend too (bare member names;
since item 195 — before, the query filtered FIELD out and returned [] for an interface), and
counterparts?kind=field rows are named Broker.ebene.
COUNTERPART_OF: the same thing in another project. A project setting
counterparts: ["pur"] (ac project create purfe … --counterpart pur, ac project update purfe --counterpart pur, GET /projects/purfe shows it) makes enrichment link, after every refresh of
either side:
| this project | → counterpart | matched on |
|---|---|---|
outbound endpoint FUNCTION |
backend handler FUNCTION |
httpMethod + path shape (every {param} segment compares as {}, class + method @Path composed as rest-endpoints does); several matches → the handler whose source lives under the frontend's backend name |
DATA_STRUCTURE in a generated=true module |
Java MODULE with the same simpleName |
unique name only; an ambiguous name stays unlinked |
its FIELDs |
the class's FIELDs |
name |
The edges are rebuilt from scratch on each run (never accumulated) and re-run for the frontend when the backend refreshes, because a refreshed handler node is deleted together with the edges pointing at it. Roadmap item 143 (Natural ↔ Java counterparts) will use the same edge.
GET /projects/{p}/counterparts?module=&kind=rest|dto|field&unmatched=&countOnly=&limit=&offset=
(CLI ac counterparts [--module] [--kind] [--unmatched] [--count-only]) lists
{kind, name, module, sourceFile, startLine, httpMethod, path, counterpartProject, counterpartName, counterpartModule, counterpartSourceFile, counterpartStartLine}; the counterpart* fields are null
for an unlinked row, and unmatched=true is the planning question: which calls does nothing
serve, which generated DTOs / fields have no backend twin. Paged with X-AC-Total-Count /
X-AC-Truncated like the search endpoints; unknown kind → 400 KIND_UNSUPPORTED; a project
naming itself as counterpart → 400 COUNTERPART_SELF.
The Redux store (item 194)
What is modelled. Every createSlice / createAppSlice in a typescript project is a
STORE_SLICE node named by the reducer key it is mounted under in the project's
configureStore (state.<key>; gruppenprovision for generalAgreementSlice, whose RTK name is
generalAgreement) — the key is what every selector path starts with, so it is the identity; the
RTK name is kept as sliceName. The sidecar traces each reducer: { key: xReducer } entry through
xReducer = xSlice.reducer (also export default xSlice.reducer) back to the slice, across
workspaces; a slice no store mounts is named by its own sliceName. Under the slice, one store
FIELD per top-level state key, named <key>.<field> (schluesseltabelle.sucheStatus, from the
checker's type of initialState, so keys only the state type declares are present too) with
dataType = the TS type, optional, store=true, slice, field. Every reducer is a FUNCTION
of kind=reducer in the slice's module, named by the action type it handles:
schluesseltabelle/updateX for a case reducer (reducerKind=reducer), schluesseltabelle/suche/fulfilled
for builder.addCase(sucheByServer.fulfilled, …) (reducerKind=case, trigger = the expression),
schluesseltabelle/matcher:isSlicePending(sliceName) for addMatcher (reducerKind=matcher). A
thunk whose lifecycle action a case handles CALLS that case (callSyntax=extraReducer), when both
live in the same file.
Reads and writes. Inside a reducer every state.a.b chain is a READS/WRITES edge from the
reducer FUNCTION to the store FIELD of its first key (state alone → the STORE_SLICE), carrying
path (the full sub-path as written, keyTableUseCaseSvcResult.result.tableId), lineNo,
via=reducer. A write is an assignment target (compound assignments also read), ++/--, delete,
a mutating method on the chain (push, splice, set, delete, …), Object.assign(state.x, …),
or a return { … } (each key written) / return other (the whole slice). Outside reducers every
store read is a READS edge from the reading function (component, hook, thunk; the module when at
top level) to a placeholder <key>.<field> that the finalize step resolve-store-placeholder
redirects onto the real field by name and then drops — so a read in KeyTablePage.tsx lands on the
field declared in keytableSlice.ts without either file knowing the other. Recognised read forms:
useSelector/useAppSelector((state) => state.a.b) (every chain rooted at the arrow's parameter;
const { x, y } = useAppSelector((s) => s.a) reads a.x and a.y), wrapper hooks such as
useSchluesseltabelleSelector((useCase) => useCase?.result?.purMode) — the sidecar finds the wrapper's
inner useAppSelector((state) => selector(state.a.b)), so the read is a.b.result.purMode with
via=useSchluesseltabelleSelector (the wrapper's own inner read is recorded too) — and
store.getState().a.b / const s = thunkAPI.getState(); s.a.b chains (via=getState; 81 sites on
pur-ui). A read of a key the store does not declare stays a placeholder and is listed by
search/identifier with an empty sourceFile.
Dispatch. A call to a slice action creator (dispatch(updateX(…)), a binding of
xSlice.actions) carries actionType = <sliceName>/updateX and is a cross-module CALLS edge to the
slice module with calleeMethod = <sliceName>/updateX — the reducer FUNCTION — so module
callees of a component list the slices it dispatches into; a thunk call keeps the thunk FUNCTION
as target and carries the thunk's type prefix as actionType. (Function-level exposure of these
cross-module edges is item 197.)
Endpoints. GET /projects/{p}/store?slice= (CLI ac store [--slice]) → [{slice, sliceName, module, sourceFile, startLine, endLine, stateType, fields: [{name, type, optional, reads, writes}], reducers, reads, writes}], all slices or one by reducer key / RTK name.
GET /projects/{p}/store/{slice}/accesses?field=&mode=reads|writes&module=&countOnly=&limit=&offset=
(CLI ac store-accesses <slice> [--field] [--mode] [--module] [--count-only]) → [{mode, slice, field, path, function, functionType, functionKind, module, sourceFile, lineNo, via}], one row per
access site, field null for a whole-slice access; paged like counterparts; unknown mode →
400 MODE_UNSUPPORTED, unknown slice → 404 SLICE_NOT_FOUND. Because store fields are FIELD
nodes with a unique name, the generic GET /variables/<key>.<field>/reads|writes and
search/identifier?name=<key>.<field> answer too.
Limits. Reads through getState() aliases are followed only inside the file that created the
alias; a thunk→case CALLS edge is emitted only when thunk and slice share a file (the pur slices
do); a case whose trigger cannot be folded to an action type (a predicate matcher) is named by its
expression text; the sidecar reads the store of every workspace in scope — a key mounted only by a
workspace outside the project (excludeDirs) falls back to the slice's own name (matches on pur).
No USES_TYPE from a store field to the DTO it holds yet — the field's dataType says
SvcResult<KeyTableUseCase>, the link is item 195. Only a top-level createSlice is a slice: a
slice built inside a factory function (pur-ui-common's filetransferSlice, created per instance
and not mounted in the pur-ui store) is not modelled.
Verified on purfe (2026-09-22, server 318, recreated + refresh --deep, 285 files, 0
failures, 0 placeholders left): 9 slices — error, global, healthTables, metadata
(pur-ui-common) and gruppenprovisionSuche, gruppenprovision (RTK name generalAgreement),
schluesseltabelle, multilinguism, translationdata (pur-ui) — 43 reducer functions, 216 read
and 82 write sites. schluesseltabelle alone: 62 sites, 25 through useSchluesseltabelleSelector,
19 through getState(), 5 through useAppSelector, 13 in reducers.
DTO field bindings (item 195)
What a binding is. The generator emits, next to every DTO interface, a Fields class tree
(AgstammUseCaseField = new AgstammUseCaseFields<AgstammUseCase, never>(), members
broker = new BrokerFields<TRoot, Broker>(this, "broker"), list members as keyTableList = (index?) => new KeyTableDOFields(...)). A path expression on it — AgstammUseCaseField.broker.ebene on a
<SmartInput field={…}>, <SmartOutput field={…}>, a table's fieldTermForRowData, a column's
field:, or a Fields-typed prop such as useCaseFieldPrefix — is a binding. The sidecar types
every hop through the checker (XFields<TRoot, TSelf>): the root DTO is TRoot, the owner
of the leaf is the TSelf of the hop before it, the leaf name is the field. Each binding becomes a
READS edge (plus a WRITES edge when the component tag matches Input$|Dropzone$|Editor$) from
the binding function (component/hook; the module at top level) to the FIELD of the generated
interface that item 193 already creates (Broker → ebene), carrying path (dotted hops from the
root, brokerList[] for a list hop), rootDto, kind, partial, component, attribute,
via=binding, lineNo.
kind=field: the leaf is a scalar (ebene);kind=prefix: a whole sub-object is handed on (useCaseFieldPrefix={X.tab.translationData},fieldTermForRowData={X.keyTableList()}) — recorded as a read of the container field so nothing is silently dropped.partial=true: the expression is rooted at a prop or local (props.useCaseFieldPrefix.dataName,tabPrefix.x(idx).gausVal), so the leaf and its owner are exact but the prefix of the path is unknown (only the tail is given). A carrier prop's own name is not part of the path.
Cross-file targets are placeholders <Dto>.<field> (binding=true, owner, field,
targetModule) resolved by the finalize step resolve-binding-placeholder exactly — module →
DATA_STRUCTURE → FIELD — and dropped afterwards; a leaf the interface does not declare stays a
placeholder (listed by search/identifier with empty sourceFile). The item-74 stale-edge sweep
covers binding targets on a deep re-ingest.
Endpoints. GET /projects/{p}/bindings?dto=&field=&mode=reads|writes&module=&partial=&countOnly=&limit=&offset=
(CLI ac bindings [--dto] [--field] [--mode] [--module] [--partial] [--count-only]) → one row per
binding site: {mode, dto, field, path, rootDto, kind, partial, component, attribute, function, functionType, functionKind, module, sourceFile, lineNo, counterpartProject, counterpartModule, counterpartField}. The counterpart* columns are the field's COUNTERPART_OF twin (item 193), so
"which page edits Java Broker.ebene" is ac bindings --dto Broker --field ebene --mode writes -p purfe and needs no query on the backend project. dto is the declaring interface (Broker),
not the root (AgstammUseCase) — filter on rootDto client-side when you need the latter. Paged like
counterparts; unknown mode → 400 MODE_UNSUPPORTED. GET /data-structures/{dto}/fields now
returns TypeScript interface fields (type=FIELD) and carries boundReads/boundWrites per field
(0 for Natural/Java). The generic variables/<Interface>.<field>/reads|writes sees binding edges
too. The leaf's owner is the interface that declares the member — datStart bound through
GeneralAgreementDO lands on AbstractHistorizedDO.datStart.
Limits. String-form field="…" bindings and the lodash-path bindings of pur-r-vbuch (excluded
workspace) are not modelled; a partial binding cannot say which list element or tab; no
USES_TYPE from the component to the root DTO (the rootDto edge property answers that). A
binding placeholder and a store placeholder share the FIELD type, so a DTO named exactly like a
store key would merge their placeholders (Dto.field vs key.field) — not the case on pur.
Verified on purfe (2026-09-22, server 322, recreated + refresh --deep, 285 files, 0
failures): 257 binding sites (196 reads, 61 writes; 226 scalar leaves, 31 prefixes; 84 partial)
over 23 declaring DTOs, every one linked to its pur counterpart field; SmartInput 122,
SmartOutput 64, tables 32, column definitions 36. counterparts?kind=field: 744 fields, exactly
one twin each (six-fold fan-out before the qualified names), 112 unmatched.
Styling: theme tokens and the style inventory (item 196)
The theme. The file with createTheme({...}) (pur-ui-common/src/theme.ts) gets a
DATA_STRUCTURE theme (kind=theme) with one FIELD per token: theme.<path> for every leaf of the
literal (palette.primary.dark, typography.h1.fontWeight, shape.borderRadius, sizes.*; props
token, tokenKind=path, value folded through constants — #0054A2 — and constant when the
leaf names one, PRIMARY_DARK) and theme.<NAME> for each exported string/number constant of that
file (tokenKind=constant, PRIMARY). MUI components.styleOverrides are recorded as tokens too, not
interpreted.
Style blocks. Every sx={…}, style={…} and styled(X)(…) block is a STYLE node under its
component FUNCTION (a styled block under the kind=styled function), named
<function>.<sx|style|styled>@<line>:<col>, with styleKind, element (the JSX tag or styled base:
Box, 'div'), properties (the CSS keys, nested selectors flattened: &:hover.color,
& .MuiPaper-root.background), literals (hard-coded colours/lengths: #005CA9, 17px, -2%,
calc(100% - 16px) — mt: 2 is theme-relative and no literal), dynamic (a value the sidecar could
not classify, or a whole sx={props.sx}), spread. Plain .css files get one STYLE per rule from
the Tier-1 scanner (body@7, @font-face@2; styleKind=css, selector, properties, literals).
Token reads. Inside a style block every chain on a theme value is a REFERENCES edge STYLE → theme.<token> with property (the CSS key it feeds) and via=theme; a theme value is anything typed
Theme (useTheme(), a ({ theme }) => styled parameter, the theme object imported under any name)
or named theme. theme.spacing(2) ends at spacing, theme.palette.grey['200'] is
palette.grey.200; a theme constant (color: PRIMARY) references theme.PRIMARY. A token read outside
a style block (borderColor={theme.palette.grey['200']}, code) is the same edge from the enclosing
FUNCTION with context = the JSX attribute or code. Cross-file targets are placeholders
theme.<token> (theme=true) resolved by exact name at finalize when exactly one theme declares the
token; a token no theme declares keeps its placeholder on purpose — GET /theme lists it with
declared=false (MUI defaults such as palette.grey.200, palette.common.white, the spacing
function, or a typo).
Endpoints. GET /projects/{p}/theme?unused= (CLI ac theme [--unused]) → [{token, kind, value, constant, declared, module, lineNo, uses}], declared tokens first. uses counts project
references (style blocks and code); MUI's own consumption of a token is invisible, so unused=true
means "no project code references it", never "safe to delete" (palette.primary.main colours every
Button whether or not a component names it). GET /projects/{p}/theme/{token}/usages (ac theme-usages palette.primary.dark, token = dotted path or constant name) → [{function, functionKind, module, sourceFile, lineNo, styleKind, element, property, context}]; unknown token →
404 TOKEN_NOT_FOUND.
GET /projects/{p}/styles?module=&kind=sx|style|styled|css&withLiterals=&countOnly=&limit=&offset=
(ac styles [--module] [--kind] [--with-literals]) → [{name, styleKind, element, selector, function, module, sourceFile, lineNo, properties, literals, dynamic, tokens}], paged like
counterparts; withLiterals=true is the review question: which blocks hard-code colours and
lengths instead of using the theme. Unknown kind → 400 KIND_UNSUPPORTED.
Several themes (item 199). Every createTheme({...}) in a file is read; a token two themes of
one file declare is one row with the first theme's value and variants: 2. A token declared by two
theme files (light/dark) is one row per file, and a read of it resolves onto both, so each row
counts the use and ?unused=true stays honest; theme/{token}/usages and styles[].tokens
report such a read once. Two reads of one token on one line of a style block (color: PRIMARY, borderColor: PRIMARY) are one usage whose property is the comma list of the keys they feed.
CSS rules: a block-less @import/@charset line is not part of the next selector, braces inside
string values do not open blocks, a nested rule head inside an at-rule body is not a declaration,
and a selector repeated on one line (minified CSS) gets a :col suffix in its name.
Limits. className strings are not matched to CSS rules; Emotion css templates and MUI
styleOverrides are not modelled; a token read from a component-level function (not a style
block) that survives a re-parse keeps its resolved edge until the file's nodes are re-created
(same class as item 198).
Verified on purfe (2026-09-22, server 326, recreated + refresh --deep, 285 files, 0
failures): 129 declared tokens (111 paths, 18 constants), 93 of them with no project reference;
12 undeclared tokens the code reads (palette.common.white 4, palette.grey.200 4, spacing 3,
palette.divider, palette.text.secondary, applyStyles, transitions.create, …);
palette.primary.dark is the most-used token (26 reads: 17 style, 3 sx, 1 styled, 5 as a
plain prop such as confirmColor). Style inventory: 304 blocks (196 sx, 78 style, 27 styled,
3 CSS rules), 109 with hard-coded literals (100% 29, 1px 21, 12px 14, 17px 14, …), 75
reading theme tokens, 41 dynamic. The only placeholders left in the project are the 12 undeclared
theme tokens — by design.
Missing capability?
If the API/CLI genuinely cannot answer a question (not just
unreachable — the capability doesn't exist), finish the task via
grep/Explore as a fallback, then use AskUserQuestion to flag the gap and
ask whether it should become a roadmap item in x-docs/roadmap.md. Don't
silently fall back and move on.