Files
agenticCode/x-docs/agent-api-usage-ac-implementation.md
Ingo Schnabel 234ea76911 Roadmap
2026-09-23 15:31:45 +02:00

1447 lines
211 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Using AgenticCode on This Repo (Dogfooding)
This repo is ingested as project **`ac`** at `http://localhost:8787`. Per
CLAUDE.md's dogfooding rule, prefer the AgenticCode API over grep/Explore for
call graphs, callers/callees, DB access, dataflow, and module overviews when
working *on this repo* — it's exactly the tool this project builds.
**For full API semantics (params, response shapes, error codes, language
applicability, curl examples)**, see
`x-docs/agent-api-system-prompt.md` — that file is the canonical reference and
is not duplicated here. This file only covers what's specific to using the
API *as Claude Code, on this checkout*.
## Tool priority
REST first (`GET /api/projects/ac/...`) → `ac`
CLI (`ac callers`, `ac callees`, `ac call-tree`, `ac context`, `ac
db-accesses`, ...) → grep/Explore. Only fall back past REST/CLI when the
question genuinely isn't answerable by this API at all (see "Missing
capability" below). If the server is unreachable, try
`./manage-ac.sh deploy` before falling back further.
## Re-ingest before trusting results
Query results reflect the last ingest, not the current working tree.
**Refresh after code changes** before trusting query results:
`ac refresh` or `POST /api/projects/ac/refresh` (add `--deep` / `?deep=true` for a
full field-level pass). `refresh` is the single (re-)ingest surface (item 42) — the
eager `ingest-all`/`ingest-module`/`ingest-call-graph` endpoints were removed.
`ac refresh <name>` deep-ingests one module + its callees/data areas; add
`--neighborhood` (`POST /refresh/{name}?scope=neighborhood`) to also pull in the
module's transitive **callers** (whole call-graph neighbourhood).
**Reconciliation on re-ingest (item 58).** A `refresh` now **purges stale nodes**:
for every re-parsed file it deletes the nodes the fresh parse no longer produces
(renamed/removed fields, moved statements) rather than leaving them to shadow the
new ones — so identifier counts and `/search/identifier` results stay clean after a
parser change or an edited source file. Applies to every full-parse path (whole-root
`refresh` with or without `--deep`, and per-module `refresh/{name}`); the coarse
Tier-1 scan run at project creation does not reconcile (it only ever runs on an empty
graph).
**Auto-invalidation (item 43).** The graph self-heals when files change or are deleted on
disk — you rarely need a manual `refresh` for staleness:
- **Changed file:** a field-level query on a module whose source changed re-ingests it
automatically (its stored `sourceHash` no longer matches), so results reflect the current
file. Granularity is per-module: a changed *dependency* is picked up when that dependency
is itself queried by name.
- **Deleted file:** a **whole-project** `refresh` now also **removes nodes for files deleted
from disk** (reconciled against the filesystem), closing the earlier gap — no project
re-create needed. (A per-module `refresh/{name}` does not sweep; it only touches its own
tree.)
- Controlled by `agenticcode.auto-invalidate.enabled` (default `true`).
## Reading source in this repo
You have direct filesystem access to this checkout — **never** call
`nodes/{id}/source` or `modules/{name}/source`. Use `sourceFile`/`startLine`/
`endLine` from a graph response (context, digest, search/identifier,
nodes/{id}, ...) and read the file directly. This is always cheaper and gives
full surrounding context; the `/source` endpoints exist only for API-only
agents with no filesystem access.
**Stale-source check (item 41).** The `/source` endpoints compare the file on
disk against the content hash (`sourceHash`) stored at ingest. If the file
changed since the last ingest they return `409 STALE_SOURCE` instead of slicing current text against old line
numbers — re-ingest (refresh) the project to update the graph. Line ranges you
read directly off disk are of course always current; this only guards the API's
own slicing. Copycode/INCLUDE slices are raw pre-expansion file text.
## Comments: reachable, but never by default (item 141)
**Every default query sees code only.** `search/identifier`, `search/annotation`,
`search/references` and a plain `search/value` never match comment text — a Natural `* ...` banner,
a trailing `/* ...`, a Java `//` line or a Javadoc block.
That silence used to be a **wrong answer, not a missing one**, wherever a convention records
something in a comment. The UPMS→PUR case: a reengineered service carries its Natural origin in a
Javadoc block (`ServiceEndpoint:` / `UPMSFunction:` / `UpmsObject:`), and the documented
Natural→Java lookup searches the program name in `pur`, where an empty result is read as *"not yet
reengineered"* — which it answered for services reengineered months earlier.
Three routes now reach comment text. Pick one before concluding "not present":
| Question | Call |
|---------------------------------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| "What does this module's header/change log say?" | `GET /modules/{name}/comments` (`ac comments <module>`) — blocks with the declaration each documents |
| "Where is this query / SQL / JSON text defined?" (Java) | `GET /search/value?value=…&contains=true` — since item 139 a `static final String` built from text blocks, literals, same-class constants and `+` carries its full text as `value`; a `"…".formatted(...)` constant carries its template (with `%s`) and `valueKind: template`. Anything built from a method call or another class's constant stays unresolved (no value) |
| "Does this string appear anywhere, code **or** comment?" | `GET /search/value?value=…&includeComments=true` (`ac search-value --include-comments`) — comment hits carry `kind: "COMMENT"` |
| "…and in text the parsers do not model at all, or in a module that is not deeply ingested?" | `GET /search/source?regex=…` — raw grep over the files on disk |
```
GET /pur/search/value?value=WPARTX0S&contains=true → [] (code only)
GET /pur/search/value?value=WPARTX0S&contains=true&includeComments=true → the Javadoc origin block
GET /upms/modules/WAGNTX0S/comments → `* #01 … Bug 266`, …
```
Comments stay **opt-in** deliberately: a comment hit is not the same evidence as a literal in code,
and folding them into the default result set would move every existing completeness count (item
131's lesson). The flip side is the rule to remember — **an empty default search says nothing about
comments.**
## Diagnosing a slow refresh (items 153/155)
Two instrumentation layers, both aimed at the same question — *where does the time go?*
* **Always on:** every persist batch logs one line with its statement breakdown
(`Persist deep 200 files: total 8740 ms = prepare 108, merge-nodes 1630, ..., commit 20`), and every
enrichment step logs duration plus created/deleted rows. The residual `commit` is deliberate: total
minus the labels is the transaction commit, so nothing hides in an unnamed remainder.
* **Opt-in per run:** `POST /refresh?deep=true&profile=true` (`ac refresh --deep --profile`) runs the
enrichment steps under Cypher `PROFILE` and logs the five heaviest operators of every step slower
than 5 s. Diagnostic only — it answers "is this step matching or writing?", which the step timing
alone cannot. Measured overhead on `upms`: none worth reporting (908 s vs 905 s).
* **Build-time diagnostic:** `agenticcode.ingest.comments.enabled=false` drops `COMMENT` nodes and
their `DOCUMENTS` edges just before persist. It exists to measure what comments cost (item 179:
~17 % of a deep refresh) and is **not** a supported operating mode — with it off, `/comments`
answers empty. It is a property, not a request parameter, so it needs a rebuild and cannot be set
per run.
What that measured, so nobody re-derives it. The figures below are **current** (2026-09-06, two clean
runs, roadmap item 175); the campaign of items 153-175 took an `upms` deep refresh from **1 225 s to
301 s**, so any older number quoted elsewhere is stale by a factor of four.
| | now |
|---------------------------------|-------------------------------------------------------------------------------------------------------------------|
| deep refresh `upms`, end to end | **301 s** (run-to-run spread 3.7 %) |
| persist | ~149 s — `merge-edges` 47.6 s, `commit` 34.1 s, `merge-nodes` 24.0 s |
| finalize | ~137 s — `resolve-field-placeholder` W/R 43.4 s, `link-args-to-params` 23.7 s, `resolve-bare-included` W/R 37.9 s |
| parsing | ~9 s (interleaved with persist) |
The original diagnosis still holds and is why those steps shrank: the five field-resolution steps
were dominated by *matching*, not writing — `resolve-bare-included` once spent 273 M database hits to
produce 43 753 rows. See roadmap items 156 and 175, and item 179 for what comment nodes cost.
## `limit`/`offset` work on some endpoints and are silently ignored on others
**Every endpoint answers completely.** The difference is whether it lets you ask for less. 19 list
endpoints declare no `limit`/`offset` at all, and JAX-RS drops an undeclared query parameter without
a word — so `?limit=2` there is not an error, it simply has no effect:
```
GET /upms/modules?limit=2 -> 3 587 rows
GET /upms/modules/WAGNTX0S/functions?limit=2 -> 22 rows
GET /upms/modules/WAGNTX0S/data-structures?limit=1 -> 16 rows
```
**Ignore `limit`/`offset` (always the full list):**
`/modules` · `/modules/{name}/functions` · `/modules/{name}/functions/overrides` ·
`/modules/{name}/functions/{function}/overrides` · `/modules/{name}/data-structures` ·
`/modules/{name}/columns` · `/modules/{name}/payload` · `/modules/{name}/dispatch-table` ·
`/modules/{name}/sql-statements` · `/data-structures/{name}/fields` · `/db-tables/{name}/columns` ·
`/variables/{name}/reads` · `/variables/{name}/writes` · `/variables/{name}/flow-forward` ·
`/variables/{name}/flow-backward` · `/variables/{name}/field-flow` · `/duplicates` ·
`/dynamic-calls/unresolved` · `/dynamic-calls/overrides`
**Honour them** (14, verified against the method signatures): the search endpoints
(`search/identifier`, `search/value`, `search/annotation`, `search/references`, `search/source`),
`rest-endpoints`, and per module `callers`, `callees`, `context`, `db-accesses`, `graph`, `reaches`,
`workfile-accesses`, `comments`.
`call-tree` is bounded differently again — by `depth`, not by row count — which is the right shape
for a tree but means `limit` does nothing there either.
Which direction the mistake runs matters, so be precise about it: a dropped `limit` means you get
**more** than you asked for, never less. It costs tokens, never correctness — the opposite of item
131's failure, where a silent 50-row cap was read as the complete set. Nothing here can under-report.
The practical consequence is budget, not trust: `/modules` on `upms` is 3 587 rows in one response.
Narrow with the filters those endpoints *do* have (`?kind=`, `?sourceFile=`, `?module=`,
`?extendsName=`) rather than with a `limit` that will be ignored, and prefer `/modules/{name}/digest`
or `/context` when you want an overview rather than an enumeration.
*(Implementing `limit` on those 19 was considered and deliberately not done: the endpoints are honest
as they stand, and a `limit` that ever acquired a default would reintroduce exactly the silent
truncation item 131 removed.)*
## `callers` / `callees` say when they are cut (item 181)
`callers`, `callees` and `functions/{fn}/callers` now send `X-AC-Total-Count` and `X-AC-Truncated`
like the search endpoints, and their body carries `total` and `truncated` next to `sourceFiles` /
`items` (also for `fields=name`; `ac callers` / `ac callees` print the usual "truncated" warning).
The default page is still 50: `upms/modules/DPARTFN0/callees` answers 50 of 58 with
`X-AC-Truncated: true` — before, the 8 missing callees (among them the module that writes the
partner) looked like an inconsistency with `digest`. **Read `truncated` before concluding "X does not
call Y"**; ask with `limit=1000` or narrow with `scope=external`.
## Truncation is now visible on the search endpoints (item 131)
`search/identifier`, `search/value`, `search/annotation`, `search/references` and
`rest-endpoints` send two headers with every answer:
| Header | Meaning |
|--------------------|---------------------------------------------------------|
| `X-AC-Total-Count` | how many rows match in total, ignoring `limit`/`offset` |
| `X-AC-Truncated` | `true` when this page leaves some out |
and all five accept **`?countOnly=true`** (CLI `--count-only`), returning `{"count": n}` instead of
rows — a completeness question is a counting question, and `@Column` on `pur` is 3 630 rows ≈ 250 k
tokens if you ask for them.
This closes item 103's own follow-up. The bodies stay bare arrays (no contract change), for the same
reason as item 130's scope headers. Why it matters: `search/annotation?name=Immutable` returned 50 of
95 rows with no total, no flag and no `Link`/`X-Total-Count` — and a real UPMS→PUR audit read that
page as the whole set, recording that 17 entities had lost `@Immutable` when **zero** had. Re-checked
against all 114 rows of `Tables_meta.csv`: 94 non-writable carry it, 20 writable do not, no
deviations. The finding cost a day and was pure artefact of the cut.
**`search/references` and `rest-endpoints` joined this late (item 135, 2026-08-20).** Item 131 was
written about the annotation search and both were overlooked. `search/references` was the damaging
one: it capped at the default 50 and said nothing at all, so a rename scoped from that page missed
every site past the fiftieth and looked complete doing it. `rest-endpoints` defaults to an uncapped
limit and so never lost rows, but it was equally silent about how many there are. Both were found by
`x-scripts/verify-api.sh` on its first run, not by a test.
Two mechanics worth knowing: the total costs a **second query only when the page comes back full**
(a short page is provably the end, so the total is arithmetic), and a total that divides evenly by
`limit` makes the last full page report `truncated` with the next page empty — one wasted call, never
a wrong answer. `ac` prints a note to **stderr** when a response is flagged truncated, so piping the
body into `jq` stays clean.
## REST surface and scope headers (item 130)
`GET /api/projects/{p}/rest-endpoints?module=&countOnly=&limit=&offset=` (CLI `ac rest-endpoints`) lists
`{httpMethod, path, module, moduleSimpleName, handler, sourceFile, startLine}` — the **composed**
path (class-level `@Path` + method-level `@Path`), so "which code runs for `POST /partners`" is one
call. Previously the two halves had to be joined by hand from two `/search/annotation` calls, because
annotations are stored by name without their arguments; the parser now persists `restPath` and
`httpMethod`, including a `@Path` written as a constant reference. A method with no HTTP-verb
annotation is not an endpoint and is excluded. A class with no `@Path` of its own inherits the
nearest one from its `extends`/`implements` ancestry, as JAX-RS does. Rows carry `outbound: true`
when the declaring type is a `@RegisterRestClient` interface — a call the application *makes*, not one
it serves; its path is usually empty because the base URI comes from configuration.
**Scope and freshness now ride on every project-scoped response as headers:**
| Header | Meaning |
|--------------------------|-----------------------------------------------------------------------------------------------------------------------------|
| `X-AC-Exclude-Dirs` | directories the ingest skipped, or `(none)` |
| `X-AC-Ingested-At` | when the graph was last walked (item 126) |
| `X-AC-Ingest-Incomplete` | `true` while a whole-root pass runs or after one that never finished (item 129); `unknown` when no ingest was ever recorded |
Read `X-AC-Exclude-Dirs` before trusting an **empty** answer: "no callers" means "none outside tests"
in a project excluding `test` (`app`) and "none at all" in one that does not (`pur`, `ac`) — the
bodies are identical.
They are **headers, not body fields**, because most endpoints answer with a bare JSON array
(`db-accesses`, `functions`, `search/identifier`, …); adding a field there would mean restructuring
array → object and breaking the web UI's generated client, the CLI printers and any agent that
indexes `[0]`. The trade-off is that an agent reading only the JSON body will not see them — so if
you consume this API programmatically, read the headers too. The project shell is cached for ~10 s to
keep this off the request's critical path, and the ingest path invalidates that cache explicitly, so
`X-AC-Ingest-Incomplete` flips as soon as a refresh starts rather than up to 10 s later.
## Every reference site of a name (item 128)
`GET /api/projects/{p}/search/references?name=&kind=&countOnly=&limit=&offset=` (CLI `ac references <name>`)
returns `{sourceFile, lineNo, kind, inModule, target}` per **mention** of a type — not just per call:
| `kind` | Where it comes from |
|-------------------------|--------------------------------------------|
| `CALL` | a call site (`CALLS`) |
| `IMPORT` | an `import` of the type |
| `TYPE` | a declared field / parameter / return type |
| `ANNOTATION` | the type used as an annotation |
| `EXTENDS`, `IMPLEMENTS` | inheritance |
| `INJECTS` | CDI wiring |
| `CLASS_LITERAL` | `X.class` in argument position |
| `INCLUDE` | Natural copycode inclusion |
Use it to **scope a rename**. `callers` sees calls alone, so a file that only imports the class,
declares a field of it, or names it in an annotation was invisible — and the rename that missed it
looked complete. The `name` may be the identity (FQN) or the short form; `target` echoes what it
resolved to. An unknown `kind` is `400 INVALID_KIND`, never an empty list.
**Known limits, by design:**
* **Local-variable types and generic type arguments are not indexed** — `List<Target> x` records
`List`, not `Target`. They multiply edge volume for much less value than the positions above.
* **Same-package references have no import**, so within one package the index rests on declared-type
positions alone.
* **Imports are only indexed when they look project-internal** (they share the first two package
segments with the importing file). Otherwise every `java.util`/framework import would mint a
placeholder node on every ingest, just for the finalize sweep to delete it again.
* **Natural has no import or type-position concept.** It contributes `CALL`, `INCLUDE` and inheritance
kinds only; this is not parity with Java and should not be read as such.
* Reference edges are written **at parse time**, so they only exist for files re-parsed since this
landed — a project needs a `refresh` before the index is complete.
* Mentions use their own `MENTIONS` edge type, kept out of `CALLS`/`REFERENCES` deliberately: the
call-graph traversals (`callers`, `callees`, `call-tree`, `ego-graph`) follow `REFERENCES` as
wiring, so folding imports into it made an `import` surface as a **caller**. `/search/references` is
the only endpoint that reads `MENTIONS`; the call graph is unchanged.
## Refreshing only what changed (item 129)
`POST /api/projects/{p}/refresh` (CLI `ac refresh`) has two ways to avoid re-walking a whole root:
| Form | What it does |
|---------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `?paths=a/B.java,c/D.java` (`ac refresh --paths a/B.java,c/D.java`) | Re-ingests exactly those relative paths, deep, plus their dependencies. Paths that match no file — **or that are not ingestible source files at all**, like `pom.xml` — come back in `unresolved`; a typo'd path is never silently dropped. Does **not** run the deleted-file sweep and does **not** move `ingestedAt`: both need a whole-root walk. |
| `?changedOnly=true` (`ac refresh --changed-only`) | Whole-root walk, but re-parses only files whose content hash differs from the graph's (files with no stored hash count as changed). |
**`changedOnly` is opt-in on purpose.** Three whole-walk behaviours are reduced, and one of them would
be outright corruption if it were hidden:
* **A changed Natural copycode disables skipping for that entire run** (logged). Natural projects
only — a `.cpy` sitting in a Java project (as a test fixture, say) is not inlined by anything and no
longer stands the optimisation down. Copycode text is
inlined into the including module *at parse time*, so a module whose `.cpy` changed parses
differently while its own hash is unchanged — skipping it would leave a stale expansion behind with
nothing to indicate it.
* **Duplicate-identity detection** only sees the changed files, so it can confirm duplicates among
them but not discover new ones elsewhere. Existing markers are never cleared.
* **User-exit LoC annotation** (item 47) is re-stamped only on re-parsed files.
Enrichment is project-wide and still runs in full, so this cuts **parse+persist** time only — not the
finalize pass. On small projects the whole deep refresh is already ~30 s, so measure before assuming
a win.
**`ingest.incomplete`** (on `/projects` and `/projects/{p}`) is `true` while a whole-root pass runs and
**stays true if one never finished** — a crash, a container stop, an aborted deep refresh. Before
this, an interrupted deep refresh was indistinguishable from a clean graph: the enrichment steps that
already ran are committed, so queries keep answering, just from a half-updated graph. It cannot
self-heal (a killed process clears nothing) and does not distinguish "running right now" from "died an
hour ago" — both mean the same thing to a caller. A completed refresh clears it.
## Renaming a project (item 202)
`POST /api/projects/{p}/rename` with `{"newName": "…"}` (CLI `ac project rename <old> <new>`)
moves the project key on every node and override and in other projects' `counterparts` lists, then
the shell; the old name answers `404` afterwards and nothing needs re-ingesting. Refusals:
`400 INVALID_REQUEST` (blank or unchanged), `404`, `409 PROJECT_EXISTS`. The node rewrite is batched;
an interrupted rename is finished by running it again (97 s for the 939k-node `upms`). Use it to
keep a reference graph next to a fresh ingest (`ac project rename upms upms_alt`, then create `upms`
again and compare). A full `DELETE` of a project now also removes its manual overrides; `recreate`
keeps them.
## Is this project's graph any good? (item 126)
`GET /api/projects` and `GET /api/projects/{p}` (CLI `ac project list` / `ac project show <p>`)
carry an `ingest` object describing the **last whole-root ingest**:
```json
"ingest": { "ingestedAt": "2026-08-18T10:12:44Z", "mode": "full", "filesExamined": 2981,
"filesPersisted": 2977, "filesFailed": 4, "failures": ["a/B.java", "..."],
"failuresTruncated": false, "durationSeconds": 176, "serverVersion": "…" }
```
Use it before trusting a **negative** answer: without it, "no such module" and "that part of the
project was never ingested" are the same empty response. Three rules the field obeys:
* **`ingest: null` means never recorded**, not "ingested nothing" — a project last walked before
this existed reads as null rather than as a fabricated zero.
* **Only whole-root passes write it** — the create-time Tier-1 scan, `refresh`, `refresh?deep=true`.
A by-name `refresh/{name}`, a deep ingest or a fan-out warm ingests real files but sees a fraction
of the tree, so it deliberately leaves `ingestedAt` alone; otherwise deepening one module would
advertise the whole project as freshly walked.
* **`ingestedAt` is not a freshness guarantee.** It says when the walk ran, not that the graph still
matches disk — a file edited a minute later is stale while the timestamp still looks recent. For
the real check, read a file through `GET /{p}/source?file=…`, which answers `409 STALE_SOURCE`
when the content no longer matches the ingested hash. (Targeted/incremental refresh is item 129.)
`failures` is capped at 200 paths while `filesFailed` stays exact; `failuresTruncated` says whether
the list was cut, so a short list is never mistaken for the whole story.
## Tier-1 coarse scan on project create (item 36)
Creating a project (`POST /api/projects/{p}`) now runs a **Tier-1 coarse reference scan** of the
root before returning, so the project is immediately queryable — no separate ingest call. The scan
is lexer-level (Natural) / a coarse projection of the full parse (Java): it records module/function
/data-structure **shells**, the **identifier index** (declared fields, class members), and coarse
`CALLS`/`READS`/`WRITES`/`INCLUDES` references (Natural `PERFORM`, `CALLNAT '...'`, dynamic `CALLNAT
PGM-VAR`, `PARAMETER`/`LOCAL USING` copybooks; Java resolved calls/type refs) — but **no** deep
bodies (control flow, statement-level dataflow, arg→param). Scanned modules land
`CALL_GRAPH`/`NOT_INGESTED`; field-level detail is filled in by the on-demand deep ingest below. Each
module shell carries a `sourceHash`. Disable with `agenticcode.tier1.scan-on-create=false` (creates
an empty project shell to scan later). A missing/unscannable root is non-fatal — the project is
created empty and can be re-scanned.
## Unresolved references (item 40)
A reference whose target isn't (yet) ingested — a `CALLNAT`/`PERFORM`/`USING` to a module/copybook
absent from the project, or a dynamic `CALLNAT PGM-VAR` whose literal can't be recovered — is stored
as a **deduped placeholder node** (blank `sourceFile`). Enrichment stamps each with an `unresolved`
boolean: `true` while genuinely dangling, `false` once a real definition of that name is ingested.
`/search/identifier` returns it as `unresolved` on each `IdentifierMatch`, and `GET /nodes/{id}` carries it in the
node's properties — so an agent can tell a dangling/dynamic
reference apart from a resolved one. In `callers`/`callees` such targets already appear as entries
with a blank `sourceFile`.
### Querying a module endpoint: four answers, not one (items 107, 115, 114)
Every `GET /api/projects/{p}/modules/{name}/…` endpoint used to answer `200` with an all-zeros shell
for a name that exists nowhere in the graph, byte-identical to a real-but-empty module's answer. That
is not cosmetic: `call-tree?depth=4` returning `200` with 0 modules reads as **"analysed, nothing
found"** when the truth is **"not analysable"** — the module's source was never in the checkout. The
graph knows four states and each now gets its own status:
| state | answer | what it means |
|-------------------------------------------------------------------------------|---------------------------------------------------------------|------------------------------------------------------------------------|
| no `MODULE` node for the name | `404 MODULE_NOT_FOUND` | unknown name (typo, or not in this project) |
| **placeholder** — node exists because something calls it, source never parsed | `409` + `{status:"NOT_INGESTED", module, detail, nextAction}` | knowable in principle, not analysed yet |
| **ambiguous** — several real modules share the name (Java) | `409` + `{code:"AMBIGUOUS_NAME", details:{candidates:[...]}}` | the name does not identify a module; pick one |
| **duplicate identity** — the same name in several files, skipped at ingest | `409` + `{code:"DUPLICATE_IDENTITY", details:{paths:[...]}}` | exists twice, deliberately not ingested (item 114); see `/duplicates` |
| real, ingested module | `200` | the body is the answer — an empty body genuinely means "nothing found" |
**Ambiguous names (item 115).** A Java simple name is not unique: nested `@Nested` test classes,
`Builder`, `Config`, `WorkingStorage`. Measured on `pur`, **163 names covering 385 modules (~8%)** were
addressable only ambiguously, and the endpoints used to answer with the *union* across unrelated
classes — `/modules/BrokerHistoryTests/functions` returned 627 functions for a class that has 107.
They now refuse and list the candidates. Repeat the request with `?sourceFile=<candidate>`:
```
GET /modules/Shared/functions → 409 AMBIGUOUS_NAME, candidates ["a/Shared.java","b/Shared.java"]
GET /modules/Shared/functions?sourceFile=a/Shared.java → 200
```
**A Java module's name IS its fully-qualified name** (item 117) — `com.example.OrderService`, and
`com.example.Outer.Inner` for a nested class. That is what makes same-simple-name classes
distinguishable at all. Every module endpoint accepts **either** form:
```
GET /modules/com.example.OrderService/digest → 200, always exact
GET /modules/OrderService/digest → 200 when unique, else 409 AMBIGUOUS_NAME
```
Responses carry `simpleName` alongside `name` for display. The `?module=` and `?extends=` filters and
the `ac` CLI take either form too; `--source-file` remains available on every module command.
**Which Java types are modules** (item 119). Classes, interfaces, **enums, records and annotation
types** — `moduleKind` is one of `CLASS | INTERFACE | ENUM | RECORD | ANNOTATION` (Natural adds
`PROGRAM | SUBPROGRAM | …`), and `?moduleKind=` filters on it. Their content is modelled the way each
kind carries it: a record's components and an annotation type's members are `FIELD`s (the latter with
`defaultValue` where declared), an enum's constants are `CONSTANT`s, and an enum's or record's
`implements` is a real edge — so a call against an interface fans out to an enum implementing it.
Before 119 these three kinds were not parsed at all: `GET /modules/SomeEnum/digest` answered `404`,
and a record referenced from elsewhere stayed an unresolved placeholder (`409 NOT_INGESTED`). One
gap remains by design — a record's *compact* canonical constructor is not a function node, so calls
made in its body are invisible.
Natural is unaffected throughout: its module names are file stems, and colliding identities are
skipped at ingest, so they are unique by construction (`upms` has zero ambiguous names, `pur` 163).
One limit worth knowing: only the request's **root** module is disambiguated. Inside a traversal this
no longer merges anything, because module names are unique per project after item 117 (`pur` and `upms`
have zero duplicate names) — a reference the parser could not qualify is returned flagged
`unresolved: true` rather than attached to an arbitrary candidate.
A name skipped at ingest because it exists in **more than one file** answers `409 DUPLICATE_IDENTITY`
with the conflicting `details.paths` (item 114); `GET /duplicates` lists them all. Note what that does
*not* fix: the skipped file's own calls were never parsed, so caller lists elsewhere can still be
short — they just no longer look complete.
The `409` applies to everything derived from the module's **own** source: `digest`, `context`,
`call-tree`, `callees`, `db-accesses`, `workfile-accesses`, `sql-statements`, `functions`,
`functions/overrides`, `functions/{fn}/overrides`, `functions/{fn}/callers`, `data-structures`,
`dispatch-table`, `payload`, `columns`.
`callers` and `graph` stay `200` for a placeholder — their data comes from the **calling** modules'
source and is genuine. When you get a `409`, **fall back to `/callers`**: it is the one honest answer
available for a module whose own source is missing. (`digest` no longer surfaces those callers, since
its other fields would all be structurally zero.)
`/modules/{name}/source` is unchanged: it already answered `404 MODULE_NOT_FOUND` for both an absent
module and a placeholder, since there is no source to serve either way.
**Scope limit — this guards the *root* module of a request only.** A `call-tree` that traverses *into*
placeholder targets still reports that subtree as empty without flagging it, so a dispatcher whose
targets are all placeholders still returns `200` with a silently truncated tree. Cross-check the
targets you care about individually (a `409` tells you it is unanalysed) — see item 103.
**Data literals are not call targets (item 62).** A `CALLNAT <bareword>` whose target is really a data
value — a browse key reaching the call site through a copycode/macro argument — used to leave a
permanent unresolved placeholder callee. Enrichment now reaps such a placeholder when it has no Natural
sigil (`#`/`&`/`+`), *is* a real `VARIABLE`/`CONSTANT` of the project, matches no real `MODULE`, and is
only ever reached by inferred (`CALLNAT_DYNAMIC`/`INCLUDE_MACRO`) edges. So `callees`, `call-tree`, the
ego graph and the unresolved list no longer list data names as modules. Genuine dynamic dispatch
(`CALLNAT #PGM-VAR`, sigil'd) is still reported as an unresolved target, and a static `CALLNAT 'X'` is
always trusted even when `X` collides with a field name.
**The ingest summary agrees with the graph (item 64).** A `refresh`/`refresh/{name}` response's
`unresolved` list is built during the file walk, independently of the graph — before item 64 it therefore
reported data fields as missing modules (`MODULE CO-TABLA`, `MODULE NAME-DESC-SP`) that enrichment had
already reaped, so the summary and the graph disagreed. It now applies the same item-62 gates as the graph
reap, with the "is this a real field?" test asked project-wide (a data literal is often declared outside
the ingested tree). Genuinely missing modules are still reported — including sigil'd dynamic dispatch
(`MODULE #GETSHORT-MODUL`) and a missing module whose name collides with a field name but is called
statically (the `RPC-CNTX` class). Treat `unresolved` as "dependencies that really are absent".
**Constant-folded string-assembled targets (item 83).** A dispatcher often builds the `CALLNAT <var>`
name from a base literal plus one or more `SUBSTR` overlays — e.g. `#GETSHORT-MODUL` in `YGEAGGNH`,
assembled by `MOVE 'YGEAGKEY' TO #GETSHORT-MODUL` then `MOVE 'GN0' TO SUBSTR(#GETSHORT-MODUL,6,3)` →
`YGEAGGN0`. The parser records each `SUBSTR` write as a `WRITES` carrying `substrPos`/`substrLen`
(1-based), and the `resolve-dynamic-callnat-fold` enrichment step folds the last full-var literal
written before the call site with the intervening overlays (`left`/`substring`) and MERGEs a resolved
`CALLS` edge (`callKind=CALLNAT_DYNAMIC`, `folded=true`) to the assembled module when it is a real
`MODULE`. It runs in every ingest mode (cheap, bounded by dynamic call sites) and over-approximates like
the other dynamic resolvers. This auto-recovers the `Y…GNH → Y…GN0` family with no manual override, so
folded sites drop out of `dynamic-calls/unresolved` and the assembled target appears in
`callees`/`call-tree`/`graph`/`ego graph` tagged `CALLNAT_DYNAMIC`.
**Pin what the resolvers can't: manual dynamic-`CALLNAT` overrides (item 82).** Some `CALLNAT <var>`
targets are beyond the auto-resolvers — a name never present as a recoverable literal in the reachable
code (read from a work-file record, supplied by an unlinked caller, …). Such a site stays an unresolved
placeholder. A human or agent resolves it via
`POST /api/projects/{p}/dynamic-calls/overrides` with the call site's `originFile` + `lineNo` (from
`GET .../dynamic-calls/unresolved`) and the target module name(s) — multiple targets for a genuine
branch. The override is stored as a `:DynamicCallOverride` node **outside** the `:AstNode` graph, so a
refresh never deletes it and an enrichment step (`apply-manual-dynamic-callnat`, after the auto
dynamic-CALLNAT resolvers, before the placeholder cleanup) **re-applies it automatically** — MERGEing a
`CALLS` edge (`callKind=CALLNAT_DYNAMIC`, `resolvedBy='manual'`) to each target and flagging the
placeholder `manualHidden` so `callees`/`digest`/`graph`/`call-tree` show the real target, not the
`#var`. It only applies while the site is still unresolved: once an auto-resolver catches up, the
override is skipped and listed `obsolete` — **except a constant-fold (item 83), which a manual override
outranks**: the fold skips a site carrying a `:DynamicCallOverride`, and a stale `folded` edge there is
dropped (`delete-folded-overridden-dynamic-callnat`) before `apply-manual-dynamic-callnat` runs, so the
pinned target replaces it. `DELETE .../dynamic-calls/overrides?originFile=&lineNo=`
resets one site (omit both = all), clearing the manual edges and un-hiding the placeholder inline (no
refresh). A target that is not a real `MODULE` is rejected `400 UNKNOWN_TARGET`. **Bug B fix:** the
`callees` items now carry `unresolved` (mirroring what `graph` already exposed), so an unresolved
dynamic target is machine-distinguishable from a resolved one without inspecting `sourceFile`. Since item 200 the apply
and the reset also rebuild the calling
modules' derived `CALLS_MODULE` edges in the same transaction, so `reaches` and `field-flow` honour
a pinned target without a refresh, like `callees` and `call-tree` already did.
**Dispatch guards: read `guards` — it is the only complete condition (item 72).** A `dispatch-table` row's
`guardField`/`guardValue`/`guardValues` describe the **innermost** `DECIDE` only. Natural nests
value-`DECIDE`s inside each other (824 of upms's 7,729 blocks, 178 modules), and then the innermost guard
is just one conjunct: in `VMULTMN4`, the row for `YTABLMA0.TX-TABLA` reports `#FIELD-NAME = 'TX-TABLA'`,
but the assignment also requires `#SHORT-VIEW = 'TABL'`. `guards` is the full chain — `[{field, values}]`,
**outermost first, joined by AND**, each link's `values` joined by OR. Reading only the legacy fields
**over-generalises**: port that to Java and you get a branch firing where Natural never would. For an
unnested `DECIDE` the chain has one link and says the same as the legacy fields.
- **Still incomplete for `NONE`/`ANY` branches (item 73):** an assignment in a `NONE` branch is reported
under its enclosing chain alone, but its real condition is "enclosing guard **AND NOT** any sibling
`VALUE`" — a negation a chain of equalities cannot express. `guards` is strictly better than the legacy
fields, not a total answer.
**A dispatch row's `lineNo` belongs to `sourceFile`, not to the module (item 122).** `dispatch-table`
rows now carry the same provenance quartet as `callees`/`db-accesses`/`workfile-accesses`/`functions`:
`sourceFile`, `viaCopycode`, `includedAt`, `includePath`. Resolve `lineNo` **against `sourceFile`** —
when `viaCopycode` is non-null the assignment is written in that copycode and `lineNo` is a line of the
`.cpy`, while `includedAt` is the `INCLUDE` line in the module. Before item 122 the row carried only
`lineNo`, so following it against the module file landed somewhere arbitrary: 26 of `VCOMIN50`'s 44 rows
reported line 18 or 20, which in that module is a change-history comment; the real sites are
`ISICINDE.cpy:18` and `ISICINDI.cpy:20`. Note the guard chain may **span** the include boundary — the
outer `DECIDE` in the module, the inner one in the copycode — so a row can have a multi-link `guards`
chain whose links live in different files.
**Within one guard: prefer `guardValues` over `guardValue` (item 64).** `dispatch-table` rows carry both.
`guardValue` is **lossy** and kept only for compatibility: it comma-joins the branch's `VALUE` literals,
which drops blank alternatives entirely and, for a multi-alternative branch, yields a synthetic string the
guarded field never equals (`"A1, A2"`). `guardValues` is the faithful list — every alternative in source
order, blanks included — so `VALUE 'GENAGREE-WOUT-SP', ' '` reports `["GENAGREE-WOUT-SP", " "]`, recording
that a **blank** guard field also routes into that branch. When reasoning about routing (or porting a
`DECIDE` to Java), read `guardValues`; `guardValue` will silently under-report the branch's conditions.
**Text that is not code never yields a call (items 61 & 63).** The `CALLNAT`/`CALLNAT_DYNAMIC` patterns
are unanchored (a `CALLNAT` may legally appear mid-line), so both ingest tiers first neutralise non-code
text: a full-line `*` comment is skipped, a trailing `/* …` is stripped (item 61), and a match whose
**keyword** falls inside a quoted string literal is rejected (item 63). A real `CALLNAT 'MOD'` is
unaffected — its keyword sits outside the quotes. This matters for trusting `callers`/`callees`/
`call-tree`: before item 63, prose such as `WRITE(#MSG) 'NACH CALLNAT ISINGEAG:'` or `#ERR-TYPE :=
'Callnat USIA008N'` fabricated a `CALLNAT_DYNAMIC` edge to the **real** module of that name, so a mere log
message appeared as a genuine call — and, because a real module existed, it was *not* flagged
`unresolved` and could not be reaped by the item-62 cleanup. If you query a graph ingested before
2026-07-16, re-ingest (`ac refresh`) before trusting call-graph edges into modules that are also
mentioned in log/error text.
**Calls made through a copycode's arguments (items 120/121/123).** In Natural the target of a
`CALLNAT` is often not written at the call site at all: a copycode receives the module name as a
positional `INCLUDE` argument and issues `CALLNAT &2&`. Three defects in that argument path — arguments
continued on the next line, the doubled-quote escape `'''X'''`, and the double-quote delimiter
`'"X"'` — meant such a call produced **no edge and no unresolved-dynamic-call entry**, so `callees`,
`callers`, `call-tree` and `reaches` agreed on an answer that was simply absent, with nothing saying
"not analysed". This hit the browse/access layer hardest, because that is where the idiom lives:
`YCARPBN1` and `YPOLIBN1` reported **0** callers each. Fixed 2026-08-07; re-ingest recovered 2672
copycode-derived call pairs (+49%) in `upms` with none lost. **A graph ingested before 2026-08-07
under-reports Natural callers/callees, and does so silently — re-ingest before concluding a Natural
module is unused.** A copycode parameter that genuinely has no argument now surfaces in
`/dynamic-calls/unresolved` as `&n&` rather than being dropped, so "not analysable" is visible.
**A call edge no longer outlives the call it was parsed from (item 124).** Until 2026-08-09 a refresh
only *added* the corrected call and left the old one in place, because an edge is reaped only when one
of its endpoints is — and a parser fix changes neither (the calling subroutine is unchanged, the old
target is a never-swept placeholder). So `callees`/`callers` could report a call that no source line
makes, flagged `unresolved: true` and indistinguishable from a genuine unresolved dynamic call. A
re-parsed Natural file's call edges are now reaped before the fresh ones are merged, and a call-target
placeholder left with no callers is deleted. **Two consequences for a graph ingested before
2026-08-09:** an `unresolved: true` callee may be an artefact of an already-fixed parser bug rather
than a real dynamic call, and `search/identifier` may list module names that exist nowhere in the
source. Both clear on the next deep refresh. Note the reap deliberately spares `CALLNAT_DYNAMIC` edges
onto *real* modules — those are the dynamic-call resolvers' output, not the parser's.
## LoC / SLoC metrics (item 46)
Every file-level node (a `MODULE` program/class, or a `DATA_STRUCTURE` for a Natural `.lda`/`.pda`
data area) is stamped at ingest with two deterministic line metrics:
- **`loc`** — physical lines of the file (language-independent; a trailing newline adds no phantom
line).
- **`sloc`** — source lines of code: non-blank, non-comment lines, computed **per language**. Natural
drops full-line `*`/`**`/`/*` and inline `/*` comments; Java drops `//` and `/* … */` blocks while
keeping those tokens when they appear inside string literals.
Both the cheap Tier-1 coarse scan and the deep Tier-2 parse use the *same* per-language counter, so a
module's `loc`/`sloc` are identical at any ingest depth — you can sum them to get exact,
reproducible project totals.
Where to read them:
- `GET /modules` — `loc`/`sloc` on each row.
- `GET /modules/{name}/context` — `loc`/`sloc` on the module.
- `GET /nodes/{id}` — `loc`/`sloc` in the node's raw properties.
- `GET /loc` (`ac loc`) — the rollup: a per-language breakdown (`fileCount`, `loc`,
`sloc`) plus a project-wide total, optionally narrowed by `?language=` / `?sourceFile=`. Each source
file is counted once even when it yields several nodes (Java inner classes, Natural inline groups).
`null` metrics mean the node predates item 46 — re-ingest (`refresh`) to backfill.
### Generated vs. user-exit split (item 47)
A project can be created with a source **`language`** (required at creation; an attribute only — ingest
still classifies files by extension) and a **`generatedDir`**/**`userExitDir`** pair (directory *names*,
matched as path components like `excludeDirs`; both or neither). Generated modules already contain their
hand-written user-exit twin inline, so at ingest a module under `generatedDir` whose name also occurs
under `userExitDir` is annotated with that twin's LoC/SLoC (`userExitLoc`/`userExitSloc`). User-exit
files are **not** ingested as standalone modules (they would collide by name) — the walk skips
`userExitDir`.
**Consequence for all non-LoC analysis:** the `generatedDir` copy is the *canonical, sole* module for
every structural query (call graph, DB access, functions, data structures, identifiers, dataflow,
dispatch table). `userExitDir` exists **only** to compute the generated-vs-manually-written LoC split
below; it never contributes nodes/edges. So when verifying an API response against source for a Natural
module, always read the `generatedDir` file (e.g. `generated_src/subprogram/WGEAGB0S.nat`), not the
`user_exit` fragment.
`GET /loc` (`ac loc`) then reports, per language row **and** in the project total:
- **`loc`/`sloc`** — the **total** (generated, which already includes the user exits).
- **`userExitLoc`/`userExitSloc`** — the sum of the annotated user-exit twins (the hand-written part).
- **`generatedExclusiveLoc`/`generatedExclusiveSloc`** — total − user-exit, clamped ≥0 per file (the
purely generated part).
All three are `0` for projects without a generated/user-exit split. Create with
`ac project create <name> <root> -l natural -g generated_src -u user_exit`, or add the split to an
existing project via `ac project update <name> -g generated_src -u user_exit`.
## Java `DB_ACCESS` in a project without JPA entities (item 140, 2026-08-27)
A `DB_TABLE` node is only ever created from a JPA `@Entity` or a Panache active-record class. In a
Java project that has none, **no** `DB_ACCESS` candidate can resolve — and the parser's candidate
heuristic is a deliberate over-approximation: its read gate admits *any* static receiver whose method
starts with `get`/`find`/`read`/`list`/… , so `UserContext.getCurrent()` and
`TextUtils.getColumn(line, 0, 8)` become candidates. On `app` that produced **2219** `DB_ACCESS`
nodes in a codebase with no database access whatsoever.
`db-accesses` never showed them (it joins the table with a plain `MATCH`), but **`sql-statements`
did**: it joins with `OPTIONAL MATCH`, so unresolved candidates came back as rows with
`"table": null` — 76 of them on a single `app` module.
Since 2026-08-27 the enrichment step `reap-java-db-access-without-tables` deletes every Java
`DB_ACCESS` of a project that holds no `DB_TABLE`. For such a project `sql-statements` is now empty
instead of noisy. Three things to know:
- **The gate is project-level, not per node.** One entity anywhere in the project switches the reaper
off, and the unresolved candidates stay. `pur` (2014 of 3951 unresolved) and `ac` (221 of 335) are
unaffected, and still return `table: null` rows. Treat a `sql-statements` row whose `table` is
`null` as unverified, in any project that has tables.
- **Natural is untouched.** A Natural `DB_ACCESS` comes from a literal `READ`/`FIND`/`STORE` and is a
real access whether or not its view resolved (`upms`: 14302 of 14303 resolve).
- **Recovering from it needs a full refresh.** If such a project later gains its first entity, the
reaped nodes only come back for files that are actually re-parsed — `changedOnly` will not restore
them, `refresh` without it (or `recreate`) will.
## `?depth=` means module hops (item 65)
On `db-accesses` / `sql-statements` (and the `?module=` scope of `variables/{name}/reads|writes`),
`depth=N` means **N module calls away** — the same unit `/modules/{name}/graph?depth=` and `call-tree` neighbours
use. `depth=1` = the modules this one directly `CALLNAT`s, regardless of how deeply the calling
statement sits inside subroutines.
Before item 65 these endpoints bounded the traversal on raw `CALLS` edges. A `CALLS` edge starts at the
*statement* making the call, not at the `MODULE` node, so the traversal also stepped through internal
`PERFORM` jumps and `depth` measured **statement nesting**, not dependency distance. Concretely:
`WGEAGB0S` reached `YGEAGBNH`'s tables through two module calls, but the raw path is 5 edges
(`WGEAGB0S → GET-DATA → GET-MAIN-DATA → BGEAGFN0 → READ-FILE → YGEAGBNH`), so `db-accesses?depth=2`
returned `[]` — reading as "no DB access" — and only `depth=5` was truthful.
**If you scripted a depth workaround** (a deliberately large `depth` to compensate), drop it: `depth`
is now the value you'd naturally expect, and inflated values just widen the result set.
> **`call-tree`'s `depth` column is still raw-hop based** and mixes internal subroutines into the tree:
> a direct dependency called from the main body shows `depth=1` while one called two subroutines deep
> shows `depth=3`. Use the ego graph (`/modules/{name}/graph`) when you need module-level distance. Tracked as an open
> roadmap item.
## Framework-mediated DB access via `INCLUDE` macros (item 44)
Natural's generic table-access framework hides a `CALLNAT` inside a copycode member, invoked with a
statement-level macro:
```natural
INCLUDE YFRAMGC0 'YFRAML01.C-MOD-GET' '"CO-TABLA-CO-ELEMENTO-ALT"' '"YELEMGN0"' 'YELEMKEY' 'YELEMREC'
```
The `CALLNAT` to the generic accessor (`YELEMGN0`) lives in the copycode, not in the including module,
so before item 44 both `callees` and `db-accesses` were empty for such modules. The parser now
recognises the framework macro and emits a `CALLS` edge to the accessor named in the macro arguments
(de-quoted; e.g. `'"YELEMGN0"'` → `YELEMGN0`), tagged with **`edgeKind = INCLUDE_MACRO`** on
`callers`/`callees`. Because the edge is a normal `CALLS`, the accessor's own table access surfaces
transitively: `GET /modules/{name}/db-accesses?depth=N` reports the table with `via` = the accessor
module. The recognised macros and which argument names the accessor are described declaratively in
`FrameworkMacros` (ac-parser-natural). Scope: the targeted recogniser only — general `.nsc` copycode
expansion is still open.
**`db-accesses?depth=N` is a superset of `db-accesses` (item 93).** Besides `READS`/`WRITES` it also
returns the `mode: "DECLARES"` rows — a Java entity's own `MAPS_TO` table and a repository's
`repositoryEntity` table (item 32) — for every module in the closure, with `via` naming the declaring
module. Before item 93 the transitive query carried only the `READS`/`WRITES` branch, so asking the
*same* module with `depth` dropped its declared table and a Java caller's transitive `db-accesses` came
back empty although the entity it persists through maps to a real table.
**Natural view aliases are resolved to the underlying table (item 95).** A Natural DML statement names a
*view variable* (`1 VDB2-VERSIS_LITERALES VIEW OF VERSVW_LITERALES`), not the DDM. `db-accesses` reports
the **table** — `FIND VDB2-VERSIS_LITERALES`, `FIND NUMBER NEXT-VIEW` and `STORE VDB2-VERSIS_LITERALES`
in `YLITEMN0` all come back as `VERSVW_LITERALES`, matching the SQL `SELECT … FROM` rows in the same
module. Before item 95 the alias itself was the reported name, which (a) split one table across several
names, (b) made the generator's boilerplate alias `NEXT-VIEW` a single node shared by 11 modules meaning
11 different tables, and (c) hid every `VERSVW_LOGFILE` write behind 11 `VDB2-*-VLOG` aliases. Table
names are upper-cased (Natural is case-insensitive).
**…including aliases declared in a `USING` data area (item 98).** A view is often declared not in the
module but in a `LOCAL USING` area, in the data-area *export* form (`V 1VDB2-VERSIS_GENAGREE
VERSVW_GENAGREE …` — no `VIEW OF` text). Those resolve too: `YGEAGBNH`'s
`FIND (1) VDB2-VERSIS_GENAGREE` reports `VERSVW_GENAGREE`. Resolution is scoped to each module's own
`USING` set, never by name — alias names are boilerplate, and `NEXT-VIEW` alone is declared over 100
different tables in `upms`. A module whose `USING` areas give two different tables for one alias is left
unresolved rather than guessed.
**Natural `UPDATE(<label>.)` / `DELETE(<label>.)` count as writes (item 96).** These act on the current
record of the labelled `FIND`/`READ` loop, and are reported as `WRITES` on that loop's table. This is
what makes the `Y****MN0` access layer's update/delete path visible: `YLITEMN0` reports `WRITES
VERSVW_LITERALES` at the `STORE` **and** at `UPDATE(HOLD-PRIME.)` / `DELETE(HOLD-PRIME.)`, where before
item 96 it reported only the `STORE` — reading, wrongly, as an insert-only layer. A reference that
resolves to no labelled loop (an unknown label, or the numeric source-line form) records nothing rather
than guessing a table.
**`call-tree`/`graph` agree with `callees` about overridden dynamic calls (item 97).** A manual
dynamic-call override hides the placeholder marker rather than deleting it. All read paths now filter it,
so a pinned `CALLNAT <var>` shows the real target and never the variable name. Everything driven by the
call-tree BFS — `graph`, `db-accesses?depth=N`, `sql-statements?depth=N` — inherits this.
**`callers` on a dynamically-called module is an over-approximation, and says so.** A Natural web-service
module is reached by `CALLNAT #WIF`, resolved by naming pattern: `W-LST-N0.nat:362` alone resolves to 29
`W****B*S`/`W****X*S` targets, so `WGEAGB0S` lists `W-LST-N0` and `W-MNT-N0` as callers. The rows are
tagged `edgeKind: "CALLNAT_DYNAMIC"` — treat those as *may-call*, not *does-call*, and check
`dynamic-calls/overrides` / `dynamic-calls/unresolved` when the distinction matters.
## XML payload / interface schema (item 45)
Natural XML wrapper subprograms build a wire payload by mapping data-area fields to XML tags via the
`ADD-XML-LINE` idiom (`#W-TAG := '<tag>'` / `#W-VALUE := <field>` / `PERFORM ADD-XML-LINE`, where the
subroutine `COMPRESS`es `'<' #W-TAG '>' #W-VALUE`). The deep parser extracts that contract as
`PAYLOAD_FIELD` nodes and exposes it:
- `GET /modules/{name}/payload` (`ac payload <module>`) → an array of
`{tag, field, direction, lineNo, sourceFile}` triples. `direction` is `REQUEST` for an emitted
(outbound) field. `field` is the unqualified payload field name (`WXMLIN.P-COD-USUARIO` →
`P-COD-USUARIO`). **`sourceFile`** is the file `lineNo` refers to — the module's own file for
`source=IDIOM`, or the interface PDA's file for `source=PDA` (so a caller opens the right file at the
line, not the module at a stray line).
- **`source=IDIOM`** (item 45): extracted from a static `ADD-XML-LINE` emit sequence with literal tags.
- **`source=PDA`** (item 46b): the module is a *generic*, runtime-driven serializer (it calls the
`YFRAMN07` tag-builder or has an `ADD-XML-LINE`/`ADD-XML-ACT` subroutine) with **no** static tag list
in its source — real production wrappers like `WNAUTD0S` are this shape. The contract is then derived
from the module's `PARAMETER USING` **interface PDA**: each field is a payload field, the wire tag is
the field name with the framework's `EXAMINE … '#' REPLACE '_'` normalisation applied (`#`→`_`),
direction `REQUEST`. Idiom fields take precedence when both exist.
The static idiom also handles **derived tags** (`EXAMINE #W-TAG FOR '#' REPLACE '_'` → `#`→`_`) and
**both directions**: an `ADD-XML-LINE`-style emit sub is `REQUEST`; a `GET-XML-LINE`-style parse sub
(with the reverse `field := #W-VALUE` binding) is `RESPONSE`.
Empty for modules that neither use the idiom nor are a flagged XML wrapper (or are only
coarse-ingested).
## Copycode (`.cpy`) expansion (item 46a)
Natural `INCLUDE <member> <args>` is a compile-time macro: the copycode body is spliced into the
including module (with positional `&1&…` substitution), so a copycode's `CALLNAT`/`PERFORM`, DB access
and dataflow live in the copycode, not the module. The deep and coarse parsers now expand statement-level
copycode includes before parsing, so those constructs surface on the including module — e.g. a `READ`
or `CALLNAT` that only exists in a `.cpy` shows up in the host's `db-accesses`/`callees`.
- Copycode-origin nodes/edges report the **real `.cpy` file + line** (so navigation lands in the
copycode), and carry `viaCopycode=<member>` + `includedAt=<host line>`; host statements keep their own
file + line (line numbers are remapped after the splice, never shifted).
- **A line number alone is not a location (item 66).** Because of the above, one module's calls and field
accesses come from more than one file, and host and copycode lines are freely mixed — so always read the
line together with the file the endpoint gives it:
- `callers`/`callees`/`functions/{f}/callers` return **`sites: [{lineNo, callSiteFileIndex, viaCopycode,
includedAt, includePath}]`** (not a bare `lineNos` array). `callSiteFileIndex` indexes `sourceFiles` and
is the file the call is *written in*; the entry's own `sourceFileIndex` is a different thing — the file
the named module/function is *defined* in. `viaCopycode`/`includedAt`/`includePath` are set only for
copycode sites.
- `variables/{name}/reads|writes` return `sourceFile` = the file `lineNo` is in (the `.cpy` for a
copycode access), plus `viaCopycode` + `includedAt`.
**Item 109:** a `writes` row also carries `assignedValue` — the right-hand side *as written*:
`'WREQUD0S'` (literal), `*PROGRAM` (system variable), `#DISPLAY(1)` (indexed), `#SELECTED-KEY.NUM`
(qualified reference) — plus `assignedSubstrPos`/`assignedSubstrLen` for a `SUBSTR(...)` target
(item 83). It is deliberately not normalized to literals: that would drop most write sites. This is
what answers "which values does this module put into field X" without opening the source, the
dispatcher question item 108 keeps running into.
**`assignedValue` is always `null` for Java** — the Java parser does not capture the right-hand
side. Read it as "not captured for this language", not as "nothing is assigned". Natural carries it
on ~97% of its write edges. On a `reads` row it is always `null` by definition.
Before item 66 the copycode's line was paired with the host's file: `#W-OPTIONS` writes in `WGEAGB0S`
were reported at `WGEAGB0S.nat:18/20/22`, which is its generated comment banner — the writes are really
`ISICINDI.cpy:18/20/22`. Use `includedAt` when you want the spot in the host module instead.
- **`includedAt` is always a line in the module's own file, and `includePath` shows the whole chain
(item 104).** Natural `INCLUDE` nests, often through a positional argument
(`INCLUDE USIX050C 'YFRAMMC1'` → `INCLUDE &1&` → `INCLUDE YFRAMC01`), and only the *innermost* member
is named by `viaCopycode`. `includedAt` used to be the `INCLUDE` line in the *enclosing* `.cpy` — a
file the response never named — so `ISI173N0 → YFRAMN04` reported line 27, which is a comment in
`ISI173N0.nat` and in truth line 27 of `YFRAMMC1.cpy`. Now `includedAt` is the host-module line (232),
and `includePath: [{sourceFile, lineNo}, …]` lists every hop, host first, innermost last (empty for a
direct statement, one entry for a one-level include).
- `db-accesses` / `workfile-accesses` return **`sites: [{lineNo, sourceFile, viaCopycode, includedAt,
includePath}]`**
alongside the (kept, backward-compatible) `lineNos` array — one entry per statement, each tying its
line to the file it truly lives in. `sql-statements` gains **`sourceFile`** + **`viaCopycode`** on each
statement (its `startLine`/`endLine` are lines *in `sourceFile`*). Before this, a DB/work-file access
written in an `INCLUDE`d copycode reached the API as a bare copycode-local `lineNo` with nothing to
attribute it to — e.g. the DB2 sequence read `SELECT … FROM SYSIBM-SYSDUMMY1` lives in `USIX043C.cpy`
at lines 31/39/45/51/57, but `db-accesses` for the 9 including modules (YAPRFMN0, YUGRPMN0, …) reported
those as bare line numbers that land on the host's own comment/`DEFINE DATA` lines. The `sites` file
context is the same fix item 66 applied to `variables/reads|writes` and `callees`.
- **A copycode's nodes belong to the including module, not to the copycode (item 75-B, 2026-08-22).**
Until now every module that included a `.cpy` shared *one* set of nodes for its body. That is no longer
so: a copycode-resident node is keyed per including module (`ownerModule`), so what an agent sees changes
in one visible way — **counts go up, and they are now per-module**. A `READ` written in a copycode that
20 modules include is 20 access nodes, one per module, instead of one shared node; the same holds for a
`DEFINE SUBROUTINE` in a `.cpy` and for its control-flow statements. Read it as "each of these modules
really does perform this access", which is what the API always claimed but could not previously
represent. `search/identifier` for a name defined in a widely-included copycode therefore returns one hit
per including module — filter/group by `sourceFile` + the module you care about rather than expecting a
single row. Modules and DB tables are deliberately **not** per-module: a module declared inside a
copycode (`ZDTSTBP6` in `ZDTSTBC6.cpy`) and every `DB_TABLE` stay shared, so module lookups are unchanged.
**Item 75-C (same day) takes this one step further: identity is per *expansion site*, not per module.**
A copycode included several times by the same module (`JX0031N0.nat` includes `YFRAMBC0` 16 times) now
yields one set of nodes *per include site*, keyed by `includePath`. So counts rise again for those
modules, and — the point of the change — a copycode that opens a block it does not close no longer
collects every site's nesting into one node. `includedAt` alone does **not** identify a site (item 104
makes it the host's INCLUDE line at every nesting level, and `VPARTC02.cpy` includes `L4NLOGIC` 136 times
behind a single host line); use `includePath` when you need to tell two expansions apart.
Measured on `upms` after the recreate: copycode-resident nodes 27,551 -> 51,895 (project total +5.0%),
spread over 19,565 distinct owners. Endpoint latency on the heaviest module (`JX0030N0.nat`, 91 include
sites) is unaffected: `digest` 0.83 s, `context` 0.25 s, `graph` 0.19 s.
- **The same line number can legitimately appear twice (item 69).** A host statement on line 10 and a
copycode statement on line 10 are two different statements, and both are returned — as separate entries
differing only in their file. Until item 69 the graph could not hold both: an edge was identified by
`(source, target, type, lineNo)` with no file, so the second one **overwrote** the first and a real
access was missing from every answer. Treat `(file, lineNo)` as the identity of a site, never `lineNo`.
- **`call-tree`'s `depth` counts module hops (item 67).** `depth` is how many **module boundaries** the
shortest call path crosses, not raw `CALLS` edges — a `CALLNAT` made from two subroutines deep is still
one hop. The root module's own subroutines are therefore **depth 0**. Measured on `upms`: `WGEAGB0S`'s
seven direct dependencies used to report depth 1..3 (`BGEAGFN0` was 3); all seven now report 1.
- **Results at a given `depth` are larger than before.** A subroutine of a module within `depth` hops is
now inside the bound, because it crosses no further boundary. Previously `call-tree?depth=1` could hide
a `DEFINE SUBROUTINE` of *the very module you asked about*, just because it was `PERFORM`ed from
another subroutine (raw depth 2) — that is the same bug seen from the inside.
- `call-tree` also returns **`truncated`**. `true` means the *intra-module* subroutine walk stopped at
its raw-hop budget, so some `FUNCTION` items may be missing — **not** that your `depth` was exceeded
(that is a normal, complete answer). It is conservative and can be `true` for a complete result. Tune
via `agenticcode.call-tree.internal-budget` (default 20; the deepest internal chain observed in `upms`
is 9).
- **Since item 94 the budget cannot hide a module.** `MODULE` rows come from the same module-hop BFS
that `db-accesses`/`sql-statements` use, so a callee one hop away is always listed even when its call
site sits behind a long internal `PERFORM` chain (before item 94 it was dropped, and `call-tree` then
contradicted `db-accesses`). This also removed the path enumeration that made
`call-tree?followWiring=true` time out on Java projects at `depth ≥ 2`; `followWiring` is now usable
at full depth.
- **`field-flow`'s `depth` counts module hops (item 68).** Like `db-accesses`/`sql-statements` (item 65),
`variables/{name}/field-flow?depth=N` now means "up to N **module** calls apart", not N raw `CALLS`
edges. Before item 68 a consumer called from inside a subroutine sat several raw hops away and was
dropped at `depth=1`, so the endpoint answered "nothing downstream consumes this field" — read that
answer with suspicion on any graph ingested before this change.
- **`field-flow` no longer fabricates flows between same-named fields (item 77).** A bare field reference
is resolved against **the referencing module's** own `USING` includes. It used to be resolved
project-wide: an unresolved bare field is one shared node per `(name, project)`, and the resolver
aggregated over all owning modules at once, so (a) two modules including *different* data areas that
both declare the name left both unresolved on the shared node, and (b) a module with no matching include
was redirected onto another module's field. Either way the two modules ended up on one node, and
`field-flow` — which pairs a producer with a consumer only when both touch the *same* node — reported a
dataflow between modules that share nothing but a field name. In `upms`: 38 + 28 placeholders affected
(199 module-field pairs). **Read any pre-item-77 `field-flow` result for a common field name with
suspicion**, and note the answer only changes after a *deep* re-ingest, since resolution runs there.
`reads`/`writes` are unaffected — they match every node with the name and report only the accessing
side, so they never distinguished the targets in the first place.
*Residue (item 76):* a bare field shared via copycode (132 of 18,539 source nodes in `upms`) is still one
node for several modules; per-module identity is a schema change, not a query fix.
- **Copycode provenance survives field resolution (item 70).** `viaCopycode`/`includedAt` are now kept for
fields addressed by *qualified* name (`MYLDA.Q-FIELD`, i.e. a field of a `LOCAL USING` data area) as
well as bare ones. Before item 70 only bare references kept it; qualified ones silently came back with
`viaCopycode: null` and the host file, i.e. they looked exactly like host statements.
- Excluded from expansion: framework macros (handled by the item-44 targeted recogniser), data-area
`USING` includes, unknown members, and any copycode that declares `DEFINE DATA`. Recursion is
cycle-guarded. Copycodes (`.cpy`) are not standalone modules — they enter the graph only through the
including module.
- **Staleness caveat:** the item-41/43 hash check hashes the *host* file, so auto-invalidation triggers
on a change to the host — but a change to an included `.cpy` **alone** (host unchanged) is not
detected; re-ingest the host (`refresh/{host}`) to pick it up.
## Global Data Areas (`.gda`) (item 46c)
`.gda` files are now ingested as `DATA_STRUCTURE`s like `.lda`/`.pda`, and `DEFINE DATA GLOBAL USING
<gda>` resolves to them (the `INCLUDE`/`USING` recogniser now accepts `GLOBAL`, not just
`PARAMETER`/`LOCAL`).
## Deep-ingest: now automatic (lazy Tier-2)
Field-level endpoints (`flow-forward`, `flow-backward`, `field-flow`) and
cross-module dynamic `CALLNAT` resolution need a **per-module deep ingest**, not
just a whole-root refresh. This deep ingest is now **triggered
automatically on demand**: calling a field-level endpoint for a module that is
only `CALL_GRAPH`-ingested runs a scoped deep ingest of that module (and its
dependency tree) transparently, then returns the resolved result — no `409`,
no manual `POST /refresh/{name}` step. The first such call to a cold module is
therefore slower (it walks the root and parses the program tree); subsequent
calls hit the already-`FULL` graph.
The deep ingest is best-effort: if the module cannot be resolved to a source
file, the endpoint still falls back to the `409 NOT_DEEPLY_INGESTED` /
`NOT_INGESTED` hint with a `nextAction` rather than a misleading empty result.
**Flow path-ingest (auto, cross-module fixpoint).** `flow-forward`,
`flow-backward`, and `field-flow` go one step further than the single start-module
deep ingest: after deep-ingesting the start module they run an *ingest-and-re-traverse
fixpoint*. Each round deep-ingests the **frontier** — the modules the trace surfaced
*together with their direct callee modules* — in one scope, then re-traverses. This is
what lets a dataflow trace cross into a **dynamically-dispatched** callee (`CALLNAT
PGM-VAR`): that callee is not a static dependency of the start module, so it is only
pulled in and linked (`caller.arg → callee.param`) once a round resolves the dynamic
`CALLS` edge and ingests the target. The loop is bounded by
`agenticcode.deep-ingest.flow-rounds` (default 3) and the per-round
`agenticcode.deep-ingest.fanout-nodes` budget, and stops early (fixpoint) as soon as a
round pulls in nothing new — so on an already-deep graph a flow query costs one
traversal plus one cheap frontier check, no re-run.
**Fan-out warm (auto, on the result set).** The fan-out / traversal queries
`callers`, `/search/identifier`, and `call-tree` also auto-deep-ingest — but on
the **set of modules their result surfaced**, not a single named module. Each
runs against the graph as-is, deep-ingests the surfaced modules (blocking,
bounded by the fan-out node budget `agenticcode.deep-ingest.fanout-nodes`,
default 50), and — only if that warm actually deepened something — re-runs so
the response reflects newly-resolved dynamic dispatch (e.g. a `call-tree` grows
to include a dynamically-dispatched callee once the surfaced program is deep).
When everything is already `FULL` (or the warm resolves nothing) the query
returns its first result with no redundant re-run. Note `callers` warms the
*already-surfaced* callers, so it improves downstream precision but cannot
reveal a caller that was invisible at the coarse (call-graph) tier. Other
module-level endpoints (context, callees, db-accesses) work regardless of
ingest depth and do not trigger a deep ingest.
**Bounded fan-out.** A by-name deep ingest walks the transitive dependency tree
breadth-first, bounded by `maxDepth` (hops from the named module, default 5,
ceiling 20) and `maxNodes` (files, default 300). When a bound is hit the walk
stops early and the ingest response carries a `truncation` object
(`{reason: DEPTH|NODES|NODES_AND_DEPTH, maxDepth, maxNodes, hint}`) — the modules
actually reached are marked `FULL`, the remainder stays as it was. Raise the
limits on an explicit module refresh to pull in more:
`POST /refresh/{name}?maxDepth=&maxNodes=`, or CLI `ac refresh <name> --max-depth --max-nodes`. Auto-triggered ingests
use the
server defaults; if a field-level query returns partial data because the target's
deep ingest truncated, re-run the explicit refresh with higher limits. (The
auto-trigger does not yet accept per-query limit overrides.)
**Durable ingest status + coalescing (item 36).** Each real `MODULE` node carries a
durable `ingestStatus` lifecycle — `NOT_INGESTED` (only its call graph is in the
graph) → `INGESTING` (a deep ingest is in flight) → `INGESTED` (deeply ingested,
`ingestDepth = FULL`) — separate from `ingestDepth`. When two calls trigger the same
module's deep ingest at once they **coalesce** rather than both ingesting: within one
process an in-process lock serialises them; across processes/restarts a best-effort DB
claim marks the module `INGESTING` and a loser waits for the winner to reach `FULL`
(re-claiming if the claim is released or goes stale after
`agenticcode.deep-ingest.ingesting-ttl-seconds`, default 1800; wait bounded by
`claim-wait-seconds`, default 120). A crash mid-ingest leaves the module re-triggerable
(it never reached `FULL`), and the stale `INGESTING` is reclaimed on the next call.
`GET /nodes/{id}` / `/nodes/{id}/source` expose `ingestStatus`/`ingestStatusAt` on the module node.
**Warm concurrency cap (item 37).** All auto deep-ingest/warm work (by-name, fan-out,
and flow-frontier) shares a global permit pool
(`agenticcode.deep-ingest.max-concurrent-warms`, default 2), so a burst of queries
cannot spawn unbounded parallel parses/Neo4j writes. A permit is acquired only around
the actual ingest; if none frees up within
`agenticcode.deep-ingest.warm-acquire-timeout-seconds` (default 10) the warm is skipped
and the query returns its **Tier-1** (coarse) answer immediately rather than blocking —
so under sustained load a query may transiently return shallower data; retry once load
subsides, or force it with an explicit `POST /refresh/{name}`.
## OpenAPI contract & CORS (items 48/50)
The server now ships an OpenAPI 3 spec (`quarkus-smallrye-openapi`): the machine
contract the web-UI TypeScript client is generated against. All REST endpoints
carry `@APIResponse`/`@Schema` annotations, so response bodies are typed in the
spec even though the JAX-RS methods return raw `Response`. Access it at:
- `GET /q/openapi` — YAML (or `Accept: application/json` for JSON)
- `GET /q/swagger-ui` — interactive UI (dev)
CORS is enabled (`quarkus.http.cors.enabled=true`) and restricted to the UI's dev
origins (`http://localhost:5173`, `http://localhost:4173`) — extend the
`quarkus.http.cors.origins` list per deployment; never ship a wildcard.
## Endpoint quick reference
| Endpoint | Use for |
|----------------------------------------------------------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `GET /modules?sourceFile=&moduleKind=&extends=` | List/filter modules; map a source file to its module name(s). Each row carries `loc`/`sloc` (item 46) and `ingestStatus`/`ingestDepth` (item 50) for status badges without a per-module round trip |
| `GET /loc?language=&sourceFile=` | Per-language LoC/SLoC rollup (`typescript`/`css` too, item 192) (fileCount/loc/sloc) + project total; each file counted once (item 46). For a generated/user_exit project also `userExitLoc`/`userExitSloc` + `generatedExclusiveLoc`/`generatedExclusiveSloc` (item 47) |
| `GET /modules/{name}/digest` | Tiny triage view before deciding which modules to expand |
| `GET /modules/{name}/context` | One-shot overview: functions, callers, callees, DB accesses, SQL/variable summaries (`?include=` for full lists) |
| `GET /modules/{name}/callers` \| `/callees` | Direct callers/callees incl. `EXTENDS`/`IMPLEMENTS`/`INJECTS`/`REFERENCES`. `callers` `scope`: **`external` (default)** = modules that call this one (CALLNAT/inheritance), **rolled up to the calling MODULE**: a call made from inside a subroutine/method is attributed to its owning module (never the calling `FUNCTION` node), and repeated call sites from one caller collapse to a single row whose `sites` list every line — symmetric with how `callees` anchors its source side. `internal` = the module's own subroutines' `PERFORM` wiring (function-level). The default is external-only, module-typed only, and never lists the module as its own caller (no `MODULE→MODULE` self-loop); use `scope=internal` or `/functions/{fn}/callers` for intra-module / function-level wiring. `callees` is unchanged (default lists both external CALLNAT and internal PERFORM targets) |
| `GET /modules/{name}/functions/{function}/callers` | **FUNCTION-level callers** (item 52): who `PERFORM`s (Natural) or calls (Java, TypeScript — same-module **and** cross-module, item 197) a specific subroutine/method, with call-site `lineNos`. Cross-module callers come from the module-to-module `CALLS` edge's `callerFn`/`calleeMethod`, matched by name (overloads over-approximate; a call from top-level code with no enclosing function shows only in the module-level `/callers`). Finer-grained than the module-level `/callers` (which is module→module). Same `CallRefResponse` shape. CLI `ac function-callers <module> <function>` |
| `GET /modules/{name}/call-tree?depth=` | Transitive call graph to scope a feature
| `GET /modules/{name}/reaches?target=A,B,C&direction=up\|down&depth=` | **Item 110 — "can A reach B, and how?"** Returns `{reachable, paths, truncated}` with one witness route per reached target (module names, source→target). `direction=down` (default): paths from this module to each target. `up`: paths from each target to this module. The counterpart to `call-tree`, which only walks downward and returns a closure without routes — one audit hand-rolled this as ~100 `/callers` requests. **`reachable: false` means "no path over known edges", not "no path"**: the traversal runs on resolved module calls, so a route through an unresolved dynamic `CALLNAT` (item 82) is invisible. Bounded by `depth` (item 75: the call graph has cycles). CLI `ac reaches <module> --target A,B --direction up` |
| `GET /duplicates` | **Item 114 — identities skipped at ingest** because they exist in more than one file (`{name, kind, paths}`, paths relative to the project root). These are *not* in `/modules`; asking for one by name gives `409 DUPLICATE_IDENTITY`. Their own calls are absent from the graph, so caller lists elsewhere can be short. CLI `ac duplicates` | |
| `GET /dynamic-calls/unresolved` \| `/overrides` · `POST`/`DELETE /overrides` | **Manual dynamic-`CALLNAT` overrides (item 82).** `unresolved` lists open `CALLNAT <var>` sites `{module, originFile, lineNo, variable}`; `POST /overrides {originFile, lineNo, targets[], variable?, note?}` pins a site to real module(s) (applied at once, persisted across refreshes, `400 UNKNOWN_TARGET` for a non-module); `DELETE /overrides?originFile=&lineNo=` resets one site (omit both = all) and restores the placeholder inline; `GET /overrides` lists them with an `obsolete` flag. CLI `ac dynamic-calls unresolved\|overrides\|set\|reset` |
| `GET /modules/{name}/graph?direction=&depth=&limit=` | Ego graph (item 49): bounded module-level call neighbourhood as **nodes + edges** (unlike call-tree). `direction` = `out`/`in`/`both`; `limit` caps nodes (BFS order) and sets `truncated`; unresolved targets carry `unresolved=true` + empty `sourceFile`. CLI `ac ego-graph` |
| `GET /counterparts?module=&kind=&unmatched=` (`ac counterparts`) | **Item 193:** this project's web-service calls / generated DTOs / fields with their twin in the counterpart project; `unmatched=true` = what nothing serves or mirrors yet |
| `GET /store?slice=` (`ac store`) | **Item 194:** the frontend Redux store — one row per slice (reducer key, RTK `sliceName`, `stateType`, the top-level state keys with type/optional/read/write counts, reducer count, total access sites) |
| `GET /store/{slice}/accesses?field=&mode=reads\|writes&module=` (`ac store-accesses`) | **Item 194:** who reads/writes a slice — reducers (`functionKind=reducer`, `via=reducer`) and the components/hooks/thunks selecting from it (`via` = `useAppSelector`, a wrapper hook, `getState`), with the full sub-`path` and line. Store fields also answer `variables/<slice>.<field>/reads\|writes` |
| `GET /bindings?dto=&field=&mode=reads\|writes&module=&partial=` (`ac bindings`) | **Item 195:** which component reads/writes which DTO field through the generated `Fields` path objects (`<SmartInput field={X.broker.ebene}>`), with the field's backend counterpart — `--dto Broker --field ebene --mode writes` = which page edits Java `Broker.ebene`. `data-structures/{dto}/fields` carries `boundReads`/`boundWrites` |
| `GET /theme?unused=` · `GET /theme/{token}/usages` (`ac theme`, `ac theme-usages`) | **Item 196:** the MUI theme's tokens (createTheme leaves + theme constants, value, `uses`; `declared=false` = read by the code but declared by no theme) and where one token is read (style block + CSS `property`, or plain `context`) |
| `GET /styles?module=&kind=sx\|style\|styled\|css&withLiterals=` (`ac styles`) | **Item 196:** the style inventory — every `sx`/`style`/`styled` block and CSS rule with CSS keys, hard-coded `literals` and the theme `tokens` it reads; `withLiterals=true` = what bypasses the theme |
| `GET /modules/{name}/db-accesses` \| `/sql-statements` | DB tables + mode, raw statement text (pass `?depth=` for Natural). **`db-accesses`/`workfile-accesses` return every row when no `limit` is given (item 103)** — they used to default to 50, and since the response is a bare array with no total and no `truncated` flag the cut was invisible: `WGEAGB0S?depth=10` returned 50 of 64 rows and hid 7 tables outright. An explicit `limit` is still honoured exactly. `db-accesses` items carry **`sites: [{lineNo, sourceFile, viaCopycode, includedAt}]`** (+ kept `lineNos`); `sql-statements` items carry **`sourceFile`** + **`viaCopycode`** — so a copycode-sourced access (e.g. `SELECT … FROM SYSIBM-SYSDUMMY1` in `USIX043C.cpy`) reports the `.cpy` line, not a bare number that reads as a host-file line |
| `GET /modules/{name}/workfile-accesses` | Natural **work files** (sequential/flat-file I/O — `READ`/`WRITE WORK FILE n`), the work-file analogue of `db-accesses` (item 84): `[{workFile, physicalName, mode: READS\|WRITES, recordBuffers, lineNos, sites}]`, aggregated per work-file number + mode. `sites: [{lineNo, sourceFile, viaCopycode, includedAt}]` gives each access its file context (copycode-aware), like `db-accesses`. `physicalName` comes from a `DEFINE WORK FILE n '<name>'`, else `null`. **Kept separate from `db-accesses`** — a work file is not an ADABAS/SQL table (fixes a former bug where `READ WORK FILE` created a phantom `DB_TABLE 'WORK'`). CLI `ac workfile-accesses <module>` |
| `GET /modules/{name}/data-structures` | Which copybooks/inline groups a module uses. A `USING <member>` binds by **member (file) name**, never by a level-1 record inside the file (item 100) — before that, `WGEAGB0S USING W-WIF-A2` reported `old/W-WIF-A7.pda` (whose level-1 record is a copy-pasted `1W-WIF-A2`), and a data area with several level-1 records and none named after the member (`VLAYERLA.lda`, `USIX020L.lda`) resolved to nothing at all (`sourceFile: null`, `area: UNKNOWN`, `fieldCount: 0`) although the file was ingested. One row per resolved definition, `(name, sourceFile)` (item 102) — never one row blending an arbitrary file with another definition's `fieldCount` |
| `GET /modules/{name}/payload` | Natural XML wire-payload contract: `{tag, field, direction, source, lineNo, sourceFile}` — static `ADD-XML-LINE` idiom (`source=IDIOM`, item 45) or derived from the wrapper's interface PDA (`source=PDA`, item 46b). `sourceFile` is the file `lineNo` refers to (module for IDIOM, PDA for PDA) |
| `GET /modules/{name}/comments?kind=&limit=&offset=` (`ac comments`) | **Item 141:** the module's **comment blocks** — `{text, kind, sourceFile, startLine, endLine, target, targetType, truncated}`, one row per contiguous block, ordered by line. `target`/`targetType` name the declaration the block documents: the declaration immediately below it, else the one enclosing it (so a file header banner documents the `MODULE`, a `/*` comment on a field's own line documents that field). `kind` is `JAVADOC`\|`LINE`\|`BLOCK` (Java) or `NATURAL_BANNER`\|`NATURAL_INLINE`\|`SAG` (Natural); `?kind=` filters to one. **`SAG` is excluded by default** — `**SAG` directives are generator metadata, not human notes, and would otherwise be most of the answer for every generated Natural module. Text is cut at 4 000 chars (`truncated:true`); read the file for the rest. Natural copycode comments belong to the **copycode's own module**, not to each includer. **Deep-gated:** comments come from the full parse, not the Tier-1 coarse scan, so the module is deep-ingested on demand and a still-shallow module answers `409 NOT_DEEPLY_INGESTED` rather than a misleading `[]` |
| `GET /modules/{name}/dispatch-table` | Natural `DECIDE ON VALUE OF` routing table |
| `GET /modules/{name}/functions?kind=` \| `/functions/{fn}/overrides` \| `/functions/overrides` | Method list, modifier filter (Java), subclass overrides (single/bulk). Each item carries **`sourceFile`** + **`viaCopycode`** (item 84): a Natural subroutine pulled in via `INCLUDE` reports the **copycode** file and `viaCopycode:true`, so its `startLine`/`endLine` are read as offsets into that copycode — **not** into the including module's own file (which is shorter). `viaCopycode:false` = declared inline. Always `false` for Java |
| `GET /data-structures/{name}/fields` \| `/db-tables/{name}/columns` \| `/modules/{name}/columns` | Field/column schemas for DTO/entity generation. Every field carries **`sourceFile`** (item 101). When a structure name resolves to several definitions (42 level-1 names recur across `upms` data areas), the **member root** — the definition whose file basename equals the name, i.e. what a `USING <member>` binds to — wins; **`?sourceFile=`** pins a specific one. Before item 101 the definitions were silently unioned: `W-WIF-A2` returned 15 fields, the merge of `W-WIF-A2.pda` (5) and `W-WIF-A7.pda` (10), a layout that exists nowhere |
| `GET /variables/{name}/reads` \| `/writes` \| `/flow-forward` \| `/flow-backward` \| `/field-flow` | Impact analysis and dataflow tracing |
| `GET /search/identifier` \| `/search/value` \| `/search/annotation` | Cross-project lookup by name / literal value / annotation. **All three see code only by default — an empty result is not evidence that the string is absent.** `search/value` takes **`includeComments=true`** (CLI `--include-comments`, item 141) to search comment blocks as well; those hits come back as `kind: "COMMENT"`, so a comment is never read as code. It is opt-in because a comment hit is different evidence from a literal, and folding it in silently would move every existing completeness count (item 131). `search/identifier` and `search/annotation` never match comments at all — use `includeComments`, `/modules/{name}/comments` or `/search/source` before concluding "not present" (see "Comments: reachable, but never by default" above). `search/identifier` matches the **exact** declared name but is **sigil-insensitive**: a leading Natural sigil (`#` user, `&` AIV, `+` GDA) is ignored on both sides, so `name=K-OUT-MAX` finds the declared `#K-OUT-MAX` (and vice-versa). **Item 125:** a Java **type declaration** is matched by its **short name** as well as by the fully-qualified identity the graph stores (item 117) — `name=PartnerUpdateLogic` finds `com.example.PartnerUpdateLogic`; before this it answered `[]`, which reads as "no such name". Every match carries `simpleName` and `moduleKind` (`CLASS`/`INTERFACE`/`ENUM`/`RECORD`, `PROGRAM`/`SUBPROGRAM` for Natural), both `null` for non-`MODULE` hits — so "is this name a type or a method?" needs no second call. `contains=true` (CLI `--contains`) switches to a case-insensitive **substring** match, as on `/search/value`; it was previously accepted and silently dropped. It matches the FQN too, so a package fragment also hits — filter with `type=MODULE`/`moduleKind` if that is noise. `contains` without a `name` is `400 MISSING_NAME` (a substring search for nothing is a full node dump). The substring scan is **unindexed**: it is bounded to `offset+limit` rows, so keep a `limit` on large projects. Optional **scope** filters `sourceFile=<relpath>` and `module=<name>` (item 53) narrow the match to one file / one module — use them to pinpoint a module-local declaration when a name recurs across dozens of modules (the result is otherwise paginated and the local one may fall off the page). To keep the **full cross-project list** yet still guarantee a given module's own declaration is on the first page, pass `priorityModule=<name>` instead of `module=`: it does not filter, but pins that module's matches to the front (ahead of the otherwise `sourceFile`-ordered rest) so they survive the `limit`. This is what the web UI's click-to-identify sends for the open module. CLI `ac search-identifier --module --priority-module --source-file --type --contains` accept the same filters. **Latency (item 105):** a lookup whose hits lie in a Natural data area used to take 60-75 s — every fan-out query deep-ingested the surfaced `.lda`/`.pda`, which can never reach `FULL` (a data area yields no `MODULE` node), so it was re-warmed on every call and each warm dragged a whole-project finalize behind it. Data areas are now excluded from the fan-out warm; they have no deep tier to gain |
| `GET /search/source?regex=&limit=&ignoreCase=` (`ac search-source`) | Regex **grep over module source text** (item 54): `{module, sourceFile, lineNo, line}` hits + `truncated`. Case-insensitive by default. Complements `/search/identifier` (declared names) — use for code patterns (statements, table names, literals). Sees **everything in the file, comments included**, and needs no ingest depth — so it is the fallback when a module is not deeply ingested, or when the text is something the parsers do not model. For comments specifically, prefer the graph routes added by item 141 (`/modules/{name}/comments`, `search/value?includeComments=true`), which also tell you which declaration a comment belongs to |
| `GET /nodes/{id}` | Every property of one node (when a curated DTO is missing something) |
| `GET /nodes/{id}/source` \| `/modules/{name}/source` \| `/source?file=` | Source text — **only when you have no other access to the source** (you always do in this repo, see "Reading source in this repo" above). `/modules/{name}/source` returns the **whole file** when the line range is omitted (M1), or a `[startLine,endLine]` slice when both are given. `/source?file=<relpath>` (CLI `ac file-source`) serves a file by **relative path** rather than module name — for files that aren't standalone modules, e.g. a Natural data area (PDA/LDA) USING'd by a module, whose field line numbers refer to that file. Same whole-file/range + stale-source semantics; the client-supplied path is rejected (`400 INVALID_SOURCE_FILE`) if it escapes the project root |
Full endpoint list, request params, and response field details:
`x-docs/agent-api-system-prompt.md`.
## Errors are structured JSON — always (item 136)
Every failure now answers `{ "error": ..., "code": ..., "details": {} }`, including the ones nobody
planned for: an unhandled exception is mapped to `500 INTERNAL_ERROR` with an `errorId` in `details`
that matches the stack trace in the server log (the trace itself is never in the response). Before
this, an unexpected fault escaped as a plain-text Quarkus error page with no `code` to branch on —
which is exactly the moment a client most needs a machine-readable answer. Deliberate statuses
(`PROJECT_NOT_FOUND`, `MISSING_NAME`, `STALE_SOURCE`, the runtime's own routing 404s) pass through
unchanged.
The fault that exposed this: `search/identifier` coerced `startLine`/`endLine` unconditionally, and
item 114's duplicate markers were the one kind of node created without them, so any page long enough
to reach a marker (row 487 on `ac`) died. Both halves are fixed — the markers now carry lines, and the
row mapper no longer trusts that they will.
## Verifying the API after a deploy (item 134)
After `./manage-ac.sh deploy` (or `./rebuild-and-refresh.sh`), run:
```bash
./x-scripts/verify-api.sh # defaults to project 'ac'
./x-scripts/verify-api.sh -p upms # any ingested project
AC_SERVER_URL=http://host:8787 ./x-scripts/verify-api.sh
```
It answers one question in ~10 s: *does the server that is running right now still return
plausible data over the real graph?* Exit 0 = all green, 1 = at least one check failed. Every line is
`PASS`, `FAIL` or `SKIP`; `SKIP` means the endpoint family does not apply to that project (a pure
Natural project has no `rest-endpoints`, a leaf module has no callees).
What it covers: `/api/version` and the project list; the item-130 scope headers
(`X-AC-Exclude-Dirs`, `X-AC-Ingested-At`, `X-AC-Ingest-Incomplete` — a `true` there means a refresh
was aborted and every later answer is drawn from a half-updated graph); the item-131 paging contract
on the search endpoints; per-family data plausibility; and the structured-error negative cases.
What it is **not**: a substitute for `mvn test`. The integration tests pin semantics; this pins
"the deployed thing is not obviously broken". A green run is not a quality gate. All assertions are
invariants, never fixed counts — counts move with every refresh.
## TypeScript / React projects (item 192)
A project may declare `language: typescript` (`ac project create purfe --root … --language typescript`).
Its `.ts`/`.tsx` files (not `.d.ts`) and plain `.css` files are ingested; `node_modules` and `dist`
are excluded by default. **Only a `typescript` project ingests TypeScript** — a Java project with a
bundled web UI (`ac` has `ac-ui/`) never parses it. Java and Natural files stay language-agnostic.
**Identities are paths.** A TypeScript `MODULE` is named by its root-relative path without the script
extension — `pur-r-vstamm/src/store/slices/agstammSlice` — with `simpleName` = the file stem
(`agstammSlice`), `workspace` = the first path segment, `moduleKind` = `ts`/`tsx`/`css`, and
`generated=true` + `generator` (`typescript-generator` for the Java-side EndpointGenerator output
under `generated/`, `hey-api` for `@hey-api/openapi-ts`). A CSS module keeps its extension
(`pur-ui/src/index.css`). Every `/modules/{name}/…` endpoint accepts either form (item 117), so
`ac context agstammSlice` works — until two workspaces have a file with the same stem, then use the
path.
**Two tiers, like Java/Natural.** Project creation and `refresh` without `--deep` run the **Tier-1
regex outline** in Java: module shell (`sourceHash`, `loc`/`sloc`), one `FUNCTION` per top-level
function / arrow / class (`kind` = `function` | `component` | `hook` | `thunk` | `styled` | `class`,
`exported`), one `DATA_STRUCTURE` per `interface`/`type`/`enum` (`dataType` says which), and a
`REFERENCES` edge per import to the module it resolves to (`value` = the import clause, `specifier`
= as written). npm packages are not placeholders; they are listed on the module as
`externalImports`. **A deep pass** (`refresh --deep`, `ingest` by name) runs the **Node sidecar**
(`ac-parser-typescript/sidecar/extract.mjs`, TypeScript compiler API, one whole-program run per npm
workspace, 3–5 s and ~0.5 GB each on the pur frontend) and replaces the outline with the checker's
view: exact positions, imports resolved against the file system, and **calls**:
- a callee owned by a top-level declaration of the same file → `FUNCTION -CALLS-> FUNCTION`;
- a callee in another module → `MODULE -CALLS-> MODULE` carrying `callKind` (`METHOD_CALL`, or
`CONSTRUCTOR` for `new`), `callerFn` (the calling function, absent at module level), `calleeMethod`
(the owning top-level declaration in the target), `callSyntax` (`call`/`new`/`tagged`/`jsx` — a JSX
element `<HistorieDrawer/>` is a call), and on a member call `receiver` (type of the innermost
object, e.g. `AgstammControllerEndpoint`) and `member` (`saveBroker.post`);
- calls into npm packages and the language library are **not** edges.
These are the exact properties the Java parser writes, so `callers`/`callees`/`call-tree`,
`functions/{fn}/callers` (item 52) and the placeholder rewiring work unchanged. A module that got its
Tier-2 pass carries `ingestTier=2`.
**Sidecar failure is visible, not silent.** If `node`, the script or its `node_modules/typescript`
are missing, or a workspace run fails or times out, the refresh still completes at Tier-1 for those
files and the response lists a failure with the pseudo-path `sidecar` or `sidecar:<workspace>`.
Config: `agenticcode.typescript.node`, `.sidecar-script`, `.max-heap-mb` (1024), `.timeout-seconds`
(600); the image carries node and the sidecar (`Dockerfile.jvm`), dev mode expects
`npm ci` run once in `ac-parser-typescript/sidecar/`. The project root is read-only in the
container; the sidecar reads the project's own `node_modules` for library typings and writes nothing.
**Scope of the pur frontend project.** The registered project covers the `pur-ui` and
`pur-ui-common` workspaces only; `pur-r-vstamm` and `pur-r-vbuch` are excluded via `excludeDirs`,
and the sidecar does not load an excluded workspace. Both workspaces call the `pur` backend through
the legacy generated client (`generated/endpoints.ts`, backend `pur`); the hey-api client shape is
recognised too but is not in scope.
**Resolution notes (verified on `purfe`, 2026-09-22).** An import of a workspace consumed through
its package.json `exports` resolves into its build output (`pur-ui-common/dist/x.d.ts`); the sidecar
maps that to the source twin (`pur-ui-common/src/x.ts`) so the edge lands on a real module. A bare
specifier the checker resolves to neither a file nor a package (`immer`, `redux` — transitive
dependencies the project does not list) is recorded as an external import, not a placeholder.
Transitive packages are known from `root/node_modules` (directory names), so Tier-1 treats them
as external too.
**Stale parsed edges are reaped on a deep refresh (item 198).** Every edge the parser emits is
stamped with the run's `ingestGen` at merge time; after a deep re-parse of a file, the edges from
that file's nodes whose stamp is older than the run's (the fresh parse did not re-emit them) are
deleted before the node sweep, for every language and edge type. An import or call the new parse
names differently (a renamed class, a `dist`→`src` mapping fix) therefore no longer keeps its old
placeholder alive next to the fresh edge, and a placeholder left edgeless falls to the usual
placeholder sweep. Tier-1 (`changedOnly` or non-deep) refreshes do not reap, because a Tier-1 pass
emits fewer edges than a deep one. Edges persisted before the stamp was introduced carry no
generation and are never reaped; the first deep refresh after upgrading stamps them, the next one
reaps — so a project never needs recreating after a parser change any more, two deep refreshes do.
**A module node never lives in a copycode (item 201).** A program whose body is a single `INCLUDE`
used to get a second `MODULE` node named after it with the `.cpy` as `sourceFile` (`ZDTSTBP6` in
`upms`), so `modules` and `search/identifier` listed the name twice. Fixed in the parser; a graph
ingested before the fix keeps the stray node until the project is recreated, or you remove it by hand:
`MATCH (m:MODULE {project: $p}) WHERE m.sourceFile ENDS WITH '.cpy' AND m.ownerModule = '' DETACH DELETE m`
(Natural copycodes are never modules of their own, so the match is exact).
**Known limits of 192** (the later items fill them): no field bindings (195), no styles (196);
the store is item 194 below. A `changedOnly` refresh re-runs the sidecar over the
whole workspace but re-persists only the changed files, so an unchanged file's facts can lag one
refresh (same class of caveat as 46a). The by-name deep ingest resolves dependencies by file stem, so
a dependency whose stem exists in several workspaces (`index`) is reported as a duplicate and skipped
— use `refresh --deep` for the whole frontend.
## Web-service calls and the counterpart link (item 193)
**`rest-endpoints` lists the frontend's calls.** Every member of a generated Endpoint class
(`AgstammControllerEndpoint.saveBroker`) and every hey-api sdk function is a `FUNCTION` of
`kind=endpoint` carrying the same `restPath`/`httpMethod` the Java parser writes for a handler, plus
`outbound=true` — so `GET /projects/purfe/rest-endpoints` answers with `outbound: true` rows
(`handler` = `AgstammControllerEndpoint.saveBroker`, `path` = `/agstamm/ui`). Who calls it:
`modules/{generated module}/callers` names the calling slices/components (module level), and the
member call `api.saveBroker.post(...)` is retargeted from the generic `PostMethod.post` signature to
the endpoint function, so the module-to-module `CALLS` edge carries `calleeMethod =
AgstammControllerEndpoint.saveBroker` and `callerFn = <thunk>`. Since item 203 the synthetic
class-hierarchy edges (`resolvedVia: INHERITANCE`, caller → each implementation of the called
interface/base) exist once per originating call site with its real `lineNo`, `originFile`,
`calleeMethod` and `callerFn` — before, one edge per pair took whichever call line the merge met first.
So `callees` lists every real line for an implementation, and `functions/{impl-method}/callers` also
names callers that go through the interface (`RepoImpl.save` ← `Service.store` via `Repo.save`).
Since item 197
`functions/{fn}/callers` joins these module-to-module edges back to the calling function, so
`purfe/modules/pur-ui/src/generated/endpoints/functions/GeneralAgreementUiControllerEndpoint.createNew/callers`
names the thunk in `generalAgreementSlice`, and on the Java side
`pur/modules/…AgstammLogic/functions/handleMerge/callers` names `AgstammController.mergeBroker`
(a REST controller method itself has no Java callers — it is the HTTP entry point). Extra properties on the node (
`GET /nodes/{id}`): `restUrl` (as composed,
with placeholders and query string), `restBase` (an application base such as `/pur-r-vbuch/v1`
split off so paths compare with the backend's base-less `@Path`), `backend` (`pur`, `pur-r-vstamm`,
`dynamic` — from the URL builder), `queryParams`, `requestType`, `responseType`, `paramsType`,
`generator`, `owner`, `member`. A generated interface's properties are `FIELD`s under its
`DATA_STRUCTURE`, **named `<Interface>.<member>`** (`Broker.ebene`; props `field` = the bare member,
`owner` = the interface — since item 195: a node's identity is type + name + file, and one generated
file declares hundreds of interfaces, so a bare `vid` used to be a single node under six interfaces),
so `GET /data-structures/AgstammUseCase/fields` answers for the frontend too (bare member names;
since item 195 — before, the query filtered `FIELD` out and returned `[]` for an interface), and
`counterparts?kind=field` rows are named `Broker.ebene`.
**`COUNTERPART_OF`: the same thing in another project.** A project setting
`counterparts: ["pur"]` (`ac project create purfe … --counterpart pur`, `ac project update purfe
--counterpart pur`, `GET /projects/purfe` shows it) makes enrichment link, after every refresh of
either side:
| this project | → counterpart | matched on |
|-----------------------------------------------|------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| outbound endpoint `FUNCTION` | backend handler `FUNCTION` | `httpMethod` + path shape (every `{param}` segment compares as `{}`, class + method `@Path` composed as `rest-endpoints` does); several matches → the handler whose source lives under the frontend's `backend` name |
| `DATA_STRUCTURE` in a `generated=true` module | Java `MODULE` with the same `simpleName` | unique name only; an ambiguous name stays unlinked |
| its `FIELD`s | the class's `FIELD`s | name |
The edges are rebuilt from scratch on each run (never accumulated) and re-run for the frontend when
the **backend** refreshes, because a refreshed handler node is deleted together with the edges
pointing at it. Roadmap item 143 (Natural ↔ Java counterparts) will use the same edge.
`GET /projects/{p}/counterparts?module=&kind=rest|dto|field&unmatched=&countOnly=&limit=&offset=`
(CLI `ac counterparts [--module] [--kind] [--unmatched] [--count-only]`) lists
`{kind, name, module, sourceFile, startLine, httpMethod, path, counterpartProject, counterpartName,
counterpartModule, counterpartSourceFile, counterpartStartLine}`; the `counterpart*` fields are null
for an unlinked row, and **`unmatched=true` is the planning question**: which calls does nothing
serve, which generated DTOs / fields have no backend twin. Paged with `X-AC-Total-Count` /
`X-AC-Truncated` like the search endpoints; unknown `kind` → `400 KIND_UNSUPPORTED`; a project
naming itself as counterpart → `400 COUNTERPART_SELF`.
## The Redux store (item 194)
**What is modelled.** Every `createSlice` / `createAppSlice` in a `typescript` project is a
`STORE_SLICE` node **named by the reducer key it is mounted under** in the project's
`configureStore` (`state.<key>`; `gruppenprovision` for `generalAgreementSlice`, whose RTK name is
`generalAgreement`) — the key is what every selector path starts with, so it is the identity; the
RTK name is kept as `sliceName`. The sidecar traces each `reducer: { key: xReducer }` entry through
`xReducer = xSlice.reducer` (also `export default xSlice.reducer`) back to the slice, across
workspaces; a slice no store mounts is named by its own `sliceName`. Under the slice, one **store
`FIELD` per top-level state key**, named `<key>.<field>` (`schluesseltabelle.sucheStatus`, from the
checker's type of `initialState`, so keys only the state type declares are present too) with
`dataType` = the TS type, `optional`, `store=true`, `slice`, `field`. Every reducer is a **`FUNCTION`
of `kind=reducer`** in the slice's module, named by the **action type it handles**:
`schluesseltabelle/updateX` for a case reducer (`reducerKind=reducer`), `schluesseltabelle/suche/fulfilled`
for `builder.addCase(sucheByServer.fulfilled, …)` (`reducerKind=case`, `trigger` = the expression),
`schluesseltabelle/matcher:isSlicePending(sliceName)` for `addMatcher` (`reducerKind=matcher`). A
thunk whose lifecycle action a case handles `CALLS` that case (`callSyntax=extraReducer`), when both
live in the same file.
**Reads and writes.** Inside a reducer every `state.a.b` chain is a `READS`/`WRITES` edge from the
reducer `FUNCTION` to the store `FIELD` of its first key (`state` alone → the `STORE_SLICE`), carrying
`path` (the full sub-path as written, `keyTableUseCaseSvcResult.result.tableId`), `lineNo`,
`via=reducer`. A write is an assignment target (compound assignments also read), `++`/`--`, `delete`,
a mutating method on the chain (`push`, `splice`, `set`, `delete`, …), `Object.assign(state.x, …)`,
or a `return { … }` (each key written) / `return other` (the whole slice). Outside reducers every
store read is a `READS` edge from the reading function (component, hook, thunk; the module when at
top level) to a **placeholder** `<key>.<field>` that the finalize step `resolve-store-placeholder`
redirects onto the real field by name and then drops — so a read in `KeyTablePage.tsx` lands on the
field declared in `keytableSlice.ts` without either file knowing the other. Recognised read forms:
`useSelector`/`useAppSelector((state) => state.a.b)` (every chain rooted at the arrow's parameter;
`const { x, y } = useAppSelector((s) => s.a)` reads `a.x` and `a.y`), **wrapper hooks** such as
`useSchluesseltabelleSelector((useCase) => useCase?.result?.purMode)` — the sidecar finds the wrapper's
inner `useAppSelector((state) => selector(state.a.b))`, so the read is `a.b.result.purMode` with
`via=useSchluesseltabelleSelector` (the wrapper's own inner read is recorded too) — and
`store.getState().a.b` / `const s = thunkAPI.getState(); s.a.b` chains (`via=getState`; 81 sites on
`pur-ui`). A read of a key the store does not declare stays a placeholder and is listed by
`search/identifier` with an empty `sourceFile`.
**Dispatch.** A call to a slice action creator (`dispatch(updateX(…))`, a binding of
`xSlice.actions`) carries `actionType = <sliceName>/updateX` and is a cross-module `CALLS` edge to the
slice module with `calleeMethod = <sliceName>/updateX` — the reducer `FUNCTION` — so module
`callees` of a component list the slices it dispatches into; a thunk call keeps the thunk `FUNCTION`
as target and carries the thunk's type prefix as `actionType`. (Function-level exposure of these
cross-module edges is item 197.)
**Endpoints.** `GET /projects/{p}/store?slice=` (CLI `ac store [--slice]`) → `[{slice, sliceName,
module, sourceFile, startLine, endLine, stateType, fields: [{name, type, optional, reads, writes}],
reducers, reads, writes}]`, all slices or one by reducer key / RTK name.
`GET /projects/{p}/store/{slice}/accesses?field=&mode=reads|writes&module=&countOnly=&limit=&offset=`
(CLI `ac store-accesses <slice> [--field] [--mode] [--module] [--count-only]`) → `[{mode, slice,
field, path, function, functionType, functionKind, module, sourceFile, lineNo, via}]`, one row per
access site, `field` null for a whole-slice access; paged like `counterparts`; unknown `mode` →
`400 MODE_UNSUPPORTED`, unknown slice → `404 SLICE_NOT_FOUND`. Because store fields are `FIELD`
nodes with a unique name, the generic `GET /variables/<key>.<field>/reads|writes` and
`search/identifier?name=<key>.<field>` answer too.
**Limits.** Reads through `getState()` aliases are followed only inside the file that created the
alias; a thunk→case `CALLS` edge is emitted only when thunk and slice share a file (the pur slices
do); a case whose trigger cannot be folded to an action type (a predicate matcher) is named by its
expression text; the sidecar reads the store of every workspace in scope — a key mounted only by a
workspace outside the project (`excludeDirs`) falls back to the slice's own name (matches on `pur`).
No `USES_TYPE` from a store field to the DTO it holds yet — the field's `dataType` says
`SvcResult<KeyTableUseCase>`, the link is item 195. Only a **top-level** `createSlice` is a slice: a
slice built inside a factory function (`pur-ui-common`'s `filetransferSlice`, created per instance
and not mounted in the `pur-ui` store) is not modelled.
**Verified on `purfe` (2026-09-22, server 318, recreated + `refresh --deep`, 285 files, 0
failures, 0 placeholders left):** 9 slices — `error`, `global`, `healthTables`, `metadata`
(pur-ui-common) and `gruppenprovisionSuche`, `gruppenprovision` (RTK name `generalAgreement`),
`schluesseltabelle`, `multilinguism`, `translationdata` (pur-ui) — 43 reducer functions, 216 read
and 82 write sites. `schluesseltabelle` alone: 62 sites, 25 through `useSchluesseltabelleSelector`,
19 through `getState()`, 5 through `useAppSelector`, 13 in reducers.
## DTO field bindings (item 195)
**What a binding is.** The generator emits, next to every DTO interface, a `Fields` class tree
(`AgstammUseCaseField = new AgstammUseCaseFields<AgstammUseCase, never>()`, members
`broker = new BrokerFields<TRoot, Broker>(this, "broker")`, list members as `keyTableList = (index?) =>
new KeyTableDOFields(...)`). A path expression on it — `AgstammUseCaseField.broker.ebene` on a
`<SmartInput field={…}>`, `<SmartOutput field={…}>`, a table's `fieldTermForRowData`, a column's
`field:`, or a `Fields`-typed prop such as `useCaseFieldPrefix` — is a binding. The sidecar types
every hop through the checker (`XFields<TRoot, TSelf>`): the **root DTO** is `TRoot`, the **owner**
of the leaf is the `TSelf` of the hop before it, the leaf name is the field. Each binding becomes a
`READS` edge (plus a `WRITES` edge when the component tag matches `Input$|Dropzone$|Editor$`) from
the binding function (component/hook; the module at top level) to the **`FIELD` of the generated
interface** that item 193 already creates (`Broker` → `ebene`), carrying `path` (dotted hops from the
root, `brokerList[]` for a list hop), `rootDto`, `kind`, `partial`, `component`, `attribute`,
`via=binding`, `lineNo`.
- `kind=field`: the leaf is a scalar (`ebene`); `kind=prefix`: a whole sub-object is handed on
(`useCaseFieldPrefix={X.tab.translationData}`, `fieldTermForRowData={X.keyTableList()}`) —
recorded as a read of the container field so nothing is silently dropped.
- `partial=true`: the expression is rooted at a prop or local (`props.useCaseFieldPrefix.dataName`,
`tabPrefix.x(idx).gausVal`), so the leaf and its owner are exact but the prefix of the path is
unknown (only the tail is given). A carrier prop's own name is not part of the path.
Cross-file targets are placeholders `<Dto>.<field>` (`binding=true`, `owner`, `field`,
`targetModule`) resolved by the finalize step `resolve-binding-placeholder` **exactly** — module →
`DATA_STRUCTURE` → `FIELD` — and dropped afterwards; a leaf the interface does not declare stays a
placeholder (listed by `search/identifier` with empty `sourceFile`). The item-74 stale-edge sweep
covers binding targets on a deep re-ingest.
**Endpoints.** `GET /projects/{p}/bindings?dto=&field=&mode=reads|writes&module=&partial=&countOnly=&limit=&offset=`
(CLI `ac bindings [--dto] [--field] [--mode] [--module] [--partial] [--count-only]`) → one row per
binding site: `{mode, dto, field, path, rootDto, kind, partial, component, attribute, function,
functionType, functionKind, module, sourceFile, lineNo, counterpartProject, counterpartModule,
counterpartField}`. The `counterpart*` columns are the field's `COUNTERPART_OF` twin (item 193), so
**"which page edits Java `Broker.ebene`"** is `ac bindings --dto Broker --field ebene --mode writes
-p purfe` and needs no query on the backend project. `dto` is the *declaring* interface (`Broker`),
not the root (`AgstammUseCase`) — filter on `rootDto` client-side when you need the latter. Paged like
`counterparts`; unknown `mode` → `400 MODE_UNSUPPORTED`. `GET /data-structures/{dto}/fields` now
returns TypeScript interface fields (`type=FIELD`) and carries `boundReads`/`boundWrites` per field
(0 for Natural/Java). The generic `variables/<Interface>.<field>/reads|writes` sees binding edges
too. The leaf's owner is the interface that *declares* the member — `datStart` bound through
`GeneralAgreementDO` lands on `AbstractHistorizedDO.datStart`.
**Limits.** String-form `field="…"` bindings and the lodash-path bindings of `pur-r-vbuch` (excluded
workspace) are not modelled; a `partial` binding cannot say which list element or tab; no
`USES_TYPE` from the component to the root DTO (the `rootDto` edge property answers that). A
binding placeholder and a store placeholder share the `FIELD` type, so a DTO named exactly like a
store key would merge their placeholders (`Dto.field` vs `key.field`) — not the case on `pur`.
**Verified on `purfe` (2026-09-22, server 322, recreated + `refresh --deep`, 285 files, 0
failures):** 257 binding sites (196 reads, 61 writes; 226 scalar leaves, 31 prefixes; 84 partial)
over 23 declaring DTOs, every one linked to its `pur` counterpart field; `SmartInput` 122,
`SmartOutput` 64, tables 32, column definitions 36. `counterparts?kind=field`: 744 fields, exactly
one twin each (six-fold fan-out before the qualified names), 112 unmatched.
## Styling: theme tokens and the style inventory (item 196)
**The theme.** The file with `createTheme({...})` (`pur-ui-common/src/theme.ts`) gets a
`DATA_STRUCTURE theme` (`kind=theme`) with one `FIELD` per token: `theme.<path>` for every leaf of the
literal (`palette.primary.dark`, `typography.h1.fontWeight`, `shape.borderRadius`, `sizes.*`; props
`token`, `tokenKind=path`, `value` folded through constants — `#0054A2` — and `constant` when the
leaf names one, `PRIMARY_DARK`) and `theme.<NAME>` for each exported string/number constant of that
file (`tokenKind=constant`, `PRIMARY`). MUI `components.styleOverrides` are recorded as tokens too, not
interpreted.
**Style blocks.** Every `sx={…}`, `style={…}` and `styled(X)(…)` block is a `STYLE` node under its
component `FUNCTION` (a `styled` block under the `kind=styled` function), named
`<function>.<sx|style|styled>@<line>:<col>`, with `styleKind`, `element` (the JSX tag or styled base:
`Box`, `'div'`), `properties` (the CSS keys, nested selectors flattened: `&:hover.color`,
`& .MuiPaper-root.background`), `literals` (hard-coded colours/lengths: `#005CA9`, `17px`, `-2%`,
`calc(100% - 16px)` — `mt: 2` is theme-relative and no literal), `dynamic` (a value the sidecar could
not classify, or a whole `sx={props.sx}`), `spread`. Plain `.css` files get one `STYLE` per rule from
the Tier-1 scanner (`body@7`, `@font-face@2`; `styleKind=css`, `selector`, `properties`, `literals`).
**Token reads.** Inside a style block every chain on a theme value is a `REFERENCES` edge `STYLE →
theme.<token>` with `property` (the CSS key it feeds) and `via=theme`; a theme value is anything typed
`Theme` (`useTheme()`, a `({ theme }) =>` styled parameter, the theme object imported under any name)
or named `theme`. `theme.spacing(2)` ends at `spacing`, `theme.palette.grey['200']` is
`palette.grey.200`; a theme constant (`color: PRIMARY`) references `theme.PRIMARY`. A token read outside
a style block (`borderColor={theme.palette.grey['200']}`, code) is the same edge from the enclosing
`FUNCTION` with `context` = the JSX attribute or `code`. Cross-file targets are placeholders
`theme.<token>` (`theme=true`) resolved by exact name at finalize when exactly one theme declares the
token; **a token no theme declares keeps its placeholder on purpose** — `GET /theme` lists it with
`declared=false` (MUI defaults such as `palette.grey.200`, `palette.common.white`, the `spacing`
function, or a typo).
**Endpoints.** `GET /projects/{p}/theme?unused=` (CLI `ac theme [--unused]`) → `[{token, kind, value,
constant, declared, module, lineNo, uses}]`, declared tokens first. `uses` counts **project**
references (style blocks and code); MUI's own consumption of a token is invisible, so `unused=true`
means "no project code references it", never "safe to delete" (`palette.primary.main` colours every
Button whether or not a component names it). `GET /projects/{p}/theme/{token}/usages` (`ac
theme-usages palette.primary.dark`, token = dotted path or constant name) → `[{function,
functionKind, module, sourceFile, lineNo, styleKind, element, property, context}]`; unknown token →
`404 TOKEN_NOT_FOUND`.
`GET /projects/{p}/styles?module=&kind=sx|style|styled|css&withLiterals=&countOnly=&limit=&offset=`
(`ac styles [--module] [--kind] [--with-literals]`) → `[{name, styleKind, element, selector,
function, module, sourceFile, lineNo, properties, literals, dynamic, tokens}]`, paged like
`counterparts`; **`withLiterals=true` is the review question**: which blocks hard-code colours and
lengths instead of using the theme. Unknown `kind` → `400 KIND_UNSUPPORTED`.
**Several themes (item 199).** Every `createTheme({...})` in a file is read; a token two themes of
one file declare is one row with the first theme's value and `variants: 2`. A token declared by two
theme *files* (light/dark) is one row per file, and a read of it resolves onto both, so each row
counts the use and `?unused=true` stays honest; `theme/{token}/usages` and `styles[].tokens`
report such a read once. Two reads of one token on one line of a style block (`color: PRIMARY,
borderColor: PRIMARY`) are one usage whose `property` is the comma list of the keys they feed.
CSS rules: a block-less `@import`/`@charset` line is not part of the next selector, braces inside
string values do not open blocks, a nested rule head inside an at-rule body is not a declaration,
and a selector repeated on one line (minified CSS) gets a `:col` suffix in its name.
**Limits.** `className` strings are not matched to CSS rules; Emotion `css` templates and MUI
`styleOverrides` are not modelled; a token read from a component-level function (not a style
block) that survives a re-parse keeps its resolved edge until the file's nodes are re-created
(same class as item 198).
**Verified on `purfe` (2026-09-22, server 326, recreated + `refresh --deep`, 285 files, 0
failures):** 129 declared tokens (111 paths, 18 constants), 93 of them with no project reference;
12 undeclared tokens the code reads (`palette.common.white` 4, `palette.grey.200` 4, `spacing` 3,
`palette.divider`, `palette.text.secondary`, `applyStyles`, `transitions.create`, …);
`palette.primary.dark` is the most-used token (26 reads: 17 `style`, 3 `sx`, 1 `styled`, 5 as a
plain prop such as `confirmColor`). Style inventory: 304 blocks (196 `sx`, 78 `style`, 27 `styled`,
3 CSS rules), 109 with hard-coded literals (`100%` 29, `1px` 21, `12px` 14, `17px` 14, …), 75
reading theme tokens, 41 dynamic. The only placeholders left in the project are the 12 undeclared
theme tokens — by design.
## Missing capability?
If the API/CLI genuinely cannot answer a question (not just
unreachable — the capability doesn't exist), finish the task via
grep/Explore as a fallback, then use `AskUserQuestion` to flag the gap and
ask whether it should become a roadmap item in `x-docs/roadmap.md`. Don't
silently fall back and move on.