1447 lines
211 KiB
Markdown
1447 lines
211 KiB
Markdown
# Using AgenticCode on This Repo (Dogfooding)
|
||
|
||
This repo is ingested as project **`ac`** at `http://localhost:8787`. Per
|
||
CLAUDE.md's dogfooding rule, prefer the AgenticCode API over grep/Explore for
|
||
call graphs, callers/callees, DB access, dataflow, and module overviews when
|
||
working *on this repo* — it's exactly the tool this project builds.
|
||
|
||
**For full API semantics (params, response shapes, error codes, language
|
||
applicability, curl examples)**, see
|
||
`x-docs/agent-api-system-prompt.md` — that file is the canonical reference and
|
||
is not duplicated here. This file only covers what's specific to using the
|
||
API *as Claude Code, on this checkout*.
|
||
|
||
## Tool priority
|
||
|
||
REST first (`GET /api/projects/ac/...`) → `ac`
|
||
CLI (`ac callers`, `ac callees`, `ac call-tree`, `ac context`, `ac
|
||
db-accesses`, ...) → grep/Explore. Only fall back past REST/CLI when the
|
||
question genuinely isn't answerable by this API at all (see "Missing
|
||
capability" below). If the server is unreachable, try
|
||
`./manage-ac.sh deploy` before falling back further.
|
||
|
||
## Re-ingest before trusting results
|
||
|
||
Query results reflect the last ingest, not the current working tree.
|
||
**Refresh after code changes** before trusting query results:
|
||
`ac refresh` or `POST /api/projects/ac/refresh` (add `--deep` / `?deep=true` for a
|
||
full field-level pass). `refresh` is the single (re-)ingest surface (item 42) — the
|
||
eager `ingest-all`/`ingest-module`/`ingest-call-graph` endpoints were removed.
|
||
`ac refresh <name>` deep-ingests one module + its callees/data areas; add
|
||
`--neighborhood` (`POST /refresh/{name}?scope=neighborhood`) to also pull in the
|
||
module's transitive **callers** (whole call-graph neighbourhood).
|
||
|
||
**Reconciliation on re-ingest (item 58).** A `refresh` now **purges stale nodes**:
|
||
for every re-parsed file it deletes the nodes the fresh parse no longer produces
|
||
(renamed/removed fields, moved statements) rather than leaving them to shadow the
|
||
new ones — so identifier counts and `/search/identifier` results stay clean after a
|
||
parser change or an edited source file. Applies to every full-parse path (whole-root
|
||
`refresh` with or without `--deep`, and per-module `refresh/{name}`); the coarse
|
||
Tier-1 scan run at project creation does not reconcile (it only ever runs on an empty
|
||
graph).
|
||
|
||
**Auto-invalidation (item 43).** The graph self-heals when files change or are deleted on
|
||
disk — you rarely need a manual `refresh` for staleness:
|
||
|
||
- **Changed file:** a field-level query on a module whose source changed re-ingests it
|
||
automatically (its stored `sourceHash` no longer matches), so results reflect the current
|
||
file. Granularity is per-module: a changed *dependency* is picked up when that dependency
|
||
is itself queried by name.
|
||
- **Deleted file:** a **whole-project** `refresh` now also **removes nodes for files deleted
|
||
from disk** (reconciled against the filesystem), closing the earlier gap — no project
|
||
re-create needed. (A per-module `refresh/{name}` does not sweep; it only touches its own
|
||
tree.)
|
||
- Controlled by `agenticcode.auto-invalidate.enabled` (default `true`).
|
||
|
||
## Reading source in this repo
|
||
|
||
You have direct filesystem access to this checkout — **never** call
|
||
`nodes/{id}/source` or `modules/{name}/source`. Use `sourceFile`/`startLine`/
|
||
`endLine` from a graph response (context, digest, search/identifier,
|
||
nodes/{id}, ...) and read the file directly. This is always cheaper and gives
|
||
full surrounding context; the `/source` endpoints exist only for API-only
|
||
agents with no filesystem access.
|
||
|
||
**Stale-source check (item 41).** The `/source` endpoints compare the file on
|
||
disk against the content hash (`sourceHash`) stored at ingest. If the file
|
||
changed since the last ingest they return `409 STALE_SOURCE` instead of slicing current text against old line
|
||
numbers — re-ingest (refresh) the project to update the graph. Line ranges you
|
||
read directly off disk are of course always current; this only guards the API's
|
||
own slicing. Copycode/INCLUDE slices are raw pre-expansion file text.
|
||
|
||
## Comments: reachable, but never by default (item 141)
|
||
|
||
**Every default query sees code only.** `search/identifier`, `search/annotation`,
|
||
`search/references` and a plain `search/value` never match comment text — a Natural `* ...` banner,
|
||
a trailing `/* ...`, a Java `//` line or a Javadoc block.
|
||
|
||
That silence used to be a **wrong answer, not a missing one**, wherever a convention records
|
||
something in a comment. The UPMS→PUR case: a reengineered service carries its Natural origin in a
|
||
Javadoc block (`ServiceEndpoint:` / `UPMSFunction:` / `UpmsObject:`), and the documented
|
||
Natural→Java lookup searches the program name in `pur`, where an empty result is read as *"not yet
|
||
reengineered"* — which it answered for services reengineered months earlier.
|
||
|
||
Three routes now reach comment text. Pick one before concluding "not present":
|
||
|
||
| Question | Call |
|
||
|---------------------------------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||
| "What does this module's header/change log say?" | `GET /modules/{name}/comments` (`ac comments <module>`) — blocks with the declaration each documents |
|
||
| "Where is this query / SQL / JSON text defined?" (Java) | `GET /search/value?value=…&contains=true` — since item 139 a `static final String` built from text blocks, literals, same-class constants and `+` carries its full text as `value`; a `"…".formatted(...)` constant carries its template (with `%s`) and `valueKind: template`. Anything built from a method call or another class's constant stays unresolved (no value) |
|
||
| "Does this string appear anywhere, code **or** comment?" | `GET /search/value?value=…&includeComments=true` (`ac search-value --include-comments`) — comment hits carry `kind: "COMMENT"` |
|
||
| "…and in text the parsers do not model at all, or in a module that is not deeply ingested?" | `GET /search/source?regex=…` — raw grep over the files on disk |
|
||
|
||
```
|
||
GET /pur/search/value?value=WPARTX0S&contains=true → [] (code only)
|
||
GET /pur/search/value?value=WPARTX0S&contains=true&includeComments=true → the Javadoc origin block
|
||
GET /upms/modules/WAGNTX0S/comments → `* #01 … Bug 266`, …
|
||
```
|
||
|
||
Comments stay **opt-in** deliberately: a comment hit is not the same evidence as a literal in code,
|
||
and folding them into the default result set would move every existing completeness count (item
|
||
131's lesson). The flip side is the rule to remember — **an empty default search says nothing about
|
||
comments.**
|
||
|
||
## Diagnosing a slow refresh (items 153/155)
|
||
|
||
Two instrumentation layers, both aimed at the same question — *where does the time go?*
|
||
|
||
* **Always on:** every persist batch logs one line with its statement breakdown
|
||
(`Persist deep 200 files: total 8740 ms = prepare 108, merge-nodes 1630, ..., commit 20`), and every
|
||
enrichment step logs duration plus created/deleted rows. The residual `commit` is deliberate: total
|
||
minus the labels is the transaction commit, so nothing hides in an unnamed remainder.
|
||
* **Opt-in per run:** `POST /refresh?deep=true&profile=true` (`ac refresh --deep --profile`) runs the
|
||
enrichment steps under Cypher `PROFILE` and logs the five heaviest operators of every step slower
|
||
than 5 s. Diagnostic only — it answers "is this step matching or writing?", which the step timing
|
||
alone cannot. Measured overhead on `upms`: none worth reporting (908 s vs 905 s).
|
||
* **Build-time diagnostic:** `agenticcode.ingest.comments.enabled=false` drops `COMMENT` nodes and
|
||
their `DOCUMENTS` edges just before persist. It exists to measure what comments cost (item 179:
|
||
~17 % of a deep refresh) and is **not** a supported operating mode — with it off, `/comments`
|
||
answers empty. It is a property, not a request parameter, so it needs a rebuild and cannot be set
|
||
per run.
|
||
|
||
What that measured, so nobody re-derives it. The figures below are **current** (2026-09-06, two clean
|
||
runs, roadmap item 175); the campaign of items 153-175 took an `upms` deep refresh from **1 225 s to
|
||
301 s**, so any older number quoted elsewhere is stale by a factor of four.
|
||
|
||
| | now |
|
||
|---------------------------------|-------------------------------------------------------------------------------------------------------------------|
|
||
| deep refresh `upms`, end to end | **301 s** (run-to-run spread 3.7 %) |
|
||
| persist | ~149 s — `merge-edges` 47.6 s, `commit` 34.1 s, `merge-nodes` 24.0 s |
|
||
| finalize | ~137 s — `resolve-field-placeholder` W/R 43.4 s, `link-args-to-params` 23.7 s, `resolve-bare-included` W/R 37.9 s |
|
||
| parsing | ~9 s (interleaved with persist) |
|
||
|
||
The original diagnosis still holds and is why those steps shrank: the five field-resolution steps
|
||
were dominated by *matching*, not writing — `resolve-bare-included` once spent 273 M database hits to
|
||
produce 43 753 rows. See roadmap items 156 and 175, and item 179 for what comment nodes cost.
|
||
|
||
## `limit`/`offset` work on some endpoints and are silently ignored on others
|
||
|
||
**Every endpoint answers completely.** The difference is whether it lets you ask for less. 19 list
|
||
endpoints declare no `limit`/`offset` at all, and JAX-RS drops an undeclared query parameter without
|
||
a word — so `?limit=2` there is not an error, it simply has no effect:
|
||
|
||
```
|
||
GET /upms/modules?limit=2 -> 3 587 rows
|
||
GET /upms/modules/WAGNTX0S/functions?limit=2 -> 22 rows
|
||
GET /upms/modules/WAGNTX0S/data-structures?limit=1 -> 16 rows
|
||
```
|
||
|
||
**Ignore `limit`/`offset` (always the full list):**
|
||
`/modules` · `/modules/{name}/functions` · `/modules/{name}/functions/overrides` ·
|
||
`/modules/{name}/functions/{function}/overrides` · `/modules/{name}/data-structures` ·
|
||
`/modules/{name}/columns` · `/modules/{name}/payload` · `/modules/{name}/dispatch-table` ·
|
||
`/modules/{name}/sql-statements` · `/data-structures/{name}/fields` · `/db-tables/{name}/columns` ·
|
||
`/variables/{name}/reads` · `/variables/{name}/writes` · `/variables/{name}/flow-forward` ·
|
||
`/variables/{name}/flow-backward` · `/variables/{name}/field-flow` · `/duplicates` ·
|
||
`/dynamic-calls/unresolved` · `/dynamic-calls/overrides`
|
||
|
||
**Honour them** (14, verified against the method signatures): the search endpoints
|
||
(`search/identifier`, `search/value`, `search/annotation`, `search/references`, `search/source`),
|
||
`rest-endpoints`, and per module `callers`, `callees`, `context`, `db-accesses`, `graph`, `reaches`,
|
||
`workfile-accesses`, `comments`.
|
||
|
||
`call-tree` is bounded differently again — by `depth`, not by row count — which is the right shape
|
||
for a tree but means `limit` does nothing there either.
|
||
|
||
Which direction the mistake runs matters, so be precise about it: a dropped `limit` means you get
|
||
**more** than you asked for, never less. It costs tokens, never correctness — the opposite of item
|
||
131's failure, where a silent 50-row cap was read as the complete set. Nothing here can under-report.
|
||
|
||
The practical consequence is budget, not trust: `/modules` on `upms` is 3 587 rows in one response.
|
||
Narrow with the filters those endpoints *do* have (`?kind=`, `?sourceFile=`, `?module=`,
|
||
`?extendsName=`) rather than with a `limit` that will be ignored, and prefer `/modules/{name}/digest`
|
||
or `/context` when you want an overview rather than an enumeration.
|
||
|
||
*(Implementing `limit` on those 19 was considered and deliberately not done: the endpoints are honest
|
||
as they stand, and a `limit` that ever acquired a default would reintroduce exactly the silent
|
||
truncation item 131 removed.)*
|
||
|
||
## `callers` / `callees` say when they are cut (item 181)
|
||
|
||
`callers`, `callees` and `functions/{fn}/callers` now send `X-AC-Total-Count` and `X-AC-Truncated`
|
||
like the search endpoints, and their body carries `total` and `truncated` next to `sourceFiles` /
|
||
`items` (also for `fields=name`; `ac callers` / `ac callees` print the usual "truncated" warning).
|
||
The default page is still 50: `upms/modules/DPARTFN0/callees` answers 50 of 58 with
|
||
`X-AC-Truncated: true` — before, the 8 missing callees (among them the module that writes the
|
||
partner) looked like an inconsistency with `digest`. **Read `truncated` before concluding "X does not
|
||
call Y"**; ask with `limit=1000` or narrow with `scope=external`.
|
||
|
||
## Truncation is now visible on the search endpoints (item 131)
|
||
|
||
`search/identifier`, `search/value`, `search/annotation`, `search/references` and
|
||
`rest-endpoints` send two headers with every answer:
|
||
|
||
| Header | Meaning |
|
||
|--------------------|---------------------------------------------------------|
|
||
| `X-AC-Total-Count` | how many rows match in total, ignoring `limit`/`offset` |
|
||
| `X-AC-Truncated` | `true` when this page leaves some out |
|
||
|
||
and all five accept **`?countOnly=true`** (CLI `--count-only`), returning `{"count": n}` instead of
|
||
rows — a completeness question is a counting question, and `@Column` on `pur` is 3 630 rows ≈ 250 k
|
||
tokens if you ask for them.
|
||
|
||
This closes item 103's own follow-up. The bodies stay bare arrays (no contract change), for the same
|
||
reason as item 130's scope headers. Why it matters: `search/annotation?name=Immutable` returned 50 of
|
||
95 rows with no total, no flag and no `Link`/`X-Total-Count` — and a real UPMS→PUR audit read that
|
||
page as the whole set, recording that 17 entities had lost `@Immutable` when **zero** had. Re-checked
|
||
against all 114 rows of `Tables_meta.csv`: 94 non-writable carry it, 20 writable do not, no
|
||
deviations. The finding cost a day and was pure artefact of the cut.
|
||
|
||
**`search/references` and `rest-endpoints` joined this late (item 135, 2026-08-20).** Item 131 was
|
||
written about the annotation search and both were overlooked. `search/references` was the damaging
|
||
one: it capped at the default 50 and said nothing at all, so a rename scoped from that page missed
|
||
every site past the fiftieth and looked complete doing it. `rest-endpoints` defaults to an uncapped
|
||
limit and so never lost rows, but it was equally silent about how many there are. Both were found by
|
||
`x-scripts/verify-api.sh` on its first run, not by a test.
|
||
|
||
Two mechanics worth knowing: the total costs a **second query only when the page comes back full**
|
||
(a short page is provably the end, so the total is arithmetic), and a total that divides evenly by
|
||
`limit` makes the last full page report `truncated` with the next page empty — one wasted call, never
|
||
a wrong answer. `ac` prints a note to **stderr** when a response is flagged truncated, so piping the
|
||
body into `jq` stays clean.
|
||
|
||
## REST surface and scope headers (item 130)
|
||
|
||
`GET /api/projects/{p}/rest-endpoints?module=&countOnly=&limit=&offset=` (CLI `ac rest-endpoints`) lists
|
||
`{httpMethod, path, module, moduleSimpleName, handler, sourceFile, startLine}` — the **composed**
|
||
path (class-level `@Path` + method-level `@Path`), so "which code runs for `POST /partners`" is one
|
||
call. Previously the two halves had to be joined by hand from two `/search/annotation` calls, because
|
||
annotations are stored by name without their arguments; the parser now persists `restPath` and
|
||
`httpMethod`, including a `@Path` written as a constant reference. A method with no HTTP-verb
|
||
annotation is not an endpoint and is excluded. A class with no `@Path` of its own inherits the
|
||
nearest one from its `extends`/`implements` ancestry, as JAX-RS does. Rows carry `outbound: true`
|
||
when the declaring type is a `@RegisterRestClient` interface — a call the application *makes*, not one
|
||
it serves; its path is usually empty because the base URI comes from configuration.
|
||
|
||
**Scope and freshness now ride on every project-scoped response as headers:**
|
||
|
||
| Header | Meaning |
|
||
|--------------------------|-----------------------------------------------------------------------------------------------------------------------------|
|
||
| `X-AC-Exclude-Dirs` | directories the ingest skipped, or `(none)` |
|
||
| `X-AC-Ingested-At` | when the graph was last walked (item 126) |
|
||
| `X-AC-Ingest-Incomplete` | `true` while a whole-root pass runs or after one that never finished (item 129); `unknown` when no ingest was ever recorded |
|
||
|
||
Read `X-AC-Exclude-Dirs` before trusting an **empty** answer: "no callers" means "none outside tests"
|
||
in a project excluding `test` (`app`) and "none at all" in one that does not (`pur`, `ac`) — the
|
||
bodies are identical.
|
||
|
||
They are **headers, not body fields**, because most endpoints answer with a bare JSON array
|
||
(`db-accesses`, `functions`, `search/identifier`, …); adding a field there would mean restructuring
|
||
array → object and breaking the web UI's generated client, the CLI printers and any agent that
|
||
indexes `[0]`. The trade-off is that an agent reading only the JSON body will not see them — so if
|
||
you consume this API programmatically, read the headers too. The project shell is cached for ~10 s to
|
||
keep this off the request's critical path, and the ingest path invalidates that cache explicitly, so
|
||
`X-AC-Ingest-Incomplete` flips as soon as a refresh starts rather than up to 10 s later.
|
||
|
||
## Every reference site of a name (item 128)
|
||
|
||
`GET /api/projects/{p}/search/references?name=&kind=&countOnly=&limit=&offset=` (CLI `ac references <name>`)
|
||
returns `{sourceFile, lineNo, kind, inModule, target}` per **mention** of a type — not just per call:
|
||
|
||
| `kind` | Where it comes from |
|
||
|-------------------------|--------------------------------------------|
|
||
| `CALL` | a call site (`CALLS`) |
|
||
| `IMPORT` | an `import` of the type |
|
||
| `TYPE` | a declared field / parameter / return type |
|
||
| `ANNOTATION` | the type used as an annotation |
|
||
| `EXTENDS`, `IMPLEMENTS` | inheritance |
|
||
| `INJECTS` | CDI wiring |
|
||
| `CLASS_LITERAL` | `X.class` in argument position |
|
||
| `INCLUDE` | Natural copycode inclusion |
|
||
|
||
Use it to **scope a rename**. `callers` sees calls alone, so a file that only imports the class,
|
||
declares a field of it, or names it in an annotation was invisible — and the rename that missed it
|
||
looked complete. The `name` may be the identity (FQN) or the short form; `target` echoes what it
|
||
resolved to. An unknown `kind` is `400 INVALID_KIND`, never an empty list.
|
||
|
||
**Known limits, by design:**
|
||
|
||
* **Local-variable types and generic type arguments are not indexed** — `List<Target> x` records
|
||
`List`, not `Target`. They multiply edge volume for much less value than the positions above.
|
||
* **Same-package references have no import**, so within one package the index rests on declared-type
|
||
positions alone.
|
||
* **Imports are only indexed when they look project-internal** (they share the first two package
|
||
segments with the importing file). Otherwise every `java.util`/framework import would mint a
|
||
placeholder node on every ingest, just for the finalize sweep to delete it again.
|
||
* **Natural has no import or type-position concept.** It contributes `CALL`, `INCLUDE` and inheritance
|
||
kinds only; this is not parity with Java and should not be read as such.
|
||
* Reference edges are written **at parse time**, so they only exist for files re-parsed since this
|
||
landed — a project needs a `refresh` before the index is complete.
|
||
* Mentions use their own `MENTIONS` edge type, kept out of `CALLS`/`REFERENCES` deliberately: the
|
||
call-graph traversals (`callers`, `callees`, `call-tree`, `ego-graph`) follow `REFERENCES` as
|
||
wiring, so folding imports into it made an `import` surface as a **caller**. `/search/references` is
|
||
the only endpoint that reads `MENTIONS`; the call graph is unchanged.
|
||
|
||
## Refreshing only what changed (item 129)
|
||
|
||
`POST /api/projects/{p}/refresh` (CLI `ac refresh`) has two ways to avoid re-walking a whole root:
|
||
|
||
| Form | What it does |
|
||
|---------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||
| `?paths=a/B.java,c/D.java` (`ac refresh --paths a/B.java,c/D.java`) | Re-ingests exactly those relative paths, deep, plus their dependencies. Paths that match no file — **or that are not ingestible source files at all**, like `pom.xml` — come back in `unresolved`; a typo'd path is never silently dropped. Does **not** run the deleted-file sweep and does **not** move `ingestedAt`: both need a whole-root walk. |
|
||
| `?changedOnly=true` (`ac refresh --changed-only`) | Whole-root walk, but re-parses only files whose content hash differs from the graph's (files with no stored hash count as changed). |
|
||
|
||
**`changedOnly` is opt-in on purpose.** Three whole-walk behaviours are reduced, and one of them would
|
||
be outright corruption if it were hidden:
|
||
|
||
* **A changed Natural copycode disables skipping for that entire run** (logged). Natural projects
|
||
only — a `.cpy` sitting in a Java project (as a test fixture, say) is not inlined by anything and no
|
||
longer stands the optimisation down. Copycode text is
|
||
inlined into the including module *at parse time*, so a module whose `.cpy` changed parses
|
||
differently while its own hash is unchanged — skipping it would leave a stale expansion behind with
|
||
nothing to indicate it.
|
||
* **Duplicate-identity detection** only sees the changed files, so it can confirm duplicates among
|
||
them but not discover new ones elsewhere. Existing markers are never cleared.
|
||
* **User-exit LoC annotation** (item 47) is re-stamped only on re-parsed files.
|
||
|
||
Enrichment is project-wide and still runs in full, so this cuts **parse+persist** time only — not the
|
||
finalize pass. On small projects the whole deep refresh is already ~30 s, so measure before assuming
|
||
a win.
|
||
|
||
**`ingest.incomplete`** (on `/projects` and `/projects/{p}`) is `true` while a whole-root pass runs and
|
||
**stays true if one never finished** — a crash, a container stop, an aborted deep refresh. Before
|
||
this, an interrupted deep refresh was indistinguishable from a clean graph: the enrichment steps that
|
||
already ran are committed, so queries keep answering, just from a half-updated graph. It cannot
|
||
self-heal (a killed process clears nothing) and does not distinguish "running right now" from "died an
|
||
hour ago" — both mean the same thing to a caller. A completed refresh clears it.
|
||
|
||
## Renaming a project (item 202)
|
||
|
||
`POST /api/projects/{p}/rename` with `{"newName": "…"}` (CLI `ac project rename <old> <new>`)
|
||
moves the project key on every node and override and in other projects' `counterparts` lists, then
|
||
the shell; the old name answers `404` afterwards and nothing needs re-ingesting. Refusals:
|
||
`400 INVALID_REQUEST` (blank or unchanged), `404`, `409 PROJECT_EXISTS`. The node rewrite is batched;
|
||
an interrupted rename is finished by running it again (97 s for the 939k-node `upms`). Use it to
|
||
keep a reference graph next to a fresh ingest (`ac project rename upms upms_alt`, then create `upms`
|
||
again and compare). A full `DELETE` of a project now also removes its manual overrides; `recreate`
|
||
keeps them.
|
||
|
||
## Is this project's graph any good? (item 126)
|
||
|
||
`GET /api/projects` and `GET /api/projects/{p}` (CLI `ac project list` / `ac project show <p>`)
|
||
carry an `ingest` object describing the **last whole-root ingest**:
|
||
|
||
```json
|
||
"ingest": { "ingestedAt": "2026-08-18T10:12:44Z", "mode": "full", "filesExamined": 2981,
|
||
"filesPersisted": 2977, "filesFailed": 4, "failures": ["a/B.java", "..."],
|
||
"failuresTruncated": false, "durationSeconds": 176, "serverVersion": "…" }
|
||
```
|
||
|
||
Use it before trusting a **negative** answer: without it, "no such module" and "that part of the
|
||
project was never ingested" are the same empty response. Three rules the field obeys:
|
||
|
||
* **`ingest: null` means never recorded**, not "ingested nothing" — a project last walked before
|
||
this existed reads as null rather than as a fabricated zero.
|
||
* **Only whole-root passes write it** — the create-time Tier-1 scan, `refresh`, `refresh?deep=true`.
|
||
A by-name `refresh/{name}`, a deep ingest or a fan-out warm ingests real files but sees a fraction
|
||
of the tree, so it deliberately leaves `ingestedAt` alone; otherwise deepening one module would
|
||
advertise the whole project as freshly walked.
|
||
* **`ingestedAt` is not a freshness guarantee.** It says when the walk ran, not that the graph still
|
||
matches disk — a file edited a minute later is stale while the timestamp still looks recent. For
|
||
the real check, read a file through `GET /{p}/source?file=…`, which answers `409 STALE_SOURCE`
|
||
when the content no longer matches the ingested hash. (Targeted/incremental refresh is item 129.)
|
||
|
||
`failures` is capped at 200 paths while `filesFailed` stays exact; `failuresTruncated` says whether
|
||
the list was cut, so a short list is never mistaken for the whole story.
|
||
|
||
## Tier-1 coarse scan on project create (item 36)
|
||
|
||
Creating a project (`POST /api/projects/{p}`) now runs a **Tier-1 coarse reference scan** of the
|
||
root before returning, so the project is immediately queryable — no separate ingest call. The scan
|
||
is lexer-level (Natural) / a coarse projection of the full parse (Java): it records module/function
|
||
/data-structure **shells**, the **identifier index** (declared fields, class members), and coarse
|
||
`CALLS`/`READS`/`WRITES`/`INCLUDES` references (Natural `PERFORM`, `CALLNAT '...'`, dynamic `CALLNAT
|
||
PGM-VAR`, `PARAMETER`/`LOCAL USING` copybooks; Java resolved calls/type refs) — but **no** deep
|
||
bodies (control flow, statement-level dataflow, arg→param). Scanned modules land
|
||
`CALL_GRAPH`/`NOT_INGESTED`; field-level detail is filled in by the on-demand deep ingest below. Each
|
||
module shell carries a `sourceHash`. Disable with `agenticcode.tier1.scan-on-create=false` (creates
|
||
an empty project shell to scan later). A missing/unscannable root is non-fatal — the project is
|
||
created empty and can be re-scanned.
|
||
|
||
## Unresolved references (item 40)
|
||
|
||
A reference whose target isn't (yet) ingested — a `CALLNAT`/`PERFORM`/`USING` to a module/copybook
|
||
absent from the project, or a dynamic `CALLNAT PGM-VAR` whose literal can't be recovered — is stored
|
||
as a **deduped placeholder node** (blank `sourceFile`). Enrichment stamps each with an `unresolved`
|
||
boolean: `true` while genuinely dangling, `false` once a real definition of that name is ingested.
|
||
`/search/identifier` returns it as `unresolved` on each `IdentifierMatch`, and `GET /nodes/{id}` carries it in the
|
||
node's properties — so an agent can tell a dangling/dynamic
|
||
reference apart from a resolved one. In `callers`/`callees` such targets already appear as entries
|
||
with a blank `sourceFile`.
|
||
|
||
### Querying a module endpoint: four answers, not one (items 107, 115, 114)
|
||
|
||
Every `GET /api/projects/{p}/modules/{name}/…` endpoint used to answer `200` with an all-zeros shell
|
||
for a name that exists nowhere in the graph, byte-identical to a real-but-empty module's answer. That
|
||
is not cosmetic: `call-tree?depth=4` returning `200` with 0 modules reads as **"analysed, nothing
|
||
found"** when the truth is **"not analysable"** — the module's source was never in the checkout. The
|
||
graph knows four states and each now gets its own status:
|
||
|
||
| state | answer | what it means |
|
||
|-------------------------------------------------------------------------------|---------------------------------------------------------------|------------------------------------------------------------------------|
|
||
| no `MODULE` node for the name | `404 MODULE_NOT_FOUND` | unknown name (typo, or not in this project) |
|
||
| **placeholder** — node exists because something calls it, source never parsed | `409` + `{status:"NOT_INGESTED", module, detail, nextAction}` | knowable in principle, not analysed yet |
|
||
| **ambiguous** — several real modules share the name (Java) | `409` + `{code:"AMBIGUOUS_NAME", details:{candidates:[...]}}` | the name does not identify a module; pick one |
|
||
| **duplicate identity** — the same name in several files, skipped at ingest | `409` + `{code:"DUPLICATE_IDENTITY", details:{paths:[...]}}` | exists twice, deliberately not ingested (item 114); see `/duplicates` |
|
||
| real, ingested module | `200` | the body is the answer — an empty body genuinely means "nothing found" |
|
||
|
||
**Ambiguous names (item 115).** A Java simple name is not unique: nested `@Nested` test classes,
|
||
`Builder`, `Config`, `WorkingStorage`. Measured on `pur`, **163 names covering 385 modules (~8%)** were
|
||
addressable only ambiguously, and the endpoints used to answer with the *union* across unrelated
|
||
classes — `/modules/BrokerHistoryTests/functions` returned 627 functions for a class that has 107.
|
||
They now refuse and list the candidates. Repeat the request with `?sourceFile=<candidate>`:
|
||
|
||
```
|
||
GET /modules/Shared/functions → 409 AMBIGUOUS_NAME, candidates ["a/Shared.java","b/Shared.java"]
|
||
GET /modules/Shared/functions?sourceFile=a/Shared.java → 200
|
||
```
|
||
|
||
**A Java module's name IS its fully-qualified name** (item 117) — `com.example.OrderService`, and
|
||
`com.example.Outer.Inner` for a nested class. That is what makes same-simple-name classes
|
||
distinguishable at all. Every module endpoint accepts **either** form:
|
||
|
||
```
|
||
GET /modules/com.example.OrderService/digest → 200, always exact
|
||
GET /modules/OrderService/digest → 200 when unique, else 409 AMBIGUOUS_NAME
|
||
```
|
||
|
||
Responses carry `simpleName` alongside `name` for display. The `?module=` and `?extends=` filters and
|
||
the `ac` CLI take either form too; `--source-file` remains available on every module command.
|
||
|
||
**Which Java types are modules** (item 119). Classes, interfaces, **enums, records and annotation
|
||
types** — `moduleKind` is one of `CLASS | INTERFACE | ENUM | RECORD | ANNOTATION` (Natural adds
|
||
`PROGRAM | SUBPROGRAM | …`), and `?moduleKind=` filters on it. Their content is modelled the way each
|
||
kind carries it: a record's components and an annotation type's members are `FIELD`s (the latter with
|
||
`defaultValue` where declared), an enum's constants are `CONSTANT`s, and an enum's or record's
|
||
`implements` is a real edge — so a call against an interface fans out to an enum implementing it.
|
||
|
||
Before 119 these three kinds were not parsed at all: `GET /modules/SomeEnum/digest` answered `404`,
|
||
and a record referenced from elsewhere stayed an unresolved placeholder (`409 NOT_INGESTED`). One
|
||
gap remains by design — a record's *compact* canonical constructor is not a function node, so calls
|
||
made in its body are invisible.
|
||
|
||
Natural is unaffected throughout: its module names are file stems, and colliding identities are
|
||
skipped at ingest, so they are unique by construction (`upms` has zero ambiguous names, `pur` 163).
|
||
|
||
One limit worth knowing: only the request's **root** module is disambiguated. Inside a traversal this
|
||
no longer merges anything, because module names are unique per project after item 117 (`pur` and `upms`
|
||
have zero duplicate names) — a reference the parser could not qualify is returned flagged
|
||
`unresolved: true` rather than attached to an arbitrary candidate.
|
||
|
||
A name skipped at ingest because it exists in **more than one file** answers `409 DUPLICATE_IDENTITY`
|
||
with the conflicting `details.paths` (item 114); `GET /duplicates` lists them all. Note what that does
|
||
*not* fix: the skipped file's own calls were never parsed, so caller lists elsewhere can still be
|
||
short — they just no longer look complete.
|
||
|
||
The `409` applies to everything derived from the module's **own** source: `digest`, `context`,
|
||
`call-tree`, `callees`, `db-accesses`, `workfile-accesses`, `sql-statements`, `functions`,
|
||
`functions/overrides`, `functions/{fn}/overrides`, `functions/{fn}/callers`, `data-structures`,
|
||
`dispatch-table`, `payload`, `columns`.
|
||
|
||
`callers` and `graph` stay `200` for a placeholder — their data comes from the **calling** modules'
|
||
source and is genuine. When you get a `409`, **fall back to `/callers`**: it is the one honest answer
|
||
available for a module whose own source is missing. (`digest` no longer surfaces those callers, since
|
||
its other fields would all be structurally zero.)
|
||
|
||
`/modules/{name}/source` is unchanged: it already answered `404 MODULE_NOT_FOUND` for both an absent
|
||
module and a placeholder, since there is no source to serve either way.
|
||
|
||
**Scope limit — this guards the *root* module of a request only.** A `call-tree` that traverses *into*
|
||
placeholder targets still reports that subtree as empty without flagging it, so a dispatcher whose
|
||
targets are all placeholders still returns `200` with a silently truncated tree. Cross-check the
|
||
targets you care about individually (a `409` tells you it is unanalysed) — see item 103.
|
||
|
||
**Data literals are not call targets (item 62).** A `CALLNAT <bareword>` whose target is really a data
|
||
value — a browse key reaching the call site through a copycode/macro argument — used to leave a
|
||
permanent unresolved placeholder callee. Enrichment now reaps such a placeholder when it has no Natural
|
||
sigil (`#`/`&`/`+`), *is* a real `VARIABLE`/`CONSTANT` of the project, matches no real `MODULE`, and is
|
||
only ever reached by inferred (`CALLNAT_DYNAMIC`/`INCLUDE_MACRO`) edges. So `callees`, `call-tree`, the
|
||
ego graph and the unresolved list no longer list data names as modules. Genuine dynamic dispatch
|
||
(`CALLNAT #PGM-VAR`, sigil'd) is still reported as an unresolved target, and a static `CALLNAT 'X'` is
|
||
always trusted even when `X` collides with a field name.
|
||
|
||
**The ingest summary agrees with the graph (item 64).** A `refresh`/`refresh/{name}` response's
|
||
`unresolved` list is built during the file walk, independently of the graph — before item 64 it therefore
|
||
reported data fields as missing modules (`MODULE CO-TABLA`, `MODULE NAME-DESC-SP`) that enrichment had
|
||
already reaped, so the summary and the graph disagreed. It now applies the same item-62 gates as the graph
|
||
reap, with the "is this a real field?" test asked project-wide (a data literal is often declared outside
|
||
the ingested tree). Genuinely missing modules are still reported — including sigil'd dynamic dispatch
|
||
(`MODULE #GETSHORT-MODUL`) and a missing module whose name collides with a field name but is called
|
||
statically (the `RPC-CNTX` class). Treat `unresolved` as "dependencies that really are absent".
|
||
|
||
**Constant-folded string-assembled targets (item 83).** A dispatcher often builds the `CALLNAT <var>`
|
||
name from a base literal plus one or more `SUBSTR` overlays — e.g. `#GETSHORT-MODUL` in `YGEAGGNH`,
|
||
assembled by `MOVE 'YGEAGKEY' TO #GETSHORT-MODUL` then `MOVE 'GN0' TO SUBSTR(#GETSHORT-MODUL,6,3)` →
|
||
`YGEAGGN0`. The parser records each `SUBSTR` write as a `WRITES` carrying `substrPos`/`substrLen`
|
||
(1-based), and the `resolve-dynamic-callnat-fold` enrichment step folds the last full-var literal
|
||
written before the call site with the intervening overlays (`left`/`substring`) and MERGEs a resolved
|
||
`CALLS` edge (`callKind=CALLNAT_DYNAMIC`, `folded=true`) to the assembled module when it is a real
|
||
`MODULE`. It runs in every ingest mode (cheap, bounded by dynamic call sites) and over-approximates like
|
||
the other dynamic resolvers. This auto-recovers the `Y…GNH → Y…GN0` family with no manual override, so
|
||
folded sites drop out of `dynamic-calls/unresolved` and the assembled target appears in
|
||
`callees`/`call-tree`/`graph`/`ego graph` tagged `CALLNAT_DYNAMIC`.
|
||
|
||
**Pin what the resolvers can't: manual dynamic-`CALLNAT` overrides (item 82).** Some `CALLNAT <var>`
|
||
targets are beyond the auto-resolvers — a name never present as a recoverable literal in the reachable
|
||
code (read from a work-file record, supplied by an unlinked caller, …). Such a site stays an unresolved
|
||
placeholder. A human or agent resolves it via
|
||
`POST /api/projects/{p}/dynamic-calls/overrides` with the call site's `originFile` + `lineNo` (from
|
||
`GET .../dynamic-calls/unresolved`) and the target module name(s) — multiple targets for a genuine
|
||
branch. The override is stored as a `:DynamicCallOverride` node **outside** the `:AstNode` graph, so a
|
||
refresh never deletes it and an enrichment step (`apply-manual-dynamic-callnat`, after the auto
|
||
dynamic-CALLNAT resolvers, before the placeholder cleanup) **re-applies it automatically** — MERGEing a
|
||
`CALLS` edge (`callKind=CALLNAT_DYNAMIC`, `resolvedBy='manual'`) to each target and flagging the
|
||
placeholder `manualHidden` so `callees`/`digest`/`graph`/`call-tree` show the real target, not the
|
||
`#var`. It only applies while the site is still unresolved: once an auto-resolver catches up, the
|
||
override is skipped and listed `obsolete` — **except a constant-fold (item 83), which a manual override
|
||
outranks**: the fold skips a site carrying a `:DynamicCallOverride`, and a stale `folded` edge there is
|
||
dropped (`delete-folded-overridden-dynamic-callnat`) before `apply-manual-dynamic-callnat` runs, so the
|
||
pinned target replaces it. `DELETE .../dynamic-calls/overrides?originFile=&lineNo=`
|
||
resets one site (omit both = all), clearing the manual edges and un-hiding the placeholder inline (no
|
||
refresh). A target that is not a real `MODULE` is rejected `400 UNKNOWN_TARGET`. **Bug B fix:** the
|
||
`callees` items now carry `unresolved` (mirroring what `graph` already exposed), so an unresolved
|
||
dynamic target is machine-distinguishable from a resolved one without inspecting `sourceFile`. Since item 200 the apply
|
||
and the reset also rebuild the calling
|
||
modules' derived `CALLS_MODULE` edges in the same transaction, so `reaches` and `field-flow` honour
|
||
a pinned target without a refresh, like `callees` and `call-tree` already did.
|
||
|
||
**Dispatch guards: read `guards` — it is the only complete condition (item 72).** A `dispatch-table` row's
|
||
`guardField`/`guardValue`/`guardValues` describe the **innermost** `DECIDE` only. Natural nests
|
||
value-`DECIDE`s inside each other (824 of upms's 7,729 blocks, 178 modules), and then the innermost guard
|
||
is just one conjunct: in `VMULTMN4`, the row for `YTABLMA0.TX-TABLA` reports `#FIELD-NAME = 'TX-TABLA'`,
|
||
but the assignment also requires `#SHORT-VIEW = 'TABL'`. `guards` is the full chain — `[{field, values}]`,
|
||
**outermost first, joined by AND**, each link's `values` joined by OR. Reading only the legacy fields
|
||
**over-generalises**: port that to Java and you get a branch firing where Natural never would. For an
|
||
unnested `DECIDE` the chain has one link and says the same as the legacy fields.
|
||
|
||
- **Still incomplete for `NONE`/`ANY` branches (item 73):** an assignment in a `NONE` branch is reported
|
||
under its enclosing chain alone, but its real condition is "enclosing guard **AND NOT** any sibling
|
||
`VALUE`" — a negation a chain of equalities cannot express. `guards` is strictly better than the legacy
|
||
fields, not a total answer.
|
||
|
||
**A dispatch row's `lineNo` belongs to `sourceFile`, not to the module (item 122).** `dispatch-table`
|
||
rows now carry the same provenance quartet as `callees`/`db-accesses`/`workfile-accesses`/`functions`:
|
||
`sourceFile`, `viaCopycode`, `includedAt`, `includePath`. Resolve `lineNo` **against `sourceFile`** —
|
||
when `viaCopycode` is non-null the assignment is written in that copycode and `lineNo` is a line of the
|
||
`.cpy`, while `includedAt` is the `INCLUDE` line in the module. Before item 122 the row carried only
|
||
`lineNo`, so following it against the module file landed somewhere arbitrary: 26 of `VCOMIN50`'s 44 rows
|
||
reported line 18 or 20, which in that module is a change-history comment; the real sites are
|
||
`ISICINDE.cpy:18` and `ISICINDI.cpy:20`. Note the guard chain may **span** the include boundary — the
|
||
outer `DECIDE` in the module, the inner one in the copycode — so a row can have a multi-link `guards`
|
||
chain whose links live in different files.
|
||
|
||
**Within one guard: prefer `guardValues` over `guardValue` (item 64).** `dispatch-table` rows carry both.
|
||
`guardValue` is **lossy** and kept only for compatibility: it comma-joins the branch's `VALUE` literals,
|
||
which drops blank alternatives entirely and, for a multi-alternative branch, yields a synthetic string the
|
||
guarded field never equals (`"A1, A2"`). `guardValues` is the faithful list — every alternative in source
|
||
order, blanks included — so `VALUE 'GENAGREE-WOUT-SP', ' '` reports `["GENAGREE-WOUT-SP", " "]`, recording
|
||
that a **blank** guard field also routes into that branch. When reasoning about routing (or porting a
|
||
`DECIDE` to Java), read `guardValues`; `guardValue` will silently under-report the branch's conditions.
|
||
|
||
**Text that is not code never yields a call (items 61 & 63).** The `CALLNAT`/`CALLNAT_DYNAMIC` patterns
|
||
are unanchored (a `CALLNAT` may legally appear mid-line), so both ingest tiers first neutralise non-code
|
||
text: a full-line `*` comment is skipped, a trailing `/* …` is stripped (item 61), and a match whose
|
||
**keyword** falls inside a quoted string literal is rejected (item 63). A real `CALLNAT 'MOD'` is
|
||
unaffected — its keyword sits outside the quotes. This matters for trusting `callers`/`callees`/
|
||
`call-tree`: before item 63, prose such as `WRITE(#MSG) 'NACH CALLNAT ISINGEAG:'` or `#ERR-TYPE :=
|
||
'Callnat USIA008N'` fabricated a `CALLNAT_DYNAMIC` edge to the **real** module of that name, so a mere log
|
||
message appeared as a genuine call — and, because a real module existed, it was *not* flagged
|
||
`unresolved` and could not be reaped by the item-62 cleanup. If you query a graph ingested before
|
||
2026-07-16, re-ingest (`ac refresh`) before trusting call-graph edges into modules that are also
|
||
mentioned in log/error text.
|
||
|
||
**Calls made through a copycode's arguments (items 120/121/123).** In Natural the target of a
|
||
`CALLNAT` is often not written at the call site at all: a copycode receives the module name as a
|
||
positional `INCLUDE` argument and issues `CALLNAT &2&`. Three defects in that argument path — arguments
|
||
continued on the next line, the doubled-quote escape `'''X'''`, and the double-quote delimiter
|
||
`'"X"'` — meant such a call produced **no edge and no unresolved-dynamic-call entry**, so `callees`,
|
||
`callers`, `call-tree` and `reaches` agreed on an answer that was simply absent, with nothing saying
|
||
"not analysed". This hit the browse/access layer hardest, because that is where the idiom lives:
|
||
`YCARPBN1` and `YPOLIBN1` reported **0** callers each. Fixed 2026-08-07; re-ingest recovered 2672
|
||
copycode-derived call pairs (+49%) in `upms` with none lost. **A graph ingested before 2026-08-07
|
||
under-reports Natural callers/callees, and does so silently — re-ingest before concluding a Natural
|
||
module is unused.** A copycode parameter that genuinely has no argument now surfaces in
|
||
`/dynamic-calls/unresolved` as `&n&` rather than being dropped, so "not analysable" is visible.
|
||
|
||
**A call edge no longer outlives the call it was parsed from (item 124).** Until 2026-08-09 a refresh
|
||
only *added* the corrected call and left the old one in place, because an edge is reaped only when one
|
||
of its endpoints is — and a parser fix changes neither (the calling subroutine is unchanged, the old
|
||
target is a never-swept placeholder). So `callees`/`callers` could report a call that no source line
|
||
makes, flagged `unresolved: true` and indistinguishable from a genuine unresolved dynamic call. A
|
||
re-parsed Natural file's call edges are now reaped before the fresh ones are merged, and a call-target
|
||
placeholder left with no callers is deleted. **Two consequences for a graph ingested before
|
||
2026-08-09:** an `unresolved: true` callee may be an artefact of an already-fixed parser bug rather
|
||
than a real dynamic call, and `search/identifier` may list module names that exist nowhere in the
|
||
source. Both clear on the next deep refresh. Note the reap deliberately spares `CALLNAT_DYNAMIC` edges
|
||
onto *real* modules — those are the dynamic-call resolvers' output, not the parser's.
|
||
|
||
## LoC / SLoC metrics (item 46)
|
||
|
||
Every file-level node (a `MODULE` program/class, or a `DATA_STRUCTURE` for a Natural `.lda`/`.pda`
|
||
data area) is stamped at ingest with two deterministic line metrics:
|
||
|
||
- **`loc`** — physical lines of the file (language-independent; a trailing newline adds no phantom
|
||
line).
|
||
- **`sloc`** — source lines of code: non-blank, non-comment lines, computed **per language**. Natural
|
||
drops full-line `*`/`**`/`/*` and inline `/*` comments; Java drops `//` and `/* … */` blocks while
|
||
keeping those tokens when they appear inside string literals.
|
||
|
||
Both the cheap Tier-1 coarse scan and the deep Tier-2 parse use the *same* per-language counter, so a
|
||
module's `loc`/`sloc` are identical at any ingest depth — you can sum them to get exact,
|
||
reproducible project totals.
|
||
|
||
Where to read them:
|
||
|
||
- `GET /modules` — `loc`/`sloc` on each row.
|
||
- `GET /modules/{name}/context` — `loc`/`sloc` on the module.
|
||
- `GET /nodes/{id}` — `loc`/`sloc` in the node's raw properties.
|
||
- `GET /loc` (`ac loc`) — the rollup: a per-language breakdown (`fileCount`, `loc`,
|
||
`sloc`) plus a project-wide total, optionally narrowed by `?language=` / `?sourceFile=`. Each source
|
||
file is counted once even when it yields several nodes (Java inner classes, Natural inline groups).
|
||
|
||
`null` metrics mean the node predates item 46 — re-ingest (`refresh`) to backfill.
|
||
|
||
### Generated vs. user-exit split (item 47)
|
||
|
||
A project can be created with a source **`language`** (required at creation; an attribute only — ingest
|
||
still classifies files by extension) and a **`generatedDir`**/**`userExitDir`** pair (directory *names*,
|
||
matched as path components like `excludeDirs`; both or neither). Generated modules already contain their
|
||
hand-written user-exit twin inline, so at ingest a module under `generatedDir` whose name also occurs
|
||
under `userExitDir` is annotated with that twin's LoC/SLoC (`userExitLoc`/`userExitSloc`). User-exit
|
||
files are **not** ingested as standalone modules (they would collide by name) — the walk skips
|
||
`userExitDir`.
|
||
|
||
**Consequence for all non-LoC analysis:** the `generatedDir` copy is the *canonical, sole* module for
|
||
every structural query (call graph, DB access, functions, data structures, identifiers, dataflow,
|
||
dispatch table). `userExitDir` exists **only** to compute the generated-vs-manually-written LoC split
|
||
below; it never contributes nodes/edges. So when verifying an API response against source for a Natural
|
||
module, always read the `generatedDir` file (e.g. `generated_src/subprogram/WGEAGB0S.nat`), not the
|
||
`user_exit` fragment.
|
||
|
||
`GET /loc` (`ac loc`) then reports, per language row **and** in the project total:
|
||
|
||
- **`loc`/`sloc`** — the **total** (generated, which already includes the user exits).
|
||
- **`userExitLoc`/`userExitSloc`** — the sum of the annotated user-exit twins (the hand-written part).
|
||
- **`generatedExclusiveLoc`/`generatedExclusiveSloc`** — total − user-exit, clamped ≥0 per file (the
|
||
purely generated part).
|
||
|
||
All three are `0` for projects without a generated/user-exit split. Create with
|
||
`ac project create <name> <root> -l natural -g generated_src -u user_exit`, or add the split to an
|
||
existing project via `ac project update <name> -g generated_src -u user_exit`.
|
||
|
||
## Java `DB_ACCESS` in a project without JPA entities (item 140, 2026-08-27)
|
||
|
||
A `DB_TABLE` node is only ever created from a JPA `@Entity` or a Panache active-record class. In a
|
||
Java project that has none, **no** `DB_ACCESS` candidate can resolve — and the parser's candidate
|
||
heuristic is a deliberate over-approximation: its read gate admits *any* static receiver whose method
|
||
starts with `get`/`find`/`read`/`list`/… , so `UserContext.getCurrent()` and
|
||
`TextUtils.getColumn(line, 0, 8)` become candidates. On `app` that produced **2219** `DB_ACCESS`
|
||
nodes in a codebase with no database access whatsoever.
|
||
|
||
`db-accesses` never showed them (it joins the table with a plain `MATCH`), but **`sql-statements`
|
||
did**: it joins with `OPTIONAL MATCH`, so unresolved candidates came back as rows with
|
||
`"table": null` — 76 of them on a single `app` module.
|
||
|
||
Since 2026-08-27 the enrichment step `reap-java-db-access-without-tables` deletes every Java
|
||
`DB_ACCESS` of a project that holds no `DB_TABLE`. For such a project `sql-statements` is now empty
|
||
instead of noisy. Three things to know:
|
||
|
||
- **The gate is project-level, not per node.** One entity anywhere in the project switches the reaper
|
||
off, and the unresolved candidates stay. `pur` (2014 of 3951 unresolved) and `ac` (221 of 335) are
|
||
unaffected, and still return `table: null` rows. Treat a `sql-statements` row whose `table` is
|
||
`null` as unverified, in any project that has tables.
|
||
- **Natural is untouched.** A Natural `DB_ACCESS` comes from a literal `READ`/`FIND`/`STORE` and is a
|
||
real access whether or not its view resolved (`upms`: 14302 of 14303 resolve).
|
||
- **Recovering from it needs a full refresh.** If such a project later gains its first entity, the
|
||
reaped nodes only come back for files that are actually re-parsed — `changedOnly` will not restore
|
||
them, `refresh` without it (or `recreate`) will.
|
||
|
||
## `?depth=` means module hops (item 65)
|
||
|
||
On `db-accesses` / `sql-statements` (and the `?module=` scope of `variables/{name}/reads|writes`),
|
||
`depth=N` means **N module calls away** — the same unit `/modules/{name}/graph?depth=` and `call-tree` neighbours
|
||
use. `depth=1` = the modules this one directly `CALLNAT`s, regardless of how deeply the calling
|
||
statement sits inside subroutines.
|
||
|
||
Before item 65 these endpoints bounded the traversal on raw `CALLS` edges. A `CALLS` edge starts at the
|
||
*statement* making the call, not at the `MODULE` node, so the traversal also stepped through internal
|
||
`PERFORM` jumps and `depth` measured **statement nesting**, not dependency distance. Concretely:
|
||
`WGEAGB0S` reached `YGEAGBNH`'s tables through two module calls, but the raw path is 5 edges
|
||
(`WGEAGB0S → GET-DATA → GET-MAIN-DATA → BGEAGFN0 → READ-FILE → YGEAGBNH`), so `db-accesses?depth=2`
|
||
returned `[]` — reading as "no DB access" — and only `depth=5` was truthful.
|
||
|
||
**If you scripted a depth workaround** (a deliberately large `depth` to compensate), drop it: `depth`
|
||
is now the value you'd naturally expect, and inflated values just widen the result set.
|
||
|
||
> **`call-tree`'s `depth` column is still raw-hop based** and mixes internal subroutines into the tree:
|
||
> a direct dependency called from the main body shows `depth=1` while one called two subroutines deep
|
||
> shows `depth=3`. Use the ego graph (`/modules/{name}/graph`) when you need module-level distance. Tracked as an open
|
||
> roadmap item.
|
||
|
||
## Framework-mediated DB access via `INCLUDE` macros (item 44)
|
||
|
||
Natural's generic table-access framework hides a `CALLNAT` inside a copycode member, invoked with a
|
||
statement-level macro:
|
||
|
||
```natural
|
||
INCLUDE YFRAMGC0 'YFRAML01.C-MOD-GET' '"CO-TABLA-CO-ELEMENTO-ALT"' '"YELEMGN0"' 'YELEMKEY' 'YELEMREC'
|
||
```
|
||
|
||
The `CALLNAT` to the generic accessor (`YELEMGN0`) lives in the copycode, not in the including module,
|
||
so before item 44 both `callees` and `db-accesses` were empty for such modules. The parser now
|
||
recognises the framework macro and emits a `CALLS` edge to the accessor named in the macro arguments
|
||
(de-quoted; e.g. `'"YELEMGN0"'` → `YELEMGN0`), tagged with **`edgeKind = INCLUDE_MACRO`** on
|
||
`callers`/`callees`. Because the edge is a normal `CALLS`, the accessor's own table access surfaces
|
||
transitively: `GET /modules/{name}/db-accesses?depth=N` reports the table with `via` = the accessor
|
||
module. The recognised macros and which argument names the accessor are described declaratively in
|
||
`FrameworkMacros` (ac-parser-natural). Scope: the targeted recogniser only — general `.nsc` copycode
|
||
expansion is still open.
|
||
|
||
**`db-accesses?depth=N` is a superset of `db-accesses` (item 93).** Besides `READS`/`WRITES` it also
|
||
returns the `mode: "DECLARES"` rows — a Java entity's own `MAPS_TO` table and a repository's
|
||
`repositoryEntity` table (item 32) — for every module in the closure, with `via` naming the declaring
|
||
module. Before item 93 the transitive query carried only the `READS`/`WRITES` branch, so asking the
|
||
*same* module with `depth` dropped its declared table and a Java caller's transitive `db-accesses` came
|
||
back empty although the entity it persists through maps to a real table.
|
||
|
||
**Natural view aliases are resolved to the underlying table (item 95).** A Natural DML statement names a
|
||
*view variable* (`1 VDB2-VERSIS_LITERALES VIEW OF VERSVW_LITERALES`), not the DDM. `db-accesses` reports
|
||
the **table** — `FIND VDB2-VERSIS_LITERALES`, `FIND NUMBER NEXT-VIEW` and `STORE VDB2-VERSIS_LITERALES`
|
||
in `YLITEMN0` all come back as `VERSVW_LITERALES`, matching the SQL `SELECT … FROM` rows in the same
|
||
module. Before item 95 the alias itself was the reported name, which (a) split one table across several
|
||
names, (b) made the generator's boilerplate alias `NEXT-VIEW` a single node shared by 11 modules meaning
|
||
11 different tables, and (c) hid every `VERSVW_LOGFILE` write behind 11 `VDB2-*-VLOG` aliases. Table
|
||
names are upper-cased (Natural is case-insensitive).
|
||
|
||
**…including aliases declared in a `USING` data area (item 98).** A view is often declared not in the
|
||
module but in a `LOCAL USING` area, in the data-area *export* form (`V 1VDB2-VERSIS_GENAGREE
|
||
VERSVW_GENAGREE …` — no `VIEW OF` text). Those resolve too: `YGEAGBNH`'s
|
||
`FIND (1) VDB2-VERSIS_GENAGREE` reports `VERSVW_GENAGREE`. Resolution is scoped to each module's own
|
||
`USING` set, never by name — alias names are boilerplate, and `NEXT-VIEW` alone is declared over 100
|
||
different tables in `upms`. A module whose `USING` areas give two different tables for one alias is left
|
||
unresolved rather than guessed.
|
||
|
||
**Natural `UPDATE(<label>.)` / `DELETE(<label>.)` count as writes (item 96).** These act on the current
|
||
record of the labelled `FIND`/`READ` loop, and are reported as `WRITES` on that loop's table. This is
|
||
what makes the `Y****MN0` access layer's update/delete path visible: `YLITEMN0` reports `WRITES
|
||
VERSVW_LITERALES` at the `STORE` **and** at `UPDATE(HOLD-PRIME.)` / `DELETE(HOLD-PRIME.)`, where before
|
||
item 96 it reported only the `STORE` — reading, wrongly, as an insert-only layer. A reference that
|
||
resolves to no labelled loop (an unknown label, or the numeric source-line form) records nothing rather
|
||
than guessing a table.
|
||
|
||
**`call-tree`/`graph` agree with `callees` about overridden dynamic calls (item 97).** A manual
|
||
dynamic-call override hides the placeholder marker rather than deleting it. All read paths now filter it,
|
||
so a pinned `CALLNAT <var>` shows the real target and never the variable name. Everything driven by the
|
||
call-tree BFS — `graph`, `db-accesses?depth=N`, `sql-statements?depth=N` — inherits this.
|
||
|
||
**`callers` on a dynamically-called module is an over-approximation, and says so.** A Natural web-service
|
||
module is reached by `CALLNAT #WIF`, resolved by naming pattern: `W-LST-N0.nat:362` alone resolves to 29
|
||
`W****B*S`/`W****X*S` targets, so `WGEAGB0S` lists `W-LST-N0` and `W-MNT-N0` as callers. The rows are
|
||
tagged `edgeKind: "CALLNAT_DYNAMIC"` — treat those as *may-call*, not *does-call*, and check
|
||
`dynamic-calls/overrides` / `dynamic-calls/unresolved` when the distinction matters.
|
||
|
||
## XML payload / interface schema (item 45)
|
||
|
||
Natural XML wrapper subprograms build a wire payload by mapping data-area fields to XML tags via the
|
||
`ADD-XML-LINE` idiom (`#W-TAG := '<tag>'` / `#W-VALUE := <field>` / `PERFORM ADD-XML-LINE`, where the
|
||
subroutine `COMPRESS`es `'<' #W-TAG '>' #W-VALUE`). The deep parser extracts that contract as
|
||
`PAYLOAD_FIELD` nodes and exposes it:
|
||
|
||
- `GET /modules/{name}/payload` (`ac payload <module>`) → an array of
|
||
`{tag, field, direction, lineNo, sourceFile}` triples. `direction` is `REQUEST` for an emitted
|
||
(outbound) field. `field` is the unqualified payload field name (`WXMLIN.P-COD-USUARIO` →
|
||
`P-COD-USUARIO`). **`sourceFile`** is the file `lineNo` refers to — the module's own file for
|
||
`source=IDIOM`, or the interface PDA's file for `source=PDA` (so a caller opens the right file at the
|
||
line, not the module at a stray line).
|
||
|
||
- **`source=IDIOM`** (item 45): extracted from a static `ADD-XML-LINE` emit sequence with literal tags.
|
||
- **`source=PDA`** (item 46b): the module is a *generic*, runtime-driven serializer (it calls the
|
||
`YFRAMN07` tag-builder or has an `ADD-XML-LINE`/`ADD-XML-ACT` subroutine) with **no** static tag list
|
||
in its source — real production wrappers like `WNAUTD0S` are this shape. The contract is then derived
|
||
from the module's `PARAMETER USING` **interface PDA**: each field is a payload field, the wire tag is
|
||
the field name with the framework's `EXAMINE … '#' REPLACE '_'` normalisation applied (`#`→`_`),
|
||
direction `REQUEST`. Idiom fields take precedence when both exist.
|
||
|
||
The static idiom also handles **derived tags** (`EXAMINE #W-TAG FOR '#' REPLACE '_'` → `#`→`_`) and
|
||
**both directions**: an `ADD-XML-LINE`-style emit sub is `REQUEST`; a `GET-XML-LINE`-style parse sub
|
||
(with the reverse `field := #W-VALUE` binding) is `RESPONSE`.
|
||
|
||
Empty for modules that neither use the idiom nor are a flagged XML wrapper (or are only
|
||
coarse-ingested).
|
||
|
||
## Copycode (`.cpy`) expansion (item 46a)
|
||
|
||
Natural `INCLUDE <member> <args>` is a compile-time macro: the copycode body is spliced into the
|
||
including module (with positional `&1&…` substitution), so a copycode's `CALLNAT`/`PERFORM`, DB access
|
||
and dataflow live in the copycode, not the module. The deep and coarse parsers now expand statement-level
|
||
copycode includes before parsing, so those constructs surface on the including module — e.g. a `READ`
|
||
or `CALLNAT` that only exists in a `.cpy` shows up in the host's `db-accesses`/`callees`.
|
||
|
||
- Copycode-origin nodes/edges report the **real `.cpy` file + line** (so navigation lands in the
|
||
copycode), and carry `viaCopycode=<member>` + `includedAt=<host line>`; host statements keep their own
|
||
file + line (line numbers are remapped after the splice, never shifted).
|
||
- **A line number alone is not a location (item 66).** Because of the above, one module's calls and field
|
||
accesses come from more than one file, and host and copycode lines are freely mixed — so always read the
|
||
line together with the file the endpoint gives it:
|
||
- `callers`/`callees`/`functions/{f}/callers` return **`sites: [{lineNo, callSiteFileIndex, viaCopycode,
|
||
includedAt, includePath}]`** (not a bare `lineNos` array). `callSiteFileIndex` indexes `sourceFiles` and
|
||
is the file the call is *written in*; the entry's own `sourceFileIndex` is a different thing — the file
|
||
the named module/function is *defined* in. `viaCopycode`/`includedAt`/`includePath` are set only for
|
||
copycode sites.
|
||
- `variables/{name}/reads|writes` return `sourceFile` = the file `lineNo` is in (the `.cpy` for a
|
||
copycode access), plus `viaCopycode` + `includedAt`.
|
||
**Item 109:** a `writes` row also carries `assignedValue` — the right-hand side *as written*:
|
||
`'WREQUD0S'` (literal), `*PROGRAM` (system variable), `#DISPLAY(1)` (indexed), `#SELECTED-KEY.NUM`
|
||
(qualified reference) — plus `assignedSubstrPos`/`assignedSubstrLen` for a `SUBSTR(...)` target
|
||
(item 83). It is deliberately not normalized to literals: that would drop most write sites. This is
|
||
what answers "which values does this module put into field X" without opening the source, the
|
||
dispatcher question item 108 keeps running into.
|
||
**`assignedValue` is always `null` for Java** — the Java parser does not capture the right-hand
|
||
side. Read it as "not captured for this language", not as "nothing is assigned". Natural carries it
|
||
on ~97% of its write edges. On a `reads` row it is always `null` by definition.
|
||
Before item 66 the copycode's line was paired with the host's file: `#W-OPTIONS` writes in `WGEAGB0S`
|
||
were reported at `WGEAGB0S.nat:18/20/22`, which is its generated comment banner — the writes are really
|
||
`ISICINDI.cpy:18/20/22`. Use `includedAt` when you want the spot in the host module instead.
|
||
- **`includedAt` is always a line in the module's own file, and `includePath` shows the whole chain
|
||
(item 104).** Natural `INCLUDE` nests, often through a positional argument
|
||
(`INCLUDE USIX050C 'YFRAMMC1'` → `INCLUDE &1&` → `INCLUDE YFRAMC01`), and only the *innermost* member
|
||
is named by `viaCopycode`. `includedAt` used to be the `INCLUDE` line in the *enclosing* `.cpy` — a
|
||
file the response never named — so `ISI173N0 → YFRAMN04` reported line 27, which is a comment in
|
||
`ISI173N0.nat` and in truth line 27 of `YFRAMMC1.cpy`. Now `includedAt` is the host-module line (232),
|
||
and `includePath: [{sourceFile, lineNo}, …]` lists every hop, host first, innermost last (empty for a
|
||
direct statement, one entry for a one-level include).
|
||
- `db-accesses` / `workfile-accesses` return **`sites: [{lineNo, sourceFile, viaCopycode, includedAt,
|
||
includePath}]`**
|
||
alongside the (kept, backward-compatible) `lineNos` array — one entry per statement, each tying its
|
||
line to the file it truly lives in. `sql-statements` gains **`sourceFile`** + **`viaCopycode`** on each
|
||
statement (its `startLine`/`endLine` are lines *in `sourceFile`*). Before this, a DB/work-file access
|
||
written in an `INCLUDE`d copycode reached the API as a bare copycode-local `lineNo` with nothing to
|
||
attribute it to — e.g. the DB2 sequence read `SELECT … FROM SYSIBM-SYSDUMMY1` lives in `USIX043C.cpy`
|
||
at lines 31/39/45/51/57, but `db-accesses` for the 9 including modules (YAPRFMN0, YUGRPMN0, …) reported
|
||
those as bare line numbers that land on the host's own comment/`DEFINE DATA` lines. The `sites` file
|
||
context is the same fix item 66 applied to `variables/reads|writes` and `callees`.
|
||
- **A copycode's nodes belong to the including module, not to the copycode (item 75-B, 2026-08-22).**
|
||
Until now every module that included a `.cpy` shared *one* set of nodes for its body. That is no longer
|
||
so: a copycode-resident node is keyed per including module (`ownerModule`), so what an agent sees changes
|
||
in one visible way — **counts go up, and they are now per-module**. A `READ` written in a copycode that
|
||
20 modules include is 20 access nodes, one per module, instead of one shared node; the same holds for a
|
||
`DEFINE SUBROUTINE` in a `.cpy` and for its control-flow statements. Read it as "each of these modules
|
||
really does perform this access", which is what the API always claimed but could not previously
|
||
represent. `search/identifier` for a name defined in a widely-included copycode therefore returns one hit
|
||
per including module — filter/group by `sourceFile` + the module you care about rather than expecting a
|
||
single row. Modules and DB tables are deliberately **not** per-module: a module declared inside a
|
||
copycode (`ZDTSTBP6` in `ZDTSTBC6.cpy`) and every `DB_TABLE` stay shared, so module lookups are unchanged.
|
||
**Item 75-C (same day) takes this one step further: identity is per *expansion site*, not per module.**
|
||
A copycode included several times by the same module (`JX0031N0.nat` includes `YFRAMBC0` 16 times) now
|
||
yields one set of nodes *per include site*, keyed by `includePath`. So counts rise again for those
|
||
modules, and — the point of the change — a copycode that opens a block it does not close no longer
|
||
collects every site's nesting into one node. `includedAt` alone does **not** identify a site (item 104
|
||
makes it the host's INCLUDE line at every nesting level, and `VPARTC02.cpy` includes `L4NLOGIC` 136 times
|
||
behind a single host line); use `includePath` when you need to tell two expansions apart.
|
||
Measured on `upms` after the recreate: copycode-resident nodes 27,551 -> 51,895 (project total +5.0%),
|
||
spread over 19,565 distinct owners. Endpoint latency on the heaviest module (`JX0030N0.nat`, 91 include
|
||
sites) is unaffected: `digest` 0.83 s, `context` 0.25 s, `graph` 0.19 s.
|
||
- **The same line number can legitimately appear twice (item 69).** A host statement on line 10 and a
|
||
copycode statement on line 10 are two different statements, and both are returned — as separate entries
|
||
differing only in their file. Until item 69 the graph could not hold both: an edge was identified by
|
||
`(source, target, type, lineNo)` with no file, so the second one **overwrote** the first and a real
|
||
access was missing from every answer. Treat `(file, lineNo)` as the identity of a site, never `lineNo`.
|
||
- **`call-tree`'s `depth` counts module hops (item 67).** `depth` is how many **module boundaries** the
|
||
shortest call path crosses, not raw `CALLS` edges — a `CALLNAT` made from two subroutines deep is still
|
||
one hop. The root module's own subroutines are therefore **depth 0**. Measured on `upms`: `WGEAGB0S`'s
|
||
seven direct dependencies used to report depth 1..3 (`BGEAGFN0` was 3); all seven now report 1.
|
||
- **Results at a given `depth` are larger than before.** A subroutine of a module within `depth` hops is
|
||
now inside the bound, because it crosses no further boundary. Previously `call-tree?depth=1` could hide
|
||
a `DEFINE SUBROUTINE` of *the very module you asked about*, just because it was `PERFORM`ed from
|
||
another subroutine (raw depth 2) — that is the same bug seen from the inside.
|
||
- `call-tree` also returns **`truncated`**. `true` means the *intra-module* subroutine walk stopped at
|
||
its raw-hop budget, so some `FUNCTION` items may be missing — **not** that your `depth` was exceeded
|
||
(that is a normal, complete answer). It is conservative and can be `true` for a complete result. Tune
|
||
via `agenticcode.call-tree.internal-budget` (default 20; the deepest internal chain observed in `upms`
|
||
is 9).
|
||
- **Since item 94 the budget cannot hide a module.** `MODULE` rows come from the same module-hop BFS
|
||
that `db-accesses`/`sql-statements` use, so a callee one hop away is always listed even when its call
|
||
site sits behind a long internal `PERFORM` chain (before item 94 it was dropped, and `call-tree` then
|
||
contradicted `db-accesses`). This also removed the path enumeration that made
|
||
`call-tree?followWiring=true` time out on Java projects at `depth ≥ 2`; `followWiring` is now usable
|
||
at full depth.
|
||
- **`field-flow`'s `depth` counts module hops (item 68).** Like `db-accesses`/`sql-statements` (item 65),
|
||
`variables/{name}/field-flow?depth=N` now means "up to N **module** calls apart", not N raw `CALLS`
|
||
edges. Before item 68 a consumer called from inside a subroutine sat several raw hops away and was
|
||
dropped at `depth=1`, so the endpoint answered "nothing downstream consumes this field" — read that
|
||
answer with suspicion on any graph ingested before this change.
|
||
- **`field-flow` no longer fabricates flows between same-named fields (item 77).** A bare field reference
|
||
is resolved against **the referencing module's** own `USING` includes. It used to be resolved
|
||
project-wide: an unresolved bare field is one shared node per `(name, project)`, and the resolver
|
||
aggregated over all owning modules at once, so (a) two modules including *different* data areas that
|
||
both declare the name left both unresolved on the shared node, and (b) a module with no matching include
|
||
was redirected onto another module's field. Either way the two modules ended up on one node, and
|
||
`field-flow` — which pairs a producer with a consumer only when both touch the *same* node — reported a
|
||
dataflow between modules that share nothing but a field name. In `upms`: 38 + 28 placeholders affected
|
||
(199 module-field pairs). **Read any pre-item-77 `field-flow` result for a common field name with
|
||
suspicion**, and note the answer only changes after a *deep* re-ingest, since resolution runs there.
|
||
`reads`/`writes` are unaffected — they match every node with the name and report only the accessing
|
||
side, so they never distinguished the targets in the first place.
|
||
*Residue (item 76):* a bare field shared via copycode (132 of 18,539 source nodes in `upms`) is still one
|
||
node for several modules; per-module identity is a schema change, not a query fix.
|
||
- **Copycode provenance survives field resolution (item 70).** `viaCopycode`/`includedAt` are now kept for
|
||
fields addressed by *qualified* name (`MYLDA.Q-FIELD`, i.e. a field of a `LOCAL USING` data area) as
|
||
well as bare ones. Before item 70 only bare references kept it; qualified ones silently came back with
|
||
`viaCopycode: null` and the host file, i.e. they looked exactly like host statements.
|
||
- Excluded from expansion: framework macros (handled by the item-44 targeted recogniser), data-area
|
||
`USING` includes, unknown members, and any copycode that declares `DEFINE DATA`. Recursion is
|
||
cycle-guarded. Copycodes (`.cpy`) are not standalone modules — they enter the graph only through the
|
||
including module.
|
||
- **Staleness caveat:** the item-41/43 hash check hashes the *host* file, so auto-invalidation triggers
|
||
on a change to the host — but a change to an included `.cpy` **alone** (host unchanged) is not
|
||
detected; re-ingest the host (`refresh/{host}`) to pick it up.
|
||
|
||
## Global Data Areas (`.gda`) (item 46c)
|
||
|
||
`.gda` files are now ingested as `DATA_STRUCTURE`s like `.lda`/`.pda`, and `DEFINE DATA GLOBAL USING
|
||
<gda>` resolves to them (the `INCLUDE`/`USING` recogniser now accepts `GLOBAL`, not just
|
||
`PARAMETER`/`LOCAL`).
|
||
|
||
## Deep-ingest: now automatic (lazy Tier-2)
|
||
|
||
Field-level endpoints (`flow-forward`, `flow-backward`, `field-flow`) and
|
||
cross-module dynamic `CALLNAT` resolution need a **per-module deep ingest**, not
|
||
just a whole-root refresh. This deep ingest is now **triggered
|
||
automatically on demand**: calling a field-level endpoint for a module that is
|
||
only `CALL_GRAPH`-ingested runs a scoped deep ingest of that module (and its
|
||
dependency tree) transparently, then returns the resolved result — no `409`,
|
||
no manual `POST /refresh/{name}` step. The first such call to a cold module is
|
||
therefore slower (it walks the root and parses the program tree); subsequent
|
||
calls hit the already-`FULL` graph.
|
||
|
||
The deep ingest is best-effort: if the module cannot be resolved to a source
|
||
file, the endpoint still falls back to the `409 NOT_DEEPLY_INGESTED` /
|
||
`NOT_INGESTED` hint with a `nextAction` rather than a misleading empty result.
|
||
|
||
**Flow path-ingest (auto, cross-module fixpoint).** `flow-forward`,
|
||
`flow-backward`, and `field-flow` go one step further than the single start-module
|
||
deep ingest: after deep-ingesting the start module they run an *ingest-and-re-traverse
|
||
fixpoint*. Each round deep-ingests the **frontier** — the modules the trace surfaced
|
||
*together with their direct callee modules* — in one scope, then re-traverses. This is
|
||
what lets a dataflow trace cross into a **dynamically-dispatched** callee (`CALLNAT
|
||
PGM-VAR`): that callee is not a static dependency of the start module, so it is only
|
||
pulled in and linked (`caller.arg → callee.param`) once a round resolves the dynamic
|
||
`CALLS` edge and ingests the target. The loop is bounded by
|
||
`agenticcode.deep-ingest.flow-rounds` (default 3) and the per-round
|
||
`agenticcode.deep-ingest.fanout-nodes` budget, and stops early (fixpoint) as soon as a
|
||
round pulls in nothing new — so on an already-deep graph a flow query costs one
|
||
traversal plus one cheap frontier check, no re-run.
|
||
|
||
**Fan-out warm (auto, on the result set).** The fan-out / traversal queries
|
||
`callers`, `/search/identifier`, and `call-tree` also auto-deep-ingest — but on
|
||
the **set of modules their result surfaced**, not a single named module. Each
|
||
runs against the graph as-is, deep-ingests the surfaced modules (blocking,
|
||
bounded by the fan-out node budget `agenticcode.deep-ingest.fanout-nodes`,
|
||
default 50), and — only if that warm actually deepened something — re-runs so
|
||
the response reflects newly-resolved dynamic dispatch (e.g. a `call-tree` grows
|
||
to include a dynamically-dispatched callee once the surfaced program is deep).
|
||
When everything is already `FULL` (or the warm resolves nothing) the query
|
||
returns its first result with no redundant re-run. Note `callers` warms the
|
||
*already-surfaced* callers, so it improves downstream precision but cannot
|
||
reveal a caller that was invisible at the coarse (call-graph) tier. Other
|
||
module-level endpoints (context, callees, db-accesses) work regardless of
|
||
ingest depth and do not trigger a deep ingest.
|
||
|
||
**Bounded fan-out.** A by-name deep ingest walks the transitive dependency tree
|
||
breadth-first, bounded by `maxDepth` (hops from the named module, default 5,
|
||
ceiling 20) and `maxNodes` (files, default 300). When a bound is hit the walk
|
||
stops early and the ingest response carries a `truncation` object
|
||
(`{reason: DEPTH|NODES|NODES_AND_DEPTH, maxDepth, maxNodes, hint}`) — the modules
|
||
actually reached are marked `FULL`, the remainder stays as it was. Raise the
|
||
limits on an explicit module refresh to pull in more:
|
||
`POST /refresh/{name}?maxDepth=&maxNodes=`, or CLI `ac refresh <name> --max-depth --max-nodes`. Auto-triggered ingests
|
||
use the
|
||
server defaults; if a field-level query returns partial data because the target's
|
||
deep ingest truncated, re-run the explicit refresh with higher limits. (The
|
||
auto-trigger does not yet accept per-query limit overrides.)
|
||
|
||
**Durable ingest status + coalescing (item 36).** Each real `MODULE` node carries a
|
||
durable `ingestStatus` lifecycle — `NOT_INGESTED` (only its call graph is in the
|
||
graph) → `INGESTING` (a deep ingest is in flight) → `INGESTED` (deeply ingested,
|
||
`ingestDepth = FULL`) — separate from `ingestDepth`. When two calls trigger the same
|
||
module's deep ingest at once they **coalesce** rather than both ingesting: within one
|
||
process an in-process lock serialises them; across processes/restarts a best-effort DB
|
||
claim marks the module `INGESTING` and a loser waits for the winner to reach `FULL`
|
||
(re-claiming if the claim is released or goes stale after
|
||
`agenticcode.deep-ingest.ingesting-ttl-seconds`, default 1800; wait bounded by
|
||
`claim-wait-seconds`, default 120). A crash mid-ingest leaves the module re-triggerable
|
||
(it never reached `FULL`), and the stale `INGESTING` is reclaimed on the next call.
|
||
`GET /nodes/{id}` / `/nodes/{id}/source` expose `ingestStatus`/`ingestStatusAt` on the module node.
|
||
|
||
**Warm concurrency cap (item 37).** All auto deep-ingest/warm work (by-name, fan-out,
|
||
and flow-frontier) shares a global permit pool
|
||
(`agenticcode.deep-ingest.max-concurrent-warms`, default 2), so a burst of queries
|
||
cannot spawn unbounded parallel parses/Neo4j writes. A permit is acquired only around
|
||
the actual ingest; if none frees up within
|
||
`agenticcode.deep-ingest.warm-acquire-timeout-seconds` (default 10) the warm is skipped
|
||
and the query returns its **Tier-1** (coarse) answer immediately rather than blocking —
|
||
so under sustained load a query may transiently return shallower data; retry once load
|
||
subsides, or force it with an explicit `POST /refresh/{name}`.
|
||
|
||
## OpenAPI contract & CORS (items 48/50)
|
||
|
||
The server now ships an OpenAPI 3 spec (`quarkus-smallrye-openapi`): the machine
|
||
contract the web-UI TypeScript client is generated against. All REST endpoints
|
||
carry `@APIResponse`/`@Schema` annotations, so response bodies are typed in the
|
||
spec even though the JAX-RS methods return raw `Response`. Access it at:
|
||
|
||
- `GET /q/openapi` — YAML (or `Accept: application/json` for JSON)
|
||
- `GET /q/swagger-ui` — interactive UI (dev)
|
||
|
||
CORS is enabled (`quarkus.http.cors.enabled=true`) and restricted to the UI's dev
|
||
origins (`http://localhost:5173`, `http://localhost:4173`) — extend the
|
||
`quarkus.http.cors.origins` list per deployment; never ship a wildcard.
|
||
|
||
## Endpoint quick reference
|
||
|
||
| Endpoint | Use for |
|
||
|----------------------------------------------------------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||
| `GET /modules?sourceFile=&moduleKind=&extends=` | List/filter modules; map a source file to its module name(s). Each row carries `loc`/`sloc` (item 46) and `ingestStatus`/`ingestDepth` (item 50) for status badges without a per-module round trip |
|
||
| `GET /loc?language=&sourceFile=` | Per-language LoC/SLoC rollup (`typescript`/`css` too, item 192) (fileCount/loc/sloc) + project total; each file counted once (item 46). For a generated/user_exit project also `userExitLoc`/`userExitSloc` + `generatedExclusiveLoc`/`generatedExclusiveSloc` (item 47) |
|
||
| `GET /modules/{name}/digest` | Tiny triage view before deciding which modules to expand |
|
||
| `GET /modules/{name}/context` | One-shot overview: functions, callers, callees, DB accesses, SQL/variable summaries (`?include=` for full lists) |
|
||
| `GET /modules/{name}/callers` \| `/callees` | Direct callers/callees incl. `EXTENDS`/`IMPLEMENTS`/`INJECTS`/`REFERENCES`. `callers` `scope`: **`external` (default)** = modules that call this one (CALLNAT/inheritance), **rolled up to the calling MODULE**: a call made from inside a subroutine/method is attributed to its owning module (never the calling `FUNCTION` node), and repeated call sites from one caller collapse to a single row whose `sites` list every line — symmetric with how `callees` anchors its source side. `internal` = the module's own subroutines' `PERFORM` wiring (function-level). The default is external-only, module-typed only, and never lists the module as its own caller (no `MODULE→MODULE` self-loop); use `scope=internal` or `/functions/{fn}/callers` for intra-module / function-level wiring. `callees` is unchanged (default lists both external CALLNAT and internal PERFORM targets) |
|
||
| `GET /modules/{name}/functions/{function}/callers` | **FUNCTION-level callers** (item 52): who `PERFORM`s (Natural) or calls (Java, TypeScript — same-module **and** cross-module, item 197) a specific subroutine/method, with call-site `lineNos`. Cross-module callers come from the module-to-module `CALLS` edge's `callerFn`/`calleeMethod`, matched by name (overloads over-approximate; a call from top-level code with no enclosing function shows only in the module-level `/callers`). Finer-grained than the module-level `/callers` (which is module→module). Same `CallRefResponse` shape. CLI `ac function-callers <module> <function>` |
|
||
| `GET /modules/{name}/call-tree?depth=` | Transitive call graph to scope a feature
|
||
| `GET /modules/{name}/reaches?target=A,B,C&direction=up\|down&depth=` | **Item 110 — "can A reach B, and how?"** Returns `{reachable, paths, truncated}` with one witness route per reached target (module names, source→target). `direction=down` (default): paths from this module to each target. `up`: paths from each target to this module. The counterpart to `call-tree`, which only walks downward and returns a closure without routes — one audit hand-rolled this as ~100 `/callers` requests. **`reachable: false` means "no path over known edges", not "no path"**: the traversal runs on resolved module calls, so a route through an unresolved dynamic `CALLNAT` (item 82) is invisible. Bounded by `depth` (item 75: the call graph has cycles). CLI `ac reaches <module> --target A,B --direction up` |
|
||
| `GET /duplicates` | **Item 114 — identities skipped at ingest** because they exist in more than one file (`{name, kind, paths}`, paths relative to the project root). These are *not* in `/modules`; asking for one by name gives `409 DUPLICATE_IDENTITY`. Their own calls are absent from the graph, so caller lists elsewhere can be short. CLI `ac duplicates` | |
|
||
| `GET /dynamic-calls/unresolved` \| `/overrides` · `POST`/`DELETE /overrides` | **Manual dynamic-`CALLNAT` overrides (item 82).** `unresolved` lists open `CALLNAT <var>` sites `{module, originFile, lineNo, variable}`; `POST /overrides {originFile, lineNo, targets[], variable?, note?}` pins a site to real module(s) (applied at once, persisted across refreshes, `400 UNKNOWN_TARGET` for a non-module); `DELETE /overrides?originFile=&lineNo=` resets one site (omit both = all) and restores the placeholder inline; `GET /overrides` lists them with an `obsolete` flag. CLI `ac dynamic-calls unresolved\|overrides\|set\|reset` |
|
||
| `GET /modules/{name}/graph?direction=&depth=&limit=` | Ego graph (item 49): bounded module-level call neighbourhood as **nodes + edges** (unlike call-tree). `direction` = `out`/`in`/`both`; `limit` caps nodes (BFS order) and sets `truncated`; unresolved targets carry `unresolved=true` + empty `sourceFile`. CLI `ac ego-graph` |
|
||
| `GET /counterparts?module=&kind=&unmatched=` (`ac counterparts`) | **Item 193:** this project's web-service calls / generated DTOs / fields with their twin in the counterpart project; `unmatched=true` = what nothing serves or mirrors yet |
|
||
| `GET /store?slice=` (`ac store`) | **Item 194:** the frontend Redux store — one row per slice (reducer key, RTK `sliceName`, `stateType`, the top-level state keys with type/optional/read/write counts, reducer count, total access sites) |
|
||
| `GET /store/{slice}/accesses?field=&mode=reads\|writes&module=` (`ac store-accesses`) | **Item 194:** who reads/writes a slice — reducers (`functionKind=reducer`, `via=reducer`) and the components/hooks/thunks selecting from it (`via` = `useAppSelector`, a wrapper hook, `getState`), with the full sub-`path` and line. Store fields also answer `variables/<slice>.<field>/reads\|writes` |
|
||
| `GET /bindings?dto=&field=&mode=reads\|writes&module=&partial=` (`ac bindings`) | **Item 195:** which component reads/writes which DTO field through the generated `Fields` path objects (`<SmartInput field={X.broker.ebene}>`), with the field's backend counterpart — `--dto Broker --field ebene --mode writes` = which page edits Java `Broker.ebene`. `data-structures/{dto}/fields` carries `boundReads`/`boundWrites` |
|
||
| `GET /theme?unused=` · `GET /theme/{token}/usages` (`ac theme`, `ac theme-usages`) | **Item 196:** the MUI theme's tokens (createTheme leaves + theme constants, value, `uses`; `declared=false` = read by the code but declared by no theme) and where one token is read (style block + CSS `property`, or plain `context`) |
|
||
| `GET /styles?module=&kind=sx\|style\|styled\|css&withLiterals=` (`ac styles`) | **Item 196:** the style inventory — every `sx`/`style`/`styled` block and CSS rule with CSS keys, hard-coded `literals` and the theme `tokens` it reads; `withLiterals=true` = what bypasses the theme |
|
||
| `GET /modules/{name}/db-accesses` \| `/sql-statements` | DB tables + mode, raw statement text (pass `?depth=` for Natural). **`db-accesses`/`workfile-accesses` return every row when no `limit` is given (item 103)** — they used to default to 50, and since the response is a bare array with no total and no `truncated` flag the cut was invisible: `WGEAGB0S?depth=10` returned 50 of 64 rows and hid 7 tables outright. An explicit `limit` is still honoured exactly. `db-accesses` items carry **`sites: [{lineNo, sourceFile, viaCopycode, includedAt}]`** (+ kept `lineNos`); `sql-statements` items carry **`sourceFile`** + **`viaCopycode`** — so a copycode-sourced access (e.g. `SELECT … FROM SYSIBM-SYSDUMMY1` in `USIX043C.cpy`) reports the `.cpy` line, not a bare number that reads as a host-file line |
|
||
| `GET /modules/{name}/workfile-accesses` | Natural **work files** (sequential/flat-file I/O — `READ`/`WRITE WORK FILE n`), the work-file analogue of `db-accesses` (item 84): `[{workFile, physicalName, mode: READS\|WRITES, recordBuffers, lineNos, sites}]`, aggregated per work-file number + mode. `sites: [{lineNo, sourceFile, viaCopycode, includedAt}]` gives each access its file context (copycode-aware), like `db-accesses`. `physicalName` comes from a `DEFINE WORK FILE n '<name>'`, else `null`. **Kept separate from `db-accesses`** — a work file is not an ADABAS/SQL table (fixes a former bug where `READ WORK FILE` created a phantom `DB_TABLE 'WORK'`). CLI `ac workfile-accesses <module>` |
|
||
| `GET /modules/{name}/data-structures` | Which copybooks/inline groups a module uses. A `USING <member>` binds by **member (file) name**, never by a level-1 record inside the file (item 100) — before that, `WGEAGB0S USING W-WIF-A2` reported `old/W-WIF-A7.pda` (whose level-1 record is a copy-pasted `1W-WIF-A2`), and a data area with several level-1 records and none named after the member (`VLAYERLA.lda`, `USIX020L.lda`) resolved to nothing at all (`sourceFile: null`, `area: UNKNOWN`, `fieldCount: 0`) although the file was ingested. One row per resolved definition, `(name, sourceFile)` (item 102) — never one row blending an arbitrary file with another definition's `fieldCount` |
|
||
| `GET /modules/{name}/payload` | Natural XML wire-payload contract: `{tag, field, direction, source, lineNo, sourceFile}` — static `ADD-XML-LINE` idiom (`source=IDIOM`, item 45) or derived from the wrapper's interface PDA (`source=PDA`, item 46b). `sourceFile` is the file `lineNo` refers to (module for IDIOM, PDA for PDA) |
|
||
| `GET /modules/{name}/comments?kind=&limit=&offset=` (`ac comments`) | **Item 141:** the module's **comment blocks** — `{text, kind, sourceFile, startLine, endLine, target, targetType, truncated}`, one row per contiguous block, ordered by line. `target`/`targetType` name the declaration the block documents: the declaration immediately below it, else the one enclosing it (so a file header banner documents the `MODULE`, a `/*` comment on a field's own line documents that field). `kind` is `JAVADOC`\|`LINE`\|`BLOCK` (Java) or `NATURAL_BANNER`\|`NATURAL_INLINE`\|`SAG` (Natural); `?kind=` filters to one. **`SAG` is excluded by default** — `**SAG` directives are generator metadata, not human notes, and would otherwise be most of the answer for every generated Natural module. Text is cut at 4 000 chars (`truncated:true`); read the file for the rest. Natural copycode comments belong to the **copycode's own module**, not to each includer. **Deep-gated:** comments come from the full parse, not the Tier-1 coarse scan, so the module is deep-ingested on demand and a still-shallow module answers `409 NOT_DEEPLY_INGESTED` rather than a misleading `[]` |
|
||
| `GET /modules/{name}/dispatch-table` | Natural `DECIDE ON VALUE OF` routing table |
|
||
| `GET /modules/{name}/functions?kind=` \| `/functions/{fn}/overrides` \| `/functions/overrides` | Method list, modifier filter (Java), subclass overrides (single/bulk). Each item carries **`sourceFile`** + **`viaCopycode`** (item 84): a Natural subroutine pulled in via `INCLUDE` reports the **copycode** file and `viaCopycode:true`, so its `startLine`/`endLine` are read as offsets into that copycode — **not** into the including module's own file (which is shorter). `viaCopycode:false` = declared inline. Always `false` for Java |
|
||
| `GET /data-structures/{name}/fields` \| `/db-tables/{name}/columns` \| `/modules/{name}/columns` | Field/column schemas for DTO/entity generation. Every field carries **`sourceFile`** (item 101). When a structure name resolves to several definitions (42 level-1 names recur across `upms` data areas), the **member root** — the definition whose file basename equals the name, i.e. what a `USING <member>` binds to — wins; **`?sourceFile=`** pins a specific one. Before item 101 the definitions were silently unioned: `W-WIF-A2` returned 15 fields, the merge of `W-WIF-A2.pda` (5) and `W-WIF-A7.pda` (10), a layout that exists nowhere |
|
||
| `GET /variables/{name}/reads` \| `/writes` \| `/flow-forward` \| `/flow-backward` \| `/field-flow` | Impact analysis and dataflow tracing |
|
||
| `GET /search/identifier` \| `/search/value` \| `/search/annotation` | Cross-project lookup by name / literal value / annotation. **All three see code only by default — an empty result is not evidence that the string is absent.** `search/value` takes **`includeComments=true`** (CLI `--include-comments`, item 141) to search comment blocks as well; those hits come back as `kind: "COMMENT"`, so a comment is never read as code. It is opt-in because a comment hit is different evidence from a literal, and folding it in silently would move every existing completeness count (item 131). `search/identifier` and `search/annotation` never match comments at all — use `includeComments`, `/modules/{name}/comments` or `/search/source` before concluding "not present" (see "Comments: reachable, but never by default" above). `search/identifier` matches the **exact** declared name but is **sigil-insensitive**: a leading Natural sigil (`#` user, `&` AIV, `+` GDA) is ignored on both sides, so `name=K-OUT-MAX` finds the declared `#K-OUT-MAX` (and vice-versa). **Item 125:** a Java **type declaration** is matched by its **short name** as well as by the fully-qualified identity the graph stores (item 117) — `name=PartnerUpdateLogic` finds `com.example.PartnerUpdateLogic`; before this it answered `[]`, which reads as "no such name". Every match carries `simpleName` and `moduleKind` (`CLASS`/`INTERFACE`/`ENUM`/`RECORD`, `PROGRAM`/`SUBPROGRAM` for Natural), both `null` for non-`MODULE` hits — so "is this name a type or a method?" needs no second call. `contains=true` (CLI `--contains`) switches to a case-insensitive **substring** match, as on `/search/value`; it was previously accepted and silently dropped. It matches the FQN too, so a package fragment also hits — filter with `type=MODULE`/`moduleKind` if that is noise. `contains` without a `name` is `400 MISSING_NAME` (a substring search for nothing is a full node dump). The substring scan is **unindexed**: it is bounded to `offset+limit` rows, so keep a `limit` on large projects. Optional **scope** filters `sourceFile=<relpath>` and `module=<name>` (item 53) narrow the match to one file / one module — use them to pinpoint a module-local declaration when a name recurs across dozens of modules (the result is otherwise paginated and the local one may fall off the page). To keep the **full cross-project list** yet still guarantee a given module's own declaration is on the first page, pass `priorityModule=<name>` instead of `module=`: it does not filter, but pins that module's matches to the front (ahead of the otherwise `sourceFile`-ordered rest) so they survive the `limit`. This is what the web UI's click-to-identify sends for the open module. CLI `ac search-identifier --module --priority-module --source-file --type --contains` accept the same filters. **Latency (item 105):** a lookup whose hits lie in a Natural data area used to take 60-75 s — every fan-out query deep-ingested the surfaced `.lda`/`.pda`, which can never reach `FULL` (a data area yields no `MODULE` node), so it was re-warmed on every call and each warm dragged a whole-project finalize behind it. Data areas are now excluded from the fan-out warm; they have no deep tier to gain |
|
||
| `GET /search/source?regex=&limit=&ignoreCase=` (`ac search-source`) | Regex **grep over module source text** (item 54): `{module, sourceFile, lineNo, line}` hits + `truncated`. Case-insensitive by default. Complements `/search/identifier` (declared names) — use for code patterns (statements, table names, literals). Sees **everything in the file, comments included**, and needs no ingest depth — so it is the fallback when a module is not deeply ingested, or when the text is something the parsers do not model. For comments specifically, prefer the graph routes added by item 141 (`/modules/{name}/comments`, `search/value?includeComments=true`), which also tell you which declaration a comment belongs to |
|
||
| `GET /nodes/{id}` | Every property of one node (when a curated DTO is missing something) |
|
||
| `GET /nodes/{id}/source` \| `/modules/{name}/source` \| `/source?file=` | Source text — **only when you have no other access to the source** (you always do in this repo, see "Reading source in this repo" above). `/modules/{name}/source` returns the **whole file** when the line range is omitted (M1), or a `[startLine,endLine]` slice when both are given. `/source?file=<relpath>` (CLI `ac file-source`) serves a file by **relative path** rather than module name — for files that aren't standalone modules, e.g. a Natural data area (PDA/LDA) USING'd by a module, whose field line numbers refer to that file. Same whole-file/range + stale-source semantics; the client-supplied path is rejected (`400 INVALID_SOURCE_FILE`) if it escapes the project root |
|
||
|
||
Full endpoint list, request params, and response field details:
|
||
`x-docs/agent-api-system-prompt.md`.
|
||
|
||
## Errors are structured JSON — always (item 136)
|
||
|
||
Every failure now answers `{ "error": ..., "code": ..., "details": {} }`, including the ones nobody
|
||
planned for: an unhandled exception is mapped to `500 INTERNAL_ERROR` with an `errorId` in `details`
|
||
that matches the stack trace in the server log (the trace itself is never in the response). Before
|
||
this, an unexpected fault escaped as a plain-text Quarkus error page with no `code` to branch on —
|
||
which is exactly the moment a client most needs a machine-readable answer. Deliberate statuses
|
||
(`PROJECT_NOT_FOUND`, `MISSING_NAME`, `STALE_SOURCE`, the runtime's own routing 404s) pass through
|
||
unchanged.
|
||
|
||
The fault that exposed this: `search/identifier` coerced `startLine`/`endLine` unconditionally, and
|
||
item 114's duplicate markers were the one kind of node created without them, so any page long enough
|
||
to reach a marker (row 487 on `ac`) died. Both halves are fixed — the markers now carry lines, and the
|
||
row mapper no longer trusts that they will.
|
||
|
||
## Verifying the API after a deploy (item 134)
|
||
|
||
After `./manage-ac.sh deploy` (or `./rebuild-and-refresh.sh`), run:
|
||
|
||
```bash
|
||
./x-scripts/verify-api.sh # defaults to project 'ac'
|
||
./x-scripts/verify-api.sh -p upms # any ingested project
|
||
AC_SERVER_URL=http://host:8787 ./x-scripts/verify-api.sh
|
||
```
|
||
|
||
It answers one question in ~10 s: *does the server that is running right now still return
|
||
plausible data over the real graph?* Exit 0 = all green, 1 = at least one check failed. Every line is
|
||
`PASS`, `FAIL` or `SKIP`; `SKIP` means the endpoint family does not apply to that project (a pure
|
||
Natural project has no `rest-endpoints`, a leaf module has no callees).
|
||
|
||
What it covers: `/api/version` and the project list; the item-130 scope headers
|
||
(`X-AC-Exclude-Dirs`, `X-AC-Ingested-At`, `X-AC-Ingest-Incomplete` — a `true` there means a refresh
|
||
was aborted and every later answer is drawn from a half-updated graph); the item-131 paging contract
|
||
on the search endpoints; per-family data plausibility; and the structured-error negative cases.
|
||
|
||
What it is **not**: a substitute for `mvn test`. The integration tests pin semantics; this pins
|
||
"the deployed thing is not obviously broken". A green run is not a quality gate. All assertions are
|
||
invariants, never fixed counts — counts move with every refresh.
|
||
|
||
## TypeScript / React projects (item 192)
|
||
|
||
A project may declare `language: typescript` (`ac project create purfe --root … --language typescript`).
|
||
Its `.ts`/`.tsx` files (not `.d.ts`) and plain `.css` files are ingested; `node_modules` and `dist`
|
||
are excluded by default. **Only a `typescript` project ingests TypeScript** — a Java project with a
|
||
bundled web UI (`ac` has `ac-ui/`) never parses it. Java and Natural files stay language-agnostic.
|
||
|
||
**Identities are paths.** A TypeScript `MODULE` is named by its root-relative path without the script
|
||
extension — `pur-r-vstamm/src/store/slices/agstammSlice` — with `simpleName` = the file stem
|
||
(`agstammSlice`), `workspace` = the first path segment, `moduleKind` = `ts`/`tsx`/`css`, and
|
||
`generated=true` + `generator` (`typescript-generator` for the Java-side EndpointGenerator output
|
||
under `generated/`, `hey-api` for `@hey-api/openapi-ts`). A CSS module keeps its extension
|
||
(`pur-ui/src/index.css`). Every `/modules/{name}/…` endpoint accepts either form (item 117), so
|
||
`ac context agstammSlice` works — until two workspaces have a file with the same stem, then use the
|
||
path.
|
||
|
||
**Two tiers, like Java/Natural.** Project creation and `refresh` without `--deep` run the **Tier-1
|
||
regex outline** in Java: module shell (`sourceHash`, `loc`/`sloc`), one `FUNCTION` per top-level
|
||
function / arrow / class (`kind` = `function` | `component` | `hook` | `thunk` | `styled` | `class`,
|
||
`exported`), one `DATA_STRUCTURE` per `interface`/`type`/`enum` (`dataType` says which), and a
|
||
`REFERENCES` edge per import to the module it resolves to (`value` = the import clause, `specifier`
|
||
= as written). npm packages are not placeholders; they are listed on the module as
|
||
`externalImports`. **A deep pass** (`refresh --deep`, `ingest` by name) runs the **Node sidecar**
|
||
(`ac-parser-typescript/sidecar/extract.mjs`, TypeScript compiler API, one whole-program run per npm
|
||
workspace, 3–5 s and ~0.5 GB each on the pur frontend) and replaces the outline with the checker's
|
||
view: exact positions, imports resolved against the file system, and **calls**:
|
||
|
||
- a callee owned by a top-level declaration of the same file → `FUNCTION -CALLS-> FUNCTION`;
|
||
- a callee in another module → `MODULE -CALLS-> MODULE` carrying `callKind` (`METHOD_CALL`, or
|
||
`CONSTRUCTOR` for `new`), `callerFn` (the calling function, absent at module level), `calleeMethod`
|
||
(the owning top-level declaration in the target), `callSyntax` (`call`/`new`/`tagged`/`jsx` — a JSX
|
||
element `<HistorieDrawer/>` is a call), and on a member call `receiver` (type of the innermost
|
||
object, e.g. `AgstammControllerEndpoint`) and `member` (`saveBroker.post`);
|
||
- calls into npm packages and the language library are **not** edges.
|
||
|
||
These are the exact properties the Java parser writes, so `callers`/`callees`/`call-tree`,
|
||
`functions/{fn}/callers` (item 52) and the placeholder rewiring work unchanged. A module that got its
|
||
Tier-2 pass carries `ingestTier=2`.
|
||
|
||
**Sidecar failure is visible, not silent.** If `node`, the script or its `node_modules/typescript`
|
||
are missing, or a workspace run fails or times out, the refresh still completes at Tier-1 for those
|
||
files and the response lists a failure with the pseudo-path `sidecar` or `sidecar:<workspace>`.
|
||
Config: `agenticcode.typescript.node`, `.sidecar-script`, `.max-heap-mb` (1024), `.timeout-seconds`
|
||
(600); the image carries node and the sidecar (`Dockerfile.jvm`), dev mode expects
|
||
`npm ci` run once in `ac-parser-typescript/sidecar/`. The project root is read-only in the
|
||
container; the sidecar reads the project's own `node_modules` for library typings and writes nothing.
|
||
|
||
**Scope of the pur frontend project.** The registered project covers the `pur-ui` and
|
||
`pur-ui-common` workspaces only; `pur-r-vstamm` and `pur-r-vbuch` are excluded via `excludeDirs`,
|
||
and the sidecar does not load an excluded workspace. Both workspaces call the `pur` backend through
|
||
the legacy generated client (`generated/endpoints.ts`, backend `pur`); the hey-api client shape is
|
||
recognised too but is not in scope.
|
||
|
||
**Resolution notes (verified on `purfe`, 2026-09-22).** An import of a workspace consumed through
|
||
its package.json `exports` resolves into its build output (`pur-ui-common/dist/x.d.ts`); the sidecar
|
||
maps that to the source twin (`pur-ui-common/src/x.ts`) so the edge lands on a real module. A bare
|
||
specifier the checker resolves to neither a file nor a package (`immer`, `redux` — transitive
|
||
dependencies the project does not list) is recorded as an external import, not a placeholder.
|
||
|
||
Transitive packages are known from `root/node_modules` (directory names), so Tier-1 treats them
|
||
as external too.
|
||
|
||
**Stale parsed edges are reaped on a deep refresh (item 198).** Every edge the parser emits is
|
||
stamped with the run's `ingestGen` at merge time; after a deep re-parse of a file, the edges from
|
||
that file's nodes whose stamp is older than the run's (the fresh parse did not re-emit them) are
|
||
deleted before the node sweep, for every language and edge type. An import or call the new parse
|
||
names differently (a renamed class, a `dist`→`src` mapping fix) therefore no longer keeps its old
|
||
placeholder alive next to the fresh edge, and a placeholder left edgeless falls to the usual
|
||
placeholder sweep. Tier-1 (`changedOnly` or non-deep) refreshes do not reap, because a Tier-1 pass
|
||
emits fewer edges than a deep one. Edges persisted before the stamp was introduced carry no
|
||
generation and are never reaped; the first deep refresh after upgrading stamps them, the next one
|
||
reaps — so a project never needs recreating after a parser change any more, two deep refreshes do.
|
||
|
||
**A module node never lives in a copycode (item 201).** A program whose body is a single `INCLUDE`
|
||
used to get a second `MODULE` node named after it with the `.cpy` as `sourceFile` (`ZDTSTBP6` in
|
||
`upms`), so `modules` and `search/identifier` listed the name twice. Fixed in the parser; a graph
|
||
ingested before the fix keeps the stray node until the project is recreated, or you remove it by hand:
|
||
`MATCH (m:MODULE {project: $p}) WHERE m.sourceFile ENDS WITH '.cpy' AND m.ownerModule = '' DETACH DELETE m`
|
||
(Natural copycodes are never modules of their own, so the match is exact).
|
||
|
||
**Known limits of 192** (the later items fill them): no field bindings (195), no styles (196);
|
||
the store is item 194 below. A `changedOnly` refresh re-runs the sidecar over the
|
||
whole workspace but re-persists only the changed files, so an unchanged file's facts can lag one
|
||
refresh (same class of caveat as 46a). The by-name deep ingest resolves dependencies by file stem, so
|
||
a dependency whose stem exists in several workspaces (`index`) is reported as a duplicate and skipped
|
||
— use `refresh --deep` for the whole frontend.
|
||
|
||
## Web-service calls and the counterpart link (item 193)
|
||
|
||
**`rest-endpoints` lists the frontend's calls.** Every member of a generated Endpoint class
|
||
(`AgstammControllerEndpoint.saveBroker`) and every hey-api sdk function is a `FUNCTION` of
|
||
`kind=endpoint` carrying the same `restPath`/`httpMethod` the Java parser writes for a handler, plus
|
||
`outbound=true` — so `GET /projects/purfe/rest-endpoints` answers with `outbound: true` rows
|
||
(`handler` = `AgstammControllerEndpoint.saveBroker`, `path` = `/agstamm/ui`). Who calls it:
|
||
`modules/{generated module}/callers` names the calling slices/components (module level), and the
|
||
member call `api.saveBroker.post(...)` is retargeted from the generic `PostMethod.post` signature to
|
||
the endpoint function, so the module-to-module `CALLS` edge carries `calleeMethod =
|
||
AgstammControllerEndpoint.saveBroker` and `callerFn = <thunk>`. Since item 203 the synthetic
|
||
class-hierarchy edges (`resolvedVia: INHERITANCE`, caller → each implementation of the called
|
||
interface/base) exist once per originating call site with its real `lineNo`, `originFile`,
|
||
`calleeMethod` and `callerFn` — before, one edge per pair took whichever call line the merge met first.
|
||
So `callees` lists every real line for an implementation, and `functions/{impl-method}/callers` also
|
||
names callers that go through the interface (`RepoImpl.save` ← `Service.store` via `Repo.save`).
|
||
Since item 197
|
||
`functions/{fn}/callers` joins these module-to-module edges back to the calling function, so
|
||
`purfe/modules/pur-ui/src/generated/endpoints/functions/GeneralAgreementUiControllerEndpoint.createNew/callers`
|
||
names the thunk in `generalAgreementSlice`, and on the Java side
|
||
`pur/modules/…AgstammLogic/functions/handleMerge/callers` names `AgstammController.mergeBroker`
|
||
(a REST controller method itself has no Java callers — it is the HTTP entry point). Extra properties on the node (
|
||
`GET /nodes/{id}`): `restUrl` (as composed,
|
||
with placeholders and query string), `restBase` (an application base such as `/pur-r-vbuch/v1`
|
||
split off so paths compare with the backend's base-less `@Path`), `backend` (`pur`, `pur-r-vstamm`,
|
||
`dynamic` — from the URL builder), `queryParams`, `requestType`, `responseType`, `paramsType`,
|
||
`generator`, `owner`, `member`. A generated interface's properties are `FIELD`s under its
|
||
`DATA_STRUCTURE`, **named `<Interface>.<member>`** (`Broker.ebene`; props `field` = the bare member,
|
||
`owner` = the interface — since item 195: a node's identity is type + name + file, and one generated
|
||
file declares hundreds of interfaces, so a bare `vid` used to be a single node under six interfaces),
|
||
so `GET /data-structures/AgstammUseCase/fields` answers for the frontend too (bare member names;
|
||
since item 195 — before, the query filtered `FIELD` out and returned `[]` for an interface), and
|
||
`counterparts?kind=field` rows are named `Broker.ebene`.
|
||
|
||
**`COUNTERPART_OF`: the same thing in another project.** A project setting
|
||
`counterparts: ["pur"]` (`ac project create purfe … --counterpart pur`, `ac project update purfe
|
||
--counterpart pur`, `GET /projects/purfe` shows it) makes enrichment link, after every refresh of
|
||
either side:
|
||
|
||
| this project | → counterpart | matched on |
|
||
|-----------------------------------------------|------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||
| outbound endpoint `FUNCTION` | backend handler `FUNCTION` | `httpMethod` + path shape (every `{param}` segment compares as `{}`, class + method `@Path` composed as `rest-endpoints` does); several matches → the handler whose source lives under the frontend's `backend` name |
|
||
| `DATA_STRUCTURE` in a `generated=true` module | Java `MODULE` with the same `simpleName` | unique name only; an ambiguous name stays unlinked |
|
||
| its `FIELD`s | the class's `FIELD`s | name |
|
||
|
||
The edges are rebuilt from scratch on each run (never accumulated) and re-run for the frontend when
|
||
the **backend** refreshes, because a refreshed handler node is deleted together with the edges
|
||
pointing at it. Roadmap item 143 (Natural ↔ Java counterparts) will use the same edge.
|
||
|
||
`GET /projects/{p}/counterparts?module=&kind=rest|dto|field&unmatched=&countOnly=&limit=&offset=`
|
||
(CLI `ac counterparts [--module] [--kind] [--unmatched] [--count-only]`) lists
|
||
`{kind, name, module, sourceFile, startLine, httpMethod, path, counterpartProject, counterpartName,
|
||
counterpartModule, counterpartSourceFile, counterpartStartLine}`; the `counterpart*` fields are null
|
||
for an unlinked row, and **`unmatched=true` is the planning question**: which calls does nothing
|
||
serve, which generated DTOs / fields have no backend twin. Paged with `X-AC-Total-Count` /
|
||
`X-AC-Truncated` like the search endpoints; unknown `kind` → `400 KIND_UNSUPPORTED`; a project
|
||
naming itself as counterpart → `400 COUNTERPART_SELF`.
|
||
|
||
## The Redux store (item 194)
|
||
|
||
**What is modelled.** Every `createSlice` / `createAppSlice` in a `typescript` project is a
|
||
`STORE_SLICE` node **named by the reducer key it is mounted under** in the project's
|
||
`configureStore` (`state.<key>`; `gruppenprovision` for `generalAgreementSlice`, whose RTK name is
|
||
`generalAgreement`) — the key is what every selector path starts with, so it is the identity; the
|
||
RTK name is kept as `sliceName`. The sidecar traces each `reducer: { key: xReducer }` entry through
|
||
`xReducer = xSlice.reducer` (also `export default xSlice.reducer`) back to the slice, across
|
||
workspaces; a slice no store mounts is named by its own `sliceName`. Under the slice, one **store
|
||
`FIELD` per top-level state key**, named `<key>.<field>` (`schluesseltabelle.sucheStatus`, from the
|
||
checker's type of `initialState`, so keys only the state type declares are present too) with
|
||
`dataType` = the TS type, `optional`, `store=true`, `slice`, `field`. Every reducer is a **`FUNCTION`
|
||
of `kind=reducer`** in the slice's module, named by the **action type it handles**:
|
||
`schluesseltabelle/updateX` for a case reducer (`reducerKind=reducer`), `schluesseltabelle/suche/fulfilled`
|
||
for `builder.addCase(sucheByServer.fulfilled, …)` (`reducerKind=case`, `trigger` = the expression),
|
||
`schluesseltabelle/matcher:isSlicePending(sliceName)` for `addMatcher` (`reducerKind=matcher`). A
|
||
thunk whose lifecycle action a case handles `CALLS` that case (`callSyntax=extraReducer`), when both
|
||
live in the same file.
|
||
|
||
**Reads and writes.** Inside a reducer every `state.a.b` chain is a `READS`/`WRITES` edge from the
|
||
reducer `FUNCTION` to the store `FIELD` of its first key (`state` alone → the `STORE_SLICE`), carrying
|
||
`path` (the full sub-path as written, `keyTableUseCaseSvcResult.result.tableId`), `lineNo`,
|
||
`via=reducer`. A write is an assignment target (compound assignments also read), `++`/`--`, `delete`,
|
||
a mutating method on the chain (`push`, `splice`, `set`, `delete`, …), `Object.assign(state.x, …)`,
|
||
or a `return { … }` (each key written) / `return other` (the whole slice). Outside reducers every
|
||
store read is a `READS` edge from the reading function (component, hook, thunk; the module when at
|
||
top level) to a **placeholder** `<key>.<field>` that the finalize step `resolve-store-placeholder`
|
||
redirects onto the real field by name and then drops — so a read in `KeyTablePage.tsx` lands on the
|
||
field declared in `keytableSlice.ts` without either file knowing the other. Recognised read forms:
|
||
`useSelector`/`useAppSelector((state) => state.a.b)` (every chain rooted at the arrow's parameter;
|
||
`const { x, y } = useAppSelector((s) => s.a)` reads `a.x` and `a.y`), **wrapper hooks** such as
|
||
`useSchluesseltabelleSelector((useCase) => useCase?.result?.purMode)` — the sidecar finds the wrapper's
|
||
inner `useAppSelector((state) => selector(state.a.b))`, so the read is `a.b.result.purMode` with
|
||
`via=useSchluesseltabelleSelector` (the wrapper's own inner read is recorded too) — and
|
||
`store.getState().a.b` / `const s = thunkAPI.getState(); s.a.b` chains (`via=getState`; 81 sites on
|
||
`pur-ui`). A read of a key the store does not declare stays a placeholder and is listed by
|
||
`search/identifier` with an empty `sourceFile`.
|
||
|
||
**Dispatch.** A call to a slice action creator (`dispatch(updateX(…))`, a binding of
|
||
`xSlice.actions`) carries `actionType = <sliceName>/updateX` and is a cross-module `CALLS` edge to the
|
||
slice module with `calleeMethod = <sliceName>/updateX` — the reducer `FUNCTION` — so module
|
||
`callees` of a component list the slices it dispatches into; a thunk call keeps the thunk `FUNCTION`
|
||
as target and carries the thunk's type prefix as `actionType`. (Function-level exposure of these
|
||
cross-module edges is item 197.)
|
||
|
||
**Endpoints.** `GET /projects/{p}/store?slice=` (CLI `ac store [--slice]`) → `[{slice, sliceName,
|
||
module, sourceFile, startLine, endLine, stateType, fields: [{name, type, optional, reads, writes}],
|
||
reducers, reads, writes}]`, all slices or one by reducer key / RTK name.
|
||
`GET /projects/{p}/store/{slice}/accesses?field=&mode=reads|writes&module=&countOnly=&limit=&offset=`
|
||
(CLI `ac store-accesses <slice> [--field] [--mode] [--module] [--count-only]`) → `[{mode, slice,
|
||
field, path, function, functionType, functionKind, module, sourceFile, lineNo, via}]`, one row per
|
||
access site, `field` null for a whole-slice access; paged like `counterparts`; unknown `mode` →
|
||
`400 MODE_UNSUPPORTED`, unknown slice → `404 SLICE_NOT_FOUND`. Because store fields are `FIELD`
|
||
nodes with a unique name, the generic `GET /variables/<key>.<field>/reads|writes` and
|
||
`search/identifier?name=<key>.<field>` answer too.
|
||
|
||
**Limits.** Reads through `getState()` aliases are followed only inside the file that created the
|
||
alias; a thunk→case `CALLS` edge is emitted only when thunk and slice share a file (the pur slices
|
||
do); a case whose trigger cannot be folded to an action type (a predicate matcher) is named by its
|
||
expression text; the sidecar reads the store of every workspace in scope — a key mounted only by a
|
||
workspace outside the project (`excludeDirs`) falls back to the slice's own name (matches on `pur`).
|
||
No `USES_TYPE` from a store field to the DTO it holds yet — the field's `dataType` says
|
||
`SvcResult<KeyTableUseCase>`, the link is item 195. Only a **top-level** `createSlice` is a slice: a
|
||
slice built inside a factory function (`pur-ui-common`'s `filetransferSlice`, created per instance
|
||
and not mounted in the `pur-ui` store) is not modelled.
|
||
|
||
**Verified on `purfe` (2026-09-22, server 318, recreated + `refresh --deep`, 285 files, 0
|
||
failures, 0 placeholders left):** 9 slices — `error`, `global`, `healthTables`, `metadata`
|
||
(pur-ui-common) and `gruppenprovisionSuche`, `gruppenprovision` (RTK name `generalAgreement`),
|
||
`schluesseltabelle`, `multilinguism`, `translationdata` (pur-ui) — 43 reducer functions, 216 read
|
||
and 82 write sites. `schluesseltabelle` alone: 62 sites, 25 through `useSchluesseltabelleSelector`,
|
||
19 through `getState()`, 5 through `useAppSelector`, 13 in reducers.
|
||
|
||
## DTO field bindings (item 195)
|
||
|
||
**What a binding is.** The generator emits, next to every DTO interface, a `Fields` class tree
|
||
(`AgstammUseCaseField = new AgstammUseCaseFields<AgstammUseCase, never>()`, members
|
||
`broker = new BrokerFields<TRoot, Broker>(this, "broker")`, list members as `keyTableList = (index?) =>
|
||
new KeyTableDOFields(...)`). A path expression on it — `AgstammUseCaseField.broker.ebene` on a
|
||
`<SmartInput field={…}>`, `<SmartOutput field={…}>`, a table's `fieldTermForRowData`, a column's
|
||
`field:`, or a `Fields`-typed prop such as `useCaseFieldPrefix` — is a binding. The sidecar types
|
||
every hop through the checker (`XFields<TRoot, TSelf>`): the **root DTO** is `TRoot`, the **owner**
|
||
of the leaf is the `TSelf` of the hop before it, the leaf name is the field. Each binding becomes a
|
||
`READS` edge (plus a `WRITES` edge when the component tag matches `Input$|Dropzone$|Editor$`) from
|
||
the binding function (component/hook; the module at top level) to the **`FIELD` of the generated
|
||
interface** that item 193 already creates (`Broker` → `ebene`), carrying `path` (dotted hops from the
|
||
root, `brokerList[]` for a list hop), `rootDto`, `kind`, `partial`, `component`, `attribute`,
|
||
`via=binding`, `lineNo`.
|
||
|
||
- `kind=field`: the leaf is a scalar (`ebene`); `kind=prefix`: a whole sub-object is handed on
|
||
(`useCaseFieldPrefix={X.tab.translationData}`, `fieldTermForRowData={X.keyTableList()}`) —
|
||
recorded as a read of the container field so nothing is silently dropped.
|
||
- `partial=true`: the expression is rooted at a prop or local (`props.useCaseFieldPrefix.dataName`,
|
||
`tabPrefix.x(idx).gausVal`), so the leaf and its owner are exact but the prefix of the path is
|
||
unknown (only the tail is given). A carrier prop's own name is not part of the path.
|
||
|
||
Cross-file targets are placeholders `<Dto>.<field>` (`binding=true`, `owner`, `field`,
|
||
`targetModule`) resolved by the finalize step `resolve-binding-placeholder` **exactly** — module →
|
||
`DATA_STRUCTURE` → `FIELD` — and dropped afterwards; a leaf the interface does not declare stays a
|
||
placeholder (listed by `search/identifier` with empty `sourceFile`). The item-74 stale-edge sweep
|
||
covers binding targets on a deep re-ingest.
|
||
|
||
**Endpoints.** `GET /projects/{p}/bindings?dto=&field=&mode=reads|writes&module=&partial=&countOnly=&limit=&offset=`
|
||
(CLI `ac bindings [--dto] [--field] [--mode] [--module] [--partial] [--count-only]`) → one row per
|
||
binding site: `{mode, dto, field, path, rootDto, kind, partial, component, attribute, function,
|
||
functionType, functionKind, module, sourceFile, lineNo, counterpartProject, counterpartModule,
|
||
counterpartField}`. The `counterpart*` columns are the field's `COUNTERPART_OF` twin (item 193), so
|
||
**"which page edits Java `Broker.ebene`"** is `ac bindings --dto Broker --field ebene --mode writes
|
||
-p purfe` and needs no query on the backend project. `dto` is the *declaring* interface (`Broker`),
|
||
not the root (`AgstammUseCase`) — filter on `rootDto` client-side when you need the latter. Paged like
|
||
`counterparts`; unknown `mode` → `400 MODE_UNSUPPORTED`. `GET /data-structures/{dto}/fields` now
|
||
returns TypeScript interface fields (`type=FIELD`) and carries `boundReads`/`boundWrites` per field
|
||
(0 for Natural/Java). The generic `variables/<Interface>.<field>/reads|writes` sees binding edges
|
||
too. The leaf's owner is the interface that *declares* the member — `datStart` bound through
|
||
`GeneralAgreementDO` lands on `AbstractHistorizedDO.datStart`.
|
||
|
||
**Limits.** String-form `field="…"` bindings and the lodash-path bindings of `pur-r-vbuch` (excluded
|
||
workspace) are not modelled; a `partial` binding cannot say which list element or tab; no
|
||
`USES_TYPE` from the component to the root DTO (the `rootDto` edge property answers that). A
|
||
binding placeholder and a store placeholder share the `FIELD` type, so a DTO named exactly like a
|
||
store key would merge their placeholders (`Dto.field` vs `key.field`) — not the case on `pur`.
|
||
|
||
**Verified on `purfe` (2026-09-22, server 322, recreated + `refresh --deep`, 285 files, 0
|
||
failures):** 257 binding sites (196 reads, 61 writes; 226 scalar leaves, 31 prefixes; 84 partial)
|
||
over 23 declaring DTOs, every one linked to its `pur` counterpart field; `SmartInput` 122,
|
||
`SmartOutput` 64, tables 32, column definitions 36. `counterparts?kind=field`: 744 fields, exactly
|
||
one twin each (six-fold fan-out before the qualified names), 112 unmatched.
|
||
|
||
## Styling: theme tokens and the style inventory (item 196)
|
||
|
||
**The theme.** The file with `createTheme({...})` (`pur-ui-common/src/theme.ts`) gets a
|
||
`DATA_STRUCTURE theme` (`kind=theme`) with one `FIELD` per token: `theme.<path>` for every leaf of the
|
||
literal (`palette.primary.dark`, `typography.h1.fontWeight`, `shape.borderRadius`, `sizes.*`; props
|
||
`token`, `tokenKind=path`, `value` folded through constants — `#0054A2` — and `constant` when the
|
||
leaf names one, `PRIMARY_DARK`) and `theme.<NAME>` for each exported string/number constant of that
|
||
file (`tokenKind=constant`, `PRIMARY`). MUI `components.styleOverrides` are recorded as tokens too, not
|
||
interpreted.
|
||
|
||
**Style blocks.** Every `sx={…}`, `style={…}` and `styled(X)(…)` block is a `STYLE` node under its
|
||
component `FUNCTION` (a `styled` block under the `kind=styled` function), named
|
||
`<function>.<sx|style|styled>@<line>:<col>`, with `styleKind`, `element` (the JSX tag or styled base:
|
||
`Box`, `'div'`), `properties` (the CSS keys, nested selectors flattened: `&:hover.color`,
|
||
`& .MuiPaper-root.background`), `literals` (hard-coded colours/lengths: `#005CA9`, `17px`, `-2%`,
|
||
`calc(100% - 16px)` — `mt: 2` is theme-relative and no literal), `dynamic` (a value the sidecar could
|
||
not classify, or a whole `sx={props.sx}`), `spread`. Plain `.css` files get one `STYLE` per rule from
|
||
the Tier-1 scanner (`body@7`, `@font-face@2`; `styleKind=css`, `selector`, `properties`, `literals`).
|
||
|
||
**Token reads.** Inside a style block every chain on a theme value is a `REFERENCES` edge `STYLE →
|
||
theme.<token>` with `property` (the CSS key it feeds) and `via=theme`; a theme value is anything typed
|
||
`Theme` (`useTheme()`, a `({ theme }) =>` styled parameter, the theme object imported under any name)
|
||
or named `theme`. `theme.spacing(2)` ends at `spacing`, `theme.palette.grey['200']` is
|
||
`palette.grey.200`; a theme constant (`color: PRIMARY`) references `theme.PRIMARY`. A token read outside
|
||
a style block (`borderColor={theme.palette.grey['200']}`, code) is the same edge from the enclosing
|
||
`FUNCTION` with `context` = the JSX attribute or `code`. Cross-file targets are placeholders
|
||
`theme.<token>` (`theme=true`) resolved by exact name at finalize when exactly one theme declares the
|
||
token; **a token no theme declares keeps its placeholder on purpose** — `GET /theme` lists it with
|
||
`declared=false` (MUI defaults such as `palette.grey.200`, `palette.common.white`, the `spacing`
|
||
function, or a typo).
|
||
|
||
**Endpoints.** `GET /projects/{p}/theme?unused=` (CLI `ac theme [--unused]`) → `[{token, kind, value,
|
||
constant, declared, module, lineNo, uses}]`, declared tokens first. `uses` counts **project**
|
||
references (style blocks and code); MUI's own consumption of a token is invisible, so `unused=true`
|
||
means "no project code references it", never "safe to delete" (`palette.primary.main` colours every
|
||
Button whether or not a component names it). `GET /projects/{p}/theme/{token}/usages` (`ac
|
||
theme-usages palette.primary.dark`, token = dotted path or constant name) → `[{function,
|
||
functionKind, module, sourceFile, lineNo, styleKind, element, property, context}]`; unknown token →
|
||
`404 TOKEN_NOT_FOUND`.
|
||
`GET /projects/{p}/styles?module=&kind=sx|style|styled|css&withLiterals=&countOnly=&limit=&offset=`
|
||
(`ac styles [--module] [--kind] [--with-literals]`) → `[{name, styleKind, element, selector,
|
||
function, module, sourceFile, lineNo, properties, literals, dynamic, tokens}]`, paged like
|
||
`counterparts`; **`withLiterals=true` is the review question**: which blocks hard-code colours and
|
||
lengths instead of using the theme. Unknown `kind` → `400 KIND_UNSUPPORTED`.
|
||
|
||
**Several themes (item 199).** Every `createTheme({...})` in a file is read; a token two themes of
|
||
one file declare is one row with the first theme's value and `variants: 2`. A token declared by two
|
||
theme *files* (light/dark) is one row per file, and a read of it resolves onto both, so each row
|
||
counts the use and `?unused=true` stays honest; `theme/{token}/usages` and `styles[].tokens`
|
||
report such a read once. Two reads of one token on one line of a style block (`color: PRIMARY,
|
||
borderColor: PRIMARY`) are one usage whose `property` is the comma list of the keys they feed.
|
||
CSS rules: a block-less `@import`/`@charset` line is not part of the next selector, braces inside
|
||
string values do not open blocks, a nested rule head inside an at-rule body is not a declaration,
|
||
and a selector repeated on one line (minified CSS) gets a `:col` suffix in its name.
|
||
|
||
**Limits.** `className` strings are not matched to CSS rules; Emotion `css` templates and MUI
|
||
`styleOverrides` are not modelled; a token read from a component-level function (not a style
|
||
block) that survives a re-parse keeps its resolved edge until the file's nodes are re-created
|
||
(same class as item 198).
|
||
|
||
**Verified on `purfe` (2026-09-22, server 326, recreated + `refresh --deep`, 285 files, 0
|
||
failures):** 129 declared tokens (111 paths, 18 constants), 93 of them with no project reference;
|
||
12 undeclared tokens the code reads (`palette.common.white` 4, `palette.grey.200` 4, `spacing` 3,
|
||
`palette.divider`, `palette.text.secondary`, `applyStyles`, `transitions.create`, …);
|
||
`palette.primary.dark` is the most-used token (26 reads: 17 `style`, 3 `sx`, 1 `styled`, 5 as a
|
||
plain prop such as `confirmColor`). Style inventory: 304 blocks (196 `sx`, 78 `style`, 27 `styled`,
|
||
3 CSS rules), 109 with hard-coded literals (`100%` 29, `1px` 21, `12px` 14, `17px` 14, …), 75
|
||
reading theme tokens, 41 dynamic. The only placeholders left in the project are the 12 undeclared
|
||
theme tokens — by design.
|
||
|
||
## Missing capability?
|
||
|
||
If the API/CLI genuinely cannot answer a question (not just
|
||
unreachable — the capability doesn't exist), finish the task via
|
||
grep/Explore as a fallback, then use `AskUserQuestion` to flag the gap and
|
||
ask whether it should become a roadmap item in `x-docs/roadmap.md`. Don't
|
||
silently fall back and move on.
|