# Using AgenticCode on This Repo (Dogfooding) This repo is ingested as project **`ac`** at `http://localhost:8787`. Per CLAUDE.md's dogfooding rule, prefer the AgenticCode API over grep/Explore for call graphs, callers/callees, DB access, dataflow, and module overviews when working *on this repo* — it's exactly the tool this project builds. **For full API semantics (params, response shapes, error codes, language applicability, curl examples)**, see `x-docs/agent-api-system-prompt.md` — that file is the canonical reference and is not duplicated here. This file only covers what's specific to using the API *as Claude Code, on this checkout*. ## Tool priority REST first (`GET /api/projects/ac/...`) → `ac` CLI (`ac callers`, `ac callees`, `ac call-tree`, `ac context`, `ac db-accesses`, ...) → grep/Explore. Only fall back past REST/CLI when the question genuinely isn't answerable by this API at all (see "Missing capability" below). If the server is unreachable, try `./manage-ac.sh deploy` before falling back further. ## Re-ingest before trusting results Query results reflect the last ingest, not the current working tree. **Refresh after code changes** before trusting query results: `ac refresh` or `POST /api/projects/ac/refresh` (add `--deep` / `?deep=true` for a full field-level pass). `refresh` is the single (re-)ingest surface (item 42) — the eager `ingest-all`/`ingest-module`/`ingest-call-graph` endpoints were removed. `ac refresh ` deep-ingests one module + its callees/data areas; add `--neighborhood` (`POST /refresh/{name}?scope=neighborhood`) to also pull in the module's transitive **callers** (whole call-graph neighbourhood). **Reconciliation on re-ingest (item 58).** A `refresh` now **purges stale nodes**: for every re-parsed file it deletes the nodes the fresh parse no longer produces (renamed/removed fields, moved statements) rather than leaving them to shadow the new ones — so identifier counts and `/search/identifier` results stay clean after a parser change or an edited source file. Applies to every full-parse path (whole-root `refresh` with or without `--deep`, and per-module `refresh/{name}`); the coarse Tier-1 scan run at project creation does not reconcile (it only ever runs on an empty graph). **Auto-invalidation (item 43).** The graph self-heals when files change or are deleted on disk — you rarely need a manual `refresh` for staleness: - **Changed file:** a field-level query on a module whose source changed re-ingests it automatically (its stored `sourceHash` no longer matches), so results reflect the current file. Granularity is per-module: a changed *dependency* is picked up when that dependency is itself queried by name. - **Deleted file:** a **whole-project** `refresh` now also **removes nodes for files deleted from disk** (reconciled against the filesystem), closing the earlier gap — no project re-create needed. (A per-module `refresh/{name}` does not sweep; it only touches its own tree.) - Controlled by `agenticcode.auto-invalidate.enabled` (default `true`). ## Reading source in this repo You have direct filesystem access to this checkout — **never** call `nodes/{id}/source` or `modules/{name}/source`. Use `sourceFile`/`startLine`/ `endLine` from a graph response (context, digest, search/identifier, nodes/{id}, ...) and read the file directly. This is always cheaper and gives full surrounding context; the `/source` endpoints exist only for API-only agents with no filesystem access. **Stale-source check (item 41).** The `/source` endpoints compare the file on disk against the content hash (`sourceHash`) stored at ingest. If the file changed since the last ingest they return `409 STALE_SOURCE` instead of slicing current text against old line numbers — re-ingest (refresh) the project to update the graph. Line ranges you read directly off disk are of course always current; this only guards the API's own slicing. Copycode/INCLUDE slices are raw pre-expansion file text. ## Comments: reachable, but never by default (item 141) **Every default query sees code only.** `search/identifier`, `search/annotation`, `search/references` and a plain `search/value` never match comment text — a Natural `* ...` banner, a trailing `/* ...`, a Java `//` line or a Javadoc block. That silence used to be a **wrong answer, not a missing one**, wherever a convention records something in a comment. The UPMS→PUR case: a reengineered service carries its Natural origin in a Javadoc block (`ServiceEndpoint:` / `UPMSFunction:` / `UpmsObject:`), and the documented Natural→Java lookup searches the program name in `pur`, where an empty result is read as *"not yet reengineered"* — which it answered for services reengineered months earlier. Three routes now reach comment text. Pick one before concluding "not present": | Question | Call | |---------------------------------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | "What does this module's header/change log say?" | `GET /modules/{name}/comments` (`ac comments `) — blocks with the declaration each documents | | "Where is this query / SQL / JSON text defined?" (Java) | `GET /search/value?value=…&contains=true` — since item 139 a `static final String` built from text blocks, literals, same-class constants and `+` carries its full text as `value`; a `"…".formatted(...)` constant carries its template (with `%s`) and `valueKind: template`. Anything built from a method call or another class's constant stays unresolved (no value) | | "Does this string appear anywhere, code **or** comment?" | `GET /search/value?value=…&includeComments=true` (`ac search-value --include-comments`) — comment hits carry `kind: "COMMENT"` | | "…and in text the parsers do not model at all, or in a module that is not deeply ingested?" | `GET /search/source?regex=…` — raw grep over the files on disk | ``` GET /pur/search/value?value=WPARTX0S&contains=true → [] (code only) GET /pur/search/value?value=WPARTX0S&contains=true&includeComments=true → the Javadoc origin block GET /upms/modules/WAGNTX0S/comments → `* #01 … Bug 266`, … ``` Comments stay **opt-in** deliberately: a comment hit is not the same evidence as a literal in code, and folding them into the default result set would move every existing completeness count (item 131's lesson). The flip side is the rule to remember — **an empty default search says nothing about comments.** ## Diagnosing a slow refresh (items 153/155) Two instrumentation layers, both aimed at the same question — *where does the time go?* * **Always on:** every persist batch logs one line with its statement breakdown (`Persist deep 200 files: total 8740 ms = prepare 108, merge-nodes 1630, ..., commit 20`), and every enrichment step logs duration plus created/deleted rows. The residual `commit` is deliberate: total minus the labels is the transaction commit, so nothing hides in an unnamed remainder. * **Opt-in per run:** `POST /refresh?deep=true&profile=true` (`ac refresh --deep --profile`) runs the enrichment steps under Cypher `PROFILE` and logs the five heaviest operators of every step slower than 5 s. Diagnostic only — it answers "is this step matching or writing?", which the step timing alone cannot. Measured overhead on `upms`: none worth reporting (908 s vs 905 s). * **Build-time diagnostic:** `agenticcode.ingest.comments.enabled=false` drops `COMMENT` nodes and their `DOCUMENTS` edges just before persist. It exists to measure what comments cost (item 179: ~17 % of a deep refresh) and is **not** a supported operating mode — with it off, `/comments` answers empty. It is a property, not a request parameter, so it needs a rebuild and cannot be set per run. What that measured, so nobody re-derives it. The figures below are **current** (2026-09-06, two clean runs, roadmap item 175); the campaign of items 153-175 took an `upms` deep refresh from **1 225 s to 301 s**, so any older number quoted elsewhere is stale by a factor of four. | | now | |---------------------------------|-------------------------------------------------------------------------------------------------------------------| | deep refresh `upms`, end to end | **301 s** (run-to-run spread 3.7 %) | | persist | ~149 s — `merge-edges` 47.6 s, `commit` 34.1 s, `merge-nodes` 24.0 s | | finalize | ~137 s — `resolve-field-placeholder` W/R 43.4 s, `link-args-to-params` 23.7 s, `resolve-bare-included` W/R 37.9 s | | parsing | ~9 s (interleaved with persist) | The original diagnosis still holds and is why those steps shrank: the five field-resolution steps were dominated by *matching*, not writing — `resolve-bare-included` once spent 273 M database hits to produce 43 753 rows. See roadmap items 156 and 175, and item 179 for what comment nodes cost. ## `limit`/`offset` work on some endpoints and are silently ignored on others **Every endpoint answers completely.** The difference is whether it lets you ask for less. 19 list endpoints declare no `limit`/`offset` at all, and JAX-RS drops an undeclared query parameter without a word — so `?limit=2` there is not an error, it simply has no effect: ``` GET /upms/modules?limit=2 -> 3 587 rows GET /upms/modules/WAGNTX0S/functions?limit=2 -> 22 rows GET /upms/modules/WAGNTX0S/data-structures?limit=1 -> 16 rows ``` **Ignore `limit`/`offset` (always the full list):** `/modules` · `/modules/{name}/functions` · `/modules/{name}/functions/overrides` · `/modules/{name}/functions/{function}/overrides` · `/modules/{name}/data-structures` · `/modules/{name}/columns` · `/modules/{name}/payload` · `/modules/{name}/dispatch-table` · `/modules/{name}/sql-statements` · `/data-structures/{name}/fields` · `/db-tables/{name}/columns` · `/variables/{name}/reads` · `/variables/{name}/writes` · `/variables/{name}/flow-forward` · `/variables/{name}/flow-backward` · `/variables/{name}/field-flow` · `/duplicates` · `/dynamic-calls/unresolved` · `/dynamic-calls/overrides` **Honour them** (14, verified against the method signatures): the search endpoints (`search/identifier`, `search/value`, `search/annotation`, `search/references`, `search/source`), `rest-endpoints`, and per module `callers`, `callees`, `context`, `db-accesses`, `graph`, `reaches`, `workfile-accesses`, `comments`. `call-tree` is bounded differently again — by `depth`, not by row count — which is the right shape for a tree but means `limit` does nothing there either. Which direction the mistake runs matters, so be precise about it: a dropped `limit` means you get **more** than you asked for, never less. It costs tokens, never correctness — the opposite of item 131's failure, where a silent 50-row cap was read as the complete set. Nothing here can under-report. The practical consequence is budget, not trust: `/modules` on `upms` is 3 587 rows in one response. Narrow with the filters those endpoints *do* have (`?kind=`, `?sourceFile=`, `?module=`, `?extendsName=`) rather than with a `limit` that will be ignored, and prefer `/modules/{name}/digest` or `/context` when you want an overview rather than an enumeration. *(Implementing `limit` on those 19 was considered and deliberately not done: the endpoints are honest as they stand, and a `limit` that ever acquired a default would reintroduce exactly the silent truncation item 131 removed.)* ## `callers` / `callees` say when they are cut (item 181) `callers`, `callees` and `functions/{fn}/callers` now send `X-AC-Total-Count` and `X-AC-Truncated` like the search endpoints, and their body carries `total` and `truncated` next to `sourceFiles` / `items` (also for `fields=name`; `ac callers` / `ac callees` print the usual "truncated" warning). The default page is still 50: `upms/modules/DPARTFN0/callees` answers 50 of 58 with `X-AC-Truncated: true` — before, the 8 missing callees (among them the module that writes the partner) looked like an inconsistency with `digest`. **Read `truncated` before concluding "X does not call Y"**; ask with `limit=1000` or narrow with `scope=external`. ## Truncation is now visible on the search endpoints (item 131) `search/identifier`, `search/value`, `search/annotation`, `search/references` and `rest-endpoints` send two headers with every answer: | Header | Meaning | |--------------------|---------------------------------------------------------| | `X-AC-Total-Count` | how many rows match in total, ignoring `limit`/`offset` | | `X-AC-Truncated` | `true` when this page leaves some out | and all five accept **`?countOnly=true`** (CLI `--count-only`), returning `{"count": n}` instead of rows — a completeness question is a counting question, and `@Column` on `pur` is 3 630 rows ≈ 250 k tokens if you ask for them. This closes item 103's own follow-up. The bodies stay bare arrays (no contract change), for the same reason as item 130's scope headers. Why it matters: `search/annotation?name=Immutable` returned 50 of 95 rows with no total, no flag and no `Link`/`X-Total-Count` — and a real UPMS→PUR audit read that page as the whole set, recording that 17 entities had lost `@Immutable` when **zero** had. Re-checked against all 114 rows of `Tables_meta.csv`: 94 non-writable carry it, 20 writable do not, no deviations. The finding cost a day and was pure artefact of the cut. **`search/references` and `rest-endpoints` joined this late (item 135, 2026-08-20).** Item 131 was written about the annotation search and both were overlooked. `search/references` was the damaging one: it capped at the default 50 and said nothing at all, so a rename scoped from that page missed every site past the fiftieth and looked complete doing it. `rest-endpoints` defaults to an uncapped limit and so never lost rows, but it was equally silent about how many there are. Both were found by `x-scripts/verify-api.sh` on its first run, not by a test. Two mechanics worth knowing: the total costs a **second query only when the page comes back full** (a short page is provably the end, so the total is arithmetic), and a total that divides evenly by `limit` makes the last full page report `truncated` with the next page empty — one wasted call, never a wrong answer. `ac` prints a note to **stderr** when a response is flagged truncated, so piping the body into `jq` stays clean. ## REST surface and scope headers (item 130) `GET /api/projects/{p}/rest-endpoints?module=&countOnly=&limit=&offset=` (CLI `ac rest-endpoints`) lists `{httpMethod, path, module, moduleSimpleName, handler, sourceFile, startLine}` — the **composed** path (class-level `@Path` + method-level `@Path`), so "which code runs for `POST /partners`" is one call. Previously the two halves had to be joined by hand from two `/search/annotation` calls, because annotations are stored by name without their arguments; the parser now persists `restPath` and `httpMethod`, including a `@Path` written as a constant reference. A method with no HTTP-verb annotation is not an endpoint and is excluded. A class with no `@Path` of its own inherits the nearest one from its `extends`/`implements` ancestry, as JAX-RS does. Rows carry `outbound: true` when the declaring type is a `@RegisterRestClient` interface — a call the application *makes*, not one it serves; its path is usually empty because the base URI comes from configuration. **Scope and freshness now ride on every project-scoped response as headers:** | Header | Meaning | |--------------------------|-----------------------------------------------------------------------------------------------------------------------------| | `X-AC-Exclude-Dirs` | directories the ingest skipped, or `(none)` | | `X-AC-Ingested-At` | when the graph was last walked (item 126) | | `X-AC-Ingest-Incomplete` | `true` while a whole-root pass runs or after one that never finished (item 129); `unknown` when no ingest was ever recorded | Read `X-AC-Exclude-Dirs` before trusting an **empty** answer: "no callers" means "none outside tests" in a project excluding `test` (`app`) and "none at all" in one that does not (`pur`, `ac`) — the bodies are identical. They are **headers, not body fields**, because most endpoints answer with a bare JSON array (`db-accesses`, `functions`, `search/identifier`, …); adding a field there would mean restructuring array → object and breaking the web UI's generated client, the CLI printers and any agent that indexes `[0]`. The trade-off is that an agent reading only the JSON body will not see them — so if you consume this API programmatically, read the headers too. The project shell is cached for ~10 s to keep this off the request's critical path, and the ingest path invalidates that cache explicitly, so `X-AC-Ingest-Incomplete` flips as soon as a refresh starts rather than up to 10 s later. ## Every reference site of a name (item 128) `GET /api/projects/{p}/search/references?name=&kind=&countOnly=&limit=&offset=` (CLI `ac references `) returns `{sourceFile, lineNo, kind, inModule, target}` per **mention** of a type — not just per call: | `kind` | Where it comes from | |-------------------------|--------------------------------------------| | `CALL` | a call site (`CALLS`) | | `IMPORT` | an `import` of the type | | `TYPE` | a declared field / parameter / return type | | `ANNOTATION` | the type used as an annotation | | `EXTENDS`, `IMPLEMENTS` | inheritance | | `INJECTS` | CDI wiring | | `CLASS_LITERAL` | `X.class` in argument position | | `INCLUDE` | Natural copycode inclusion | Use it to **scope a rename**. `callers` sees calls alone, so a file that only imports the class, declares a field of it, or names it in an annotation was invisible — and the rename that missed it looked complete. The `name` may be the identity (FQN) or the short form; `target` echoes what it resolved to. An unknown `kind` is `400 INVALID_KIND`, never an empty list. **Known limits, by design:** * **Local-variable types and generic type arguments are not indexed** — `List x` records `List`, not `Target`. They multiply edge volume for much less value than the positions above. * **Same-package references have no import**, so within one package the index rests on declared-type positions alone. * **Imports are only indexed when they look project-internal** (they share the first two package segments with the importing file). Otherwise every `java.util`/framework import would mint a placeholder node on every ingest, just for the finalize sweep to delete it again. * **Natural has no import or type-position concept.** It contributes `CALL`, `INCLUDE` and inheritance kinds only; this is not parity with Java and should not be read as such. * Reference edges are written **at parse time**, so they only exist for files re-parsed since this landed — a project needs a `refresh` before the index is complete. * Mentions use their own `MENTIONS` edge type, kept out of `CALLS`/`REFERENCES` deliberately: the call-graph traversals (`callers`, `callees`, `call-tree`, `ego-graph`) follow `REFERENCES` as wiring, so folding imports into it made an `import` surface as a **caller**. `/search/references` is the only endpoint that reads `MENTIONS`; the call graph is unchanged. ## Refreshing only what changed (item 129) `POST /api/projects/{p}/refresh` (CLI `ac refresh`) has two ways to avoid re-walking a whole root: | Form | What it does | |---------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | `?paths=a/B.java,c/D.java` (`ac refresh --paths a/B.java,c/D.java`) | Re-ingests exactly those relative paths, deep, plus their dependencies. Paths that match no file — **or that are not ingestible source files at all**, like `pom.xml` — come back in `unresolved`; a typo'd path is never silently dropped. Does **not** run the deleted-file sweep and does **not** move `ingestedAt`: both need a whole-root walk. | | `?changedOnly=true` (`ac refresh --changed-only`) | Whole-root walk, but re-parses only files whose content hash differs from the graph's (files with no stored hash count as changed). | **`changedOnly` is opt-in on purpose.** Three whole-walk behaviours are reduced, and one of them would be outright corruption if it were hidden: * **A changed Natural copycode disables skipping for that entire run** (logged). Natural projects only — a `.cpy` sitting in a Java project (as a test fixture, say) is not inlined by anything and no longer stands the optimisation down. Copycode text is inlined into the including module *at parse time*, so a module whose `.cpy` changed parses differently while its own hash is unchanged — skipping it would leave a stale expansion behind with nothing to indicate it. * **Duplicate-identity detection** only sees the changed files, so it can confirm duplicates among them but not discover new ones elsewhere. Existing markers are never cleared. * **User-exit LoC annotation** (item 47) is re-stamped only on re-parsed files. Enrichment is project-wide and still runs in full, so this cuts **parse+persist** time only — not the finalize pass. On small projects the whole deep refresh is already ~30 s, so measure before assuming a win. **`ingest.incomplete`** (on `/projects` and `/projects/{p}`) is `true` while a whole-root pass runs and **stays true if one never finished** — a crash, a container stop, an aborted deep refresh. Before this, an interrupted deep refresh was indistinguishable from a clean graph: the enrichment steps that already ran are committed, so queries keep answering, just from a half-updated graph. It cannot self-heal (a killed process clears nothing) and does not distinguish "running right now" from "died an hour ago" — both mean the same thing to a caller. A completed refresh clears it. ## Renaming a project (item 202) `POST /api/projects/{p}/rename` with `{"newName": "…"}` (CLI `ac project rename `) moves the project key on every node and override and in other projects' `counterparts` lists, then the shell; the old name answers `404` afterwards and nothing needs re-ingesting. Refusals: `400 INVALID_REQUEST` (blank or unchanged), `404`, `409 PROJECT_EXISTS`. The node rewrite is batched; an interrupted rename is finished by running it again (97 s for the 939k-node `upms`). Use it to keep a reference graph next to a fresh ingest (`ac project rename upms upms_alt`, then create `upms` again and compare). A full `DELETE` of a project now also removes its manual overrides; `recreate` keeps them. ## Is this project's graph any good? (item 126) `GET /api/projects` and `GET /api/projects/{p}` (CLI `ac project list` / `ac project show

`) carry an `ingest` object describing the **last whole-root ingest**: ```json "ingest": { "ingestedAt": "2026-08-18T10:12:44Z", "mode": "full", "filesExamined": 2981, "filesPersisted": 2977, "filesFailed": 4, "failures": ["a/B.java", "..."], "failuresTruncated": false, "durationSeconds": 176, "serverVersion": "…" } ``` Use it before trusting a **negative** answer: without it, "no such module" and "that part of the project was never ingested" are the same empty response. Three rules the field obeys: * **`ingest: null` means never recorded**, not "ingested nothing" — a project last walked before this existed reads as null rather than as a fabricated zero. * **Only whole-root passes write it** — the create-time Tier-1 scan, `refresh`, `refresh?deep=true`. A by-name `refresh/{name}`, a deep ingest or a fan-out warm ingests real files but sees a fraction of the tree, so it deliberately leaves `ingestedAt` alone; otherwise deepening one module would advertise the whole project as freshly walked. * **`ingestedAt` is not a freshness guarantee.** It says when the walk ran, not that the graph still matches disk — a file edited a minute later is stale while the timestamp still looks recent. For the real check, read a file through `GET /{p}/source?file=…`, which answers `409 STALE_SOURCE` when the content no longer matches the ingested hash. (Targeted/incremental refresh is item 129.) `failures` is capped at 200 paths while `filesFailed` stays exact; `failuresTruncated` says whether the list was cut, so a short list is never mistaken for the whole story. ## Tier-1 coarse scan on project create (item 36) Creating a project (`POST /api/projects/{p}`) now runs a **Tier-1 coarse reference scan** of the root before returning, so the project is immediately queryable — no separate ingest call. The scan is lexer-level (Natural) / a coarse projection of the full parse (Java): it records module/function /data-structure **shells**, the **identifier index** (declared fields, class members), and coarse `CALLS`/`READS`/`WRITES`/`INCLUDES` references (Natural `PERFORM`, `CALLNAT '...'`, dynamic `CALLNAT PGM-VAR`, `PARAMETER`/`LOCAL USING` copybooks; Java resolved calls/type refs) — but **no** deep bodies (control flow, statement-level dataflow, arg→param). Scanned modules land `CALL_GRAPH`/`NOT_INGESTED`; field-level detail is filled in by the on-demand deep ingest below. Each module shell carries a `sourceHash`. Disable with `agenticcode.tier1.scan-on-create=false` (creates an empty project shell to scan later). A missing/unscannable root is non-fatal — the project is created empty and can be re-scanned. ## Unresolved references (item 40) A reference whose target isn't (yet) ingested — a `CALLNAT`/`PERFORM`/`USING` to a module/copybook absent from the project, or a dynamic `CALLNAT PGM-VAR` whose literal can't be recovered — is stored as a **deduped placeholder node** (blank `sourceFile`). Enrichment stamps each with an `unresolved` boolean: `true` while genuinely dangling, `false` once a real definition of that name is ingested. `/search/identifier` returns it as `unresolved` on each `IdentifierMatch`, and `GET /nodes/{id}` carries it in the node's properties — so an agent can tell a dangling/dynamic reference apart from a resolved one. In `callers`/`callees` such targets already appear as entries with a blank `sourceFile`. ### Querying a module endpoint: four answers, not one (items 107, 115, 114) Every `GET /api/projects/{p}/modules/{name}/…` endpoint used to answer `200` with an all-zeros shell for a name that exists nowhere in the graph, byte-identical to a real-but-empty module's answer. That is not cosmetic: `call-tree?depth=4` returning `200` with 0 modules reads as **"analysed, nothing found"** when the truth is **"not analysable"** — the module's source was never in the checkout. The graph knows four states and each now gets its own status: | state | answer | what it means | |-------------------------------------------------------------------------------|---------------------------------------------------------------|------------------------------------------------------------------------| | no `MODULE` node for the name | `404 MODULE_NOT_FOUND` | unknown name (typo, or not in this project) | | **placeholder** — node exists because something calls it, source never parsed | `409` + `{status:"NOT_INGESTED", module, detail, nextAction}` | knowable in principle, not analysed yet | | **ambiguous** — several real modules share the name (Java) | `409` + `{code:"AMBIGUOUS_NAME", details:{candidates:[...]}}` | the name does not identify a module; pick one | | **duplicate identity** — the same name in several files, skipped at ingest | `409` + `{code:"DUPLICATE_IDENTITY", details:{paths:[...]}}` | exists twice, deliberately not ingested (item 114); see `/duplicates` | | real, ingested module | `200` | the body is the answer — an empty body genuinely means "nothing found" | **Ambiguous names (item 115).** A Java simple name is not unique: nested `@Nested` test classes, `Builder`, `Config`, `WorkingStorage`. Measured on `pur`, **163 names covering 385 modules (~8%)** were addressable only ambiguously, and the endpoints used to answer with the *union* across unrelated classes — `/modules/BrokerHistoryTests/functions` returned 627 functions for a class that has 107. They now refuse and list the candidates. Repeat the request with `?sourceFile=`: ``` GET /modules/Shared/functions → 409 AMBIGUOUS_NAME, candidates ["a/Shared.java","b/Shared.java"] GET /modules/Shared/functions?sourceFile=a/Shared.java → 200 ``` **A Java module's name IS its fully-qualified name** (item 117) — `com.example.OrderService`, and `com.example.Outer.Inner` for a nested class. That is what makes same-simple-name classes distinguishable at all. Every module endpoint accepts **either** form: ``` GET /modules/com.example.OrderService/digest → 200, always exact GET /modules/OrderService/digest → 200 when unique, else 409 AMBIGUOUS_NAME ``` Responses carry `simpleName` alongside `name` for display. The `?module=` and `?extends=` filters and the `ac` CLI take either form too; `--source-file` remains available on every module command. **Which Java types are modules** (item 119). Classes, interfaces, **enums, records and annotation types** — `moduleKind` is one of `CLASS | INTERFACE | ENUM | RECORD | ANNOTATION` (Natural adds `PROGRAM | SUBPROGRAM | …`), and `?moduleKind=` filters on it. Their content is modelled the way each kind carries it: a record's components and an annotation type's members are `FIELD`s (the latter with `defaultValue` where declared), an enum's constants are `CONSTANT`s, and an enum's or record's `implements` is a real edge — so a call against an interface fans out to an enum implementing it. Before 119 these three kinds were not parsed at all: `GET /modules/SomeEnum/digest` answered `404`, and a record referenced from elsewhere stayed an unresolved placeholder (`409 NOT_INGESTED`). One gap remains by design — a record's *compact* canonical constructor is not a function node, so calls made in its body are invisible. Natural is unaffected throughout: its module names are file stems, and colliding identities are skipped at ingest, so they are unique by construction (`upms` has zero ambiguous names, `pur` 163). One limit worth knowing: only the request's **root** module is disambiguated. Inside a traversal this no longer merges anything, because module names are unique per project after item 117 (`pur` and `upms` have zero duplicate names) — a reference the parser could not qualify is returned flagged `unresolved: true` rather than attached to an arbitrary candidate. A name skipped at ingest because it exists in **more than one file** answers `409 DUPLICATE_IDENTITY` with the conflicting `details.paths` (item 114); `GET /duplicates` lists them all. Note what that does *not* fix: the skipped file's own calls were never parsed, so caller lists elsewhere can still be short — they just no longer look complete. The `409` applies to everything derived from the module's **own** source: `digest`, `context`, `call-tree`, `callees`, `db-accesses`, `workfile-accesses`, `sql-statements`, `functions`, `functions/overrides`, `functions/{fn}/overrides`, `functions/{fn}/callers`, `data-structures`, `dispatch-table`, `payload`, `columns`. `callers` and `graph` stay `200` for a placeholder — their data comes from the **calling** modules' source and is genuine. When you get a `409`, **fall back to `/callers`**: it is the one honest answer available for a module whose own source is missing. (`digest` no longer surfaces those callers, since its other fields would all be structurally zero.) `/modules/{name}/source` is unchanged: it already answered `404 MODULE_NOT_FOUND` for both an absent module and a placeholder, since there is no source to serve either way. **Scope limit — this guards the *root* module of a request only.** A `call-tree` that traverses *into* placeholder targets still reports that subtree as empty without flagging it, so a dispatcher whose targets are all placeholders still returns `200` with a silently truncated tree. Cross-check the targets you care about individually (a `409` tells you it is unanalysed) — see item 103. **Data literals are not call targets (item 62).** A `CALLNAT ` whose target is really a data value — a browse key reaching the call site through a copycode/macro argument — used to leave a permanent unresolved placeholder callee. Enrichment now reaps such a placeholder when it has no Natural sigil (`#`/`&`/`+`), *is* a real `VARIABLE`/`CONSTANT` of the project, matches no real `MODULE`, and is only ever reached by inferred (`CALLNAT_DYNAMIC`/`INCLUDE_MACRO`) edges. So `callees`, `call-tree`, the ego graph and the unresolved list no longer list data names as modules. Genuine dynamic dispatch (`CALLNAT #PGM-VAR`, sigil'd) is still reported as an unresolved target, and a static `CALLNAT 'X'` is always trusted even when `X` collides with a field name. **The ingest summary agrees with the graph (item 64).** A `refresh`/`refresh/{name}` response's `unresolved` list is built during the file walk, independently of the graph — before item 64 it therefore reported data fields as missing modules (`MODULE CO-TABLA`, `MODULE NAME-DESC-SP`) that enrichment had already reaped, so the summary and the graph disagreed. It now applies the same item-62 gates as the graph reap, with the "is this a real field?" test asked project-wide (a data literal is often declared outside the ingested tree). Genuinely missing modules are still reported — including sigil'd dynamic dispatch (`MODULE #GETSHORT-MODUL`) and a missing module whose name collides with a field name but is called statically (the `RPC-CNTX` class). Treat `unresolved` as "dependencies that really are absent". **Constant-folded string-assembled targets (item 83).** A dispatcher often builds the `CALLNAT ` name from a base literal plus one or more `SUBSTR` overlays — e.g. `#GETSHORT-MODUL` in `YGEAGGNH`, assembled by `MOVE 'YGEAGKEY' TO #GETSHORT-MODUL` then `MOVE 'GN0' TO SUBSTR(#GETSHORT-MODUL,6,3)` → `YGEAGGN0`. The parser records each `SUBSTR` write as a `WRITES` carrying `substrPos`/`substrLen` (1-based), and the `resolve-dynamic-callnat-fold` enrichment step folds the last full-var literal written before the call site with the intervening overlays (`left`/`substring`) and MERGEs a resolved `CALLS` edge (`callKind=CALLNAT_DYNAMIC`, `folded=true`) to the assembled module when it is a real `MODULE`. It runs in every ingest mode (cheap, bounded by dynamic call sites) and over-approximates like the other dynamic resolvers. This auto-recovers the `Y…GNH → Y…GN0` family with no manual override, so folded sites drop out of `dynamic-calls/unresolved` and the assembled target appears in `callees`/`call-tree`/`graph`/`ego graph` tagged `CALLNAT_DYNAMIC`. **Pin what the resolvers can't: manual dynamic-`CALLNAT` overrides (item 82).** Some `CALLNAT ` targets are beyond the auto-resolvers — a name never present as a recoverable literal in the reachable code (read from a work-file record, supplied by an unlinked caller, …). Such a site stays an unresolved placeholder. A human or agent resolves it via `POST /api/projects/{p}/dynamic-calls/overrides` with the call site's `originFile` + `lineNo` (from `GET .../dynamic-calls/unresolved`) and the target module name(s) — multiple targets for a genuine branch. The override is stored as a `:DynamicCallOverride` node **outside** the `:AstNode` graph, so a refresh never deletes it and an enrichment step (`apply-manual-dynamic-callnat`, after the auto dynamic-CALLNAT resolvers, before the placeholder cleanup) **re-applies it automatically** — MERGEing a `CALLS` edge (`callKind=CALLNAT_DYNAMIC`, `resolvedBy='manual'`) to each target and flagging the placeholder `manualHidden` so `callees`/`digest`/`graph`/`call-tree` show the real target, not the `#var`. It only applies while the site is still unresolved: once an auto-resolver catches up, the override is skipped and listed `obsolete` — **except a constant-fold (item 83), which a manual override outranks**: the fold skips a site carrying a `:DynamicCallOverride`, and a stale `folded` edge there is dropped (`delete-folded-overridden-dynamic-callnat`) before `apply-manual-dynamic-callnat` runs, so the pinned target replaces it. `DELETE .../dynamic-calls/overrides?originFile=&lineNo=` resets one site (omit both = all), clearing the manual edges and un-hiding the placeholder inline (no refresh). A target that is not a real `MODULE` is rejected `400 UNKNOWN_TARGET`. **Bug B fix:** the `callees` items now carry `unresolved` (mirroring what `graph` already exposed), so an unresolved dynamic target is machine-distinguishable from a resolved one without inspecting `sourceFile`. Since item 200 the apply and the reset also rebuild the calling modules' derived `CALLS_MODULE` edges in the same transaction, so `reaches` and `field-flow` honour a pinned target without a refresh, like `callees` and `call-tree` already did. **Dispatch guards: read `guards` — it is the only complete condition (item 72).** A `dispatch-table` row's `guardField`/`guardValue`/`guardValues` describe the **innermost** `DECIDE` only. Natural nests value-`DECIDE`s inside each other (824 of upms's 7,729 blocks, 178 modules), and then the innermost guard is just one conjunct: in `VMULTMN4`, the row for `YTABLMA0.TX-TABLA` reports `#FIELD-NAME = 'TX-TABLA'`, but the assignment also requires `#SHORT-VIEW = 'TABL'`. `guards` is the full chain — `[{field, values}]`, **outermost first, joined by AND**, each link's `values` joined by OR. Reading only the legacy fields **over-generalises**: port that to Java and you get a branch firing where Natural never would. For an unnested `DECIDE` the chain has one link and says the same as the legacy fields. - **Still incomplete for `NONE`/`ANY` branches (item 73):** an assignment in a `NONE` branch is reported under its enclosing chain alone, but its real condition is "enclosing guard **AND NOT** any sibling `VALUE`" — a negation a chain of equalities cannot express. `guards` is strictly better than the legacy fields, not a total answer. **A dispatch row's `lineNo` belongs to `sourceFile`, not to the module (item 122).** `dispatch-table` rows now carry the same provenance quartet as `callees`/`db-accesses`/`workfile-accesses`/`functions`: `sourceFile`, `viaCopycode`, `includedAt`, `includePath`. Resolve `lineNo` **against `sourceFile`** — when `viaCopycode` is non-null the assignment is written in that copycode and `lineNo` is a line of the `.cpy`, while `includedAt` is the `INCLUDE` line in the module. Before item 122 the row carried only `lineNo`, so following it against the module file landed somewhere arbitrary: 26 of `VCOMIN50`'s 44 rows reported line 18 or 20, which in that module is a change-history comment; the real sites are `ISICINDE.cpy:18` and `ISICINDI.cpy:20`. Note the guard chain may **span** the include boundary — the outer `DECIDE` in the module, the inner one in the copycode — so a row can have a multi-link `guards` chain whose links live in different files. **Within one guard: prefer `guardValues` over `guardValue` (item 64).** `dispatch-table` rows carry both. `guardValue` is **lossy** and kept only for compatibility: it comma-joins the branch's `VALUE` literals, which drops blank alternatives entirely and, for a multi-alternative branch, yields a synthetic string the guarded field never equals (`"A1, A2"`). `guardValues` is the faithful list — every alternative in source order, blanks included — so `VALUE 'GENAGREE-WOUT-SP', ' '` reports `["GENAGREE-WOUT-SP", " "]`, recording that a **blank** guard field also routes into that branch. When reasoning about routing (or porting a `DECIDE` to Java), read `guardValues`; `guardValue` will silently under-report the branch's conditions. **Text that is not code never yields a call (items 61 & 63).** The `CALLNAT`/`CALLNAT_DYNAMIC` patterns are unanchored (a `CALLNAT` may legally appear mid-line), so both ingest tiers first neutralise non-code text: a full-line `*` comment is skipped, a trailing `/* …` is stripped (item 61), and a match whose **keyword** falls inside a quoted string literal is rejected (item 63). A real `CALLNAT 'MOD'` is unaffected — its keyword sits outside the quotes. This matters for trusting `callers`/`callees`/ `call-tree`: before item 63, prose such as `WRITE(#MSG) 'NACH CALLNAT ISINGEAG:'` or `#ERR-TYPE := 'Callnat USIA008N'` fabricated a `CALLNAT_DYNAMIC` edge to the **real** module of that name, so a mere log message appeared as a genuine call — and, because a real module existed, it was *not* flagged `unresolved` and could not be reaped by the item-62 cleanup. If you query a graph ingested before 2026-07-16, re-ingest (`ac refresh`) before trusting call-graph edges into modules that are also mentioned in log/error text. **Calls made through a copycode's arguments (items 120/121/123).** In Natural the target of a `CALLNAT` is often not written at the call site at all: a copycode receives the module name as a positional `INCLUDE` argument and issues `CALLNAT &2&`. Three defects in that argument path — arguments continued on the next line, the doubled-quote escape `'''X'''`, and the double-quote delimiter `'"X"'` — meant such a call produced **no edge and no unresolved-dynamic-call entry**, so `callees`, `callers`, `call-tree` and `reaches` agreed on an answer that was simply absent, with nothing saying "not analysed". This hit the browse/access layer hardest, because that is where the idiom lives: `YCARPBN1` and `YPOLIBN1` reported **0** callers each. Fixed 2026-08-07; re-ingest recovered 2672 copycode-derived call pairs (+49%) in `upms` with none lost. **A graph ingested before 2026-08-07 under-reports Natural callers/callees, and does so silently — re-ingest before concluding a Natural module is unused.** A copycode parameter that genuinely has no argument now surfaces in `/dynamic-calls/unresolved` as `&n&` rather than being dropped, so "not analysable" is visible. **A call edge no longer outlives the call it was parsed from (item 124).** Until 2026-08-09 a refresh only *added* the corrected call and left the old one in place, because an edge is reaped only when one of its endpoints is — and a parser fix changes neither (the calling subroutine is unchanged, the old target is a never-swept placeholder). So `callees`/`callers` could report a call that no source line makes, flagged `unresolved: true` and indistinguishable from a genuine unresolved dynamic call. A re-parsed Natural file's call edges are now reaped before the fresh ones are merged, and a call-target placeholder left with no callers is deleted. **Two consequences for a graph ingested before 2026-08-09:** an `unresolved: true` callee may be an artefact of an already-fixed parser bug rather than a real dynamic call, and `search/identifier` may list module names that exist nowhere in the source. Both clear on the next deep refresh. Note the reap deliberately spares `CALLNAT_DYNAMIC` edges onto *real* modules — those are the dynamic-call resolvers' output, not the parser's. ## LoC / SLoC metrics (item 46) Every file-level node (a `MODULE` program/class, or a `DATA_STRUCTURE` for a Natural `.lda`/`.pda` data area) is stamped at ingest with two deterministic line metrics: - **`loc`** — physical lines of the file (language-independent; a trailing newline adds no phantom line). - **`sloc`** — source lines of code: non-blank, non-comment lines, computed **per language**. Natural drops full-line `*`/`**`/`/*` and inline `/*` comments; Java drops `//` and `/* … */` blocks while keeping those tokens when they appear inside string literals. Both the cheap Tier-1 coarse scan and the deep Tier-2 parse use the *same* per-language counter, so a module's `loc`/`sloc` are identical at any ingest depth — you can sum them to get exact, reproducible project totals. Where to read them: - `GET /modules` — `loc`/`sloc` on each row. - `GET /modules/{name}/context` — `loc`/`sloc` on the module. - `GET /nodes/{id}` — `loc`/`sloc` in the node's raw properties. - `GET /loc` (`ac loc`) — the rollup: a per-language breakdown (`fileCount`, `loc`, `sloc`) plus a project-wide total, optionally narrowed by `?language=` / `?sourceFile=`. Each source file is counted once even when it yields several nodes (Java inner classes, Natural inline groups). `null` metrics mean the node predates item 46 — re-ingest (`refresh`) to backfill. ### Generated vs. user-exit split (item 47) A project can be created with a source **`language`** (required at creation; an attribute only — ingest still classifies files by extension) and a **`generatedDir`**/**`userExitDir`** pair (directory *names*, matched as path components like `excludeDirs`; both or neither). Generated modules already contain their hand-written user-exit twin inline, so at ingest a module under `generatedDir` whose name also occurs under `userExitDir` is annotated with that twin's LoC/SLoC (`userExitLoc`/`userExitSloc`). User-exit files are **not** ingested as standalone modules (they would collide by name) — the walk skips `userExitDir`. **Consequence for all non-LoC analysis:** the `generatedDir` copy is the *canonical, sole* module for every structural query (call graph, DB access, functions, data structures, identifiers, dataflow, dispatch table). `userExitDir` exists **only** to compute the generated-vs-manually-written LoC split below; it never contributes nodes/edges. So when verifying an API response against source for a Natural module, always read the `generatedDir` file (e.g. `generated_src/subprogram/WGEAGB0S.nat`), not the `user_exit` fragment. `GET /loc` (`ac loc`) then reports, per language row **and** in the project total: - **`loc`/`sloc`** — the **total** (generated, which already includes the user exits). - **`userExitLoc`/`userExitSloc`** — the sum of the annotated user-exit twins (the hand-written part). - **`generatedExclusiveLoc`/`generatedExclusiveSloc`** — total − user-exit, clamped ≥0 per file (the purely generated part). All three are `0` for projects without a generated/user-exit split. Create with `ac project create -l natural -g generated_src -u user_exit`, or add the split to an existing project via `ac project update -g generated_src -u user_exit`. ## Java `DB_ACCESS` in a project without JPA entities (item 140, 2026-08-27) A `DB_TABLE` node is only ever created from a JPA `@Entity` or a Panache active-record class. In a Java project that has none, **no** `DB_ACCESS` candidate can resolve — and the parser's candidate heuristic is a deliberate over-approximation: its read gate admits *any* static receiver whose method starts with `get`/`find`/`read`/`list`/… , so `UserContext.getCurrent()` and `TextUtils.getColumn(line, 0, 8)` become candidates. On `app` that produced **2219** `DB_ACCESS` nodes in a codebase with no database access whatsoever. `db-accesses` never showed them (it joins the table with a plain `MATCH`), but **`sql-statements` did**: it joins with `OPTIONAL MATCH`, so unresolved candidates came back as rows with `"table": null` — 76 of them on a single `app` module. Since 2026-08-27 the enrichment step `reap-java-db-access-without-tables` deletes every Java `DB_ACCESS` of a project that holds no `DB_TABLE`. For such a project `sql-statements` is now empty instead of noisy. Three things to know: - **The gate is project-level, not per node.** One entity anywhere in the project switches the reaper off, and the unresolved candidates stay. `pur` (2014 of 3951 unresolved) and `ac` (221 of 335) are unaffected, and still return `table: null` rows. Treat a `sql-statements` row whose `table` is `null` as unverified, in any project that has tables. - **Natural is untouched.** A Natural `DB_ACCESS` comes from a literal `READ`/`FIND`/`STORE` and is a real access whether or not its view resolved (`upms`: 14302 of 14303 resolve). - **Recovering from it needs a full refresh.** If such a project later gains its first entity, the reaped nodes only come back for files that are actually re-parsed — `changedOnly` will not restore them, `refresh` without it (or `recreate`) will. ## `?depth=` means module hops (item 65) On `db-accesses` / `sql-statements` (and the `?module=` scope of `variables/{name}/reads|writes`), `depth=N` means **N module calls away** — the same unit `/modules/{name}/graph?depth=` and `call-tree` neighbours use. `depth=1` = the modules this one directly `CALLNAT`s, regardless of how deeply the calling statement sits inside subroutines. Before item 65 these endpoints bounded the traversal on raw `CALLS` edges. A `CALLS` edge starts at the *statement* making the call, not at the `MODULE` node, so the traversal also stepped through internal `PERFORM` jumps and `depth` measured **statement nesting**, not dependency distance. Concretely: `WGEAGB0S` reached `YGEAGBNH`'s tables through two module calls, but the raw path is 5 edges (`WGEAGB0S → GET-DATA → GET-MAIN-DATA → BGEAGFN0 → READ-FILE → YGEAGBNH`), so `db-accesses?depth=2` returned `[]` — reading as "no DB access" — and only `depth=5` was truthful. **If you scripted a depth workaround** (a deliberately large `depth` to compensate), drop it: `depth` is now the value you'd naturally expect, and inflated values just widen the result set. > **`call-tree`'s `depth` column is still raw-hop based** and mixes internal subroutines into the tree: > a direct dependency called from the main body shows `depth=1` while one called two subroutines deep > shows `depth=3`. Use the ego graph (`/modules/{name}/graph`) when you need module-level distance. Tracked as an open > roadmap item. ## Framework-mediated DB access via `INCLUDE` macros (item 44) Natural's generic table-access framework hides a `CALLNAT` inside a copycode member, invoked with a statement-level macro: ```natural INCLUDE YFRAMGC0 'YFRAML01.C-MOD-GET' '"CO-TABLA-CO-ELEMENTO-ALT"' '"YELEMGN0"' 'YELEMKEY' 'YELEMREC' ``` The `CALLNAT` to the generic accessor (`YELEMGN0`) lives in the copycode, not in the including module, so before item 44 both `callees` and `db-accesses` were empty for such modules. The parser now recognises the framework macro and emits a `CALLS` edge to the accessor named in the macro arguments (de-quoted; e.g. `'"YELEMGN0"'` → `YELEMGN0`), tagged with **`edgeKind = INCLUDE_MACRO`** on `callers`/`callees`. Because the edge is a normal `CALLS`, the accessor's own table access surfaces transitively: `GET /modules/{name}/db-accesses?depth=N` reports the table with `via` = the accessor module. The recognised macros and which argument names the accessor are described declaratively in `FrameworkMacros` (ac-parser-natural). Scope: the targeted recogniser only — general `.nsc` copycode expansion is still open. **`db-accesses?depth=N` is a superset of `db-accesses` (item 93).** Besides `READS`/`WRITES` it also returns the `mode: "DECLARES"` rows — a Java entity's own `MAPS_TO` table and a repository's `repositoryEntity` table (item 32) — for every module in the closure, with `via` naming the declaring module. Before item 93 the transitive query carried only the `READS`/`WRITES` branch, so asking the *same* module with `depth` dropped its declared table and a Java caller's transitive `db-accesses` came back empty although the entity it persists through maps to a real table. **Natural view aliases are resolved to the underlying table (item 95).** A Natural DML statement names a *view variable* (`1 VDB2-VERSIS_LITERALES VIEW OF VERSVW_LITERALES`), not the DDM. `db-accesses` reports the **table** — `FIND VDB2-VERSIS_LITERALES`, `FIND NUMBER NEXT-VIEW` and `STORE VDB2-VERSIS_LITERALES` in `YLITEMN0` all come back as `VERSVW_LITERALES`, matching the SQL `SELECT … FROM` rows in the same module. Before item 95 the alias itself was the reported name, which (a) split one table across several names, (b) made the generator's boilerplate alias `NEXT-VIEW` a single node shared by 11 modules meaning 11 different tables, and (c) hid every `VERSVW_LOGFILE` write behind 11 `VDB2-*-VLOG` aliases. Table names are upper-cased (Natural is case-insensitive). **…including aliases declared in a `USING` data area (item 98).** A view is often declared not in the module but in a `LOCAL USING` area, in the data-area *export* form (`V 1VDB2-VERSIS_GENAGREE VERSVW_GENAGREE …` — no `VIEW OF` text). Those resolve too: `YGEAGBNH`'s `FIND (1) VDB2-VERSIS_GENAGREE` reports `VERSVW_GENAGREE`. Resolution is scoped to each module's own `USING` set, never by name — alias names are boilerplate, and `NEXT-VIEW` alone is declared over 100 different tables in `upms`. A module whose `USING` areas give two different tables for one alias is left unresolved rather than guessed. **Natural `UPDATE(