Files
agenticCode/x-docs/mcp-api-usage-ac-implementation.md
Ingo Schnabel d4b14a3b6d Bug fixes
2026-07-27 16:20:42 +02:00

66 KiB
Raw Blame History

Using AgenticCode on This Repo (Dogfooding)

This repo is ingested as project ac at http://localhost:8787. Per CLAUDE.md's dogfooding rule, prefer the AgenticCode API over grep/Explore for call graphs, callers/callees, DB access, dataflow, and module overviews when working on this repo — it's exactly the tool this project builds.

For full API semantics (params, response shapes, error codes, language applicability, MCP tool mapping, curl examples), see x-docs/agent-api-system-prompt.md — that file is the canonical reference and is not duplicated here. This file only covers what's specific to using the API as Claude Code, on this checkout.

Tool priority

MCP first (mcp__agenticcode__*) → REST (GET /api/projects/ac/...) → ac CLI (ac callers, ac callees, ac call-tree, ac context, ac db-accesses, ...) → grep/Explore. Only fall back past REST/CLI when the question genuinely isn't answerable by this API at all (see "Missing capability" below). If an MCP call fails (session/protocol error), fall back to REST rather than abandoning the API. If the server is unreachable, try ./manage-ac.sh deploy before falling back further.

Re-ingest before trusting results

Query results reflect the last ingest, not the current working tree. Refresh after code changes before trusting query results: ac refresh or POST /api/projects/ac/refresh (add --deep / ?deep=true for a full field-level pass). refresh is the single (re-)ingest surface (item 42) — the eager ingest_all/ingest_module/ingest-call-graph tools were removed. ac refresh <name> deep-ingests one module + its callees/data areas; add --neighborhood (POST /refresh/{name}?scope=neighborhood) to also pull in the module's transitive callers (whole call-graph neighbourhood). refresh is REST + CLI only — mutations aren't exposed as MCP tools.

Reconciliation on re-ingest (item 58). A refresh now purges stale nodes: for every re-parsed file it deletes the nodes the fresh parse no longer produces (renamed/removed fields, moved statements) rather than leaving them to shadow the new ones — so identifier counts and search_identifier results stay clean after a parser change or an edited source file. Applies to every full-parse path (whole-root refresh with or without --deep, and per-module refresh/{name}); the coarse Tier-1 scan run at project creation does not reconcile (it only ever runs on an empty graph).

Auto-invalidation (item 43). The graph self-heals when files change or are deleted on disk — you rarely need a manual refresh for staleness:

  • Changed file: a field-level query on a module whose source changed re-ingests it automatically (its stored sourceHash no longer matches), so results reflect the current file. Granularity is per-module: a changed dependency is picked up when that dependency is itself queried by name.
  • Deleted file: a whole-project refresh now also removes nodes for files deleted from disk (reconciled against the filesystem), closing the earlier gap — no project re-create needed. (A per-module refresh/{name} does not sweep; it only touches its own tree.)
  • Controlled by agenticcode.auto-invalidate.enabled (default true).

Reading source in this repo

You have direct filesystem access to this checkout — never call nodes/{id}/source or modules/{name}/source. Use sourceFile/startLine/ endLine from a graph response (context, digest, search/identifier, nodes/{id}, ...) and read the file directly. This is always cheaper and gives full surrounding context; the /source endpoints exist only for API-only agents with no filesystem access.

Stale-source check (item 41). The /source endpoints compare the file on disk against the content hash (sourceHash) stored at ingest. If the file changed since the last ingest they return 409 STALE_SOURCE (MCP: a STALE_SOURCE tool error) instead of slicing current text against old line numbers — re-ingest (refresh) the project to update the graph. Line ranges you read directly off disk are of course always current; this only guards the API's own slicing. Copycode/INCLUDE slices are raw pre-expansion file text.

Tier-1 coarse scan on project create (item 36)

Creating a project (POST /api/projects/{p}) now runs a Tier-1 coarse reference scan of the root before returning, so the project is immediately queryable — no separate ingest call. The scan is lexer-level (Natural) / a coarse projection of the full parse (Java): it records module/function /data-structure shells, the identifier index (declared fields, class members), and coarse CALLS/READS/WRITES/INCLUDES references (Natural PERFORM, CALLNAT '...', dynamic CALLNAT PGM-VAR, PARAMETER/LOCAL USING copybooks; Java resolved calls/type refs) — but no deep bodies (control flow, statement-level dataflow, arg→param). Scanned modules land CALL_GRAPH/NOT_INGESTED; field-level detail is filled in by the on-demand deep ingest below. Each module shell carries a sourceHash. Disable with agenticcode.tier1.scan-on-create=false (creates an empty project shell to scan later). A missing/unscannable root is non-fatal — the project is created empty and can be re-scanned.

Unresolved references (item 40)

A reference whose target isn't (yet) ingested — a CALLNAT/PERFORM/USING to a module/copybook absent from the project, or a dynamic CALLNAT PGM-VAR whose literal can't be recovered — is stored as a deduped placeholder node (blank sourceFile). Enrichment stamps each with an unresolved boolean: true while genuinely dangling, false once a real definition of that name is ingested. search_identifier returns it as unresolved on each IdentifierMatch, and inspect_node (GET /nodes/{id}) carries it in the node's properties — so an agent can tell a dangling/dynamic reference apart from a resolved one. In callers/callees such targets already appear as entries with a blank sourceFile.

Data literals are not call targets (item 62). A CALLNAT <bareword> whose target is really a data value — a browse key reaching the call site through a copycode/macro argument — used to leave a permanent unresolved placeholder callee. Enrichment now reaps such a placeholder when it has no Natural sigil (#/&/+), is a real VARIABLE/CONSTANT of the project, matches no real MODULE, and is only ever reached by inferred (CALLNAT_DYNAMIC/INCLUDE_MACRO) edges. So callees, call-tree, the ego graph and the unresolved list no longer list data names as modules. Genuine dynamic dispatch (CALLNAT #PGM-VAR, sigil'd) is still reported as an unresolved target, and a static CALLNAT 'X' is always trusted even when X collides with a field name.

The ingest summary agrees with the graph (item 64). A refresh/refresh/{name} response's unresolved list is built during the file walk, independently of the graph — before item 64 it therefore reported data fields as missing modules (MODULE CO-TABLA, MODULE NAME-DESC-SP) that enrichment had already reaped, so the summary and the graph disagreed. It now applies the same item-62 gates as the graph reap, with the "is this a real field?" test asked project-wide (a data literal is often declared outside the ingested tree). Genuinely missing modules are still reported — including sigil'd dynamic dispatch (MODULE #GETSHORT-MODUL) and a missing module whose name collides with a field name but is called statically (the RPC-CNTX class). Treat unresolved as "dependencies that really are absent".

Constant-folded string-assembled targets (item 83). A dispatcher often builds the CALLNAT <var> name from a base literal plus one or more SUBSTR overlays — e.g. #GETSHORT-MODUL in YGEAGGNH, assembled by MOVE 'YGEAGKEY' TO #GETSHORT-MODUL then MOVE 'GN0' TO SUBSTR(#GETSHORT-MODUL,6,3) → YGEAGGN0. The parser records each SUBSTR write as a WRITES carrying substrPos/substrLen (1-based), and the resolve-dynamic-callnat-fold enrichment step folds the last full-var literal written before the call site with the intervening overlays (left/substring) and MERGEs a resolved CALLS edge (callKind=CALLNAT_DYNAMIC, folded=true) to the assembled module when it is a real MODULE. It runs in every ingest mode (cheap, bounded by dynamic call sites) and over-approximates like the other dynamic resolvers. This auto-recovers the Y…GNH → Y…GN0 family with no manual override, so folded sites drop out of dynamic-calls/unresolved and the assembled target appears in callees/call-tree/graph/ego graph tagged CALLNAT_DYNAMIC.

Pin what the resolvers can't: manual dynamic-CALLNAT overrides (item 82). Some CALLNAT <var> targets are beyond the auto-resolvers — a name never present as a recoverable literal in the reachable code (read from a work-file record, supplied by an unlinked caller, …). Such a site stays an unresolved placeholder. A human or agent resolves it via POST /api/projects/{p}/dynamic-calls/overrides with the call site's originFile + lineNo (from GET .../dynamic-calls/unresolved) and the target module name(s) — multiple targets for a genuine branch. The override is stored as a :DynamicCallOverride node outside the :AstNode graph, so a refresh never deletes it and an enrichment step (apply-manual-dynamic-callnat, after the auto dynamic-CALLNAT resolvers, before the placeholder cleanup) re-applies it automatically — MERGEing a CALLS edge (callKind=CALLNAT_DYNAMIC, resolvedBy='manual') to each target and flagging the placeholder manualHidden so callees/digest/graph/call-tree show the real target, not the #var. It only applies while the site is still unresolved: once an auto-resolver catches up, the override is skipped and listed obsolete — except a constant-fold (item 83), which a manual override outranks: the fold skips a site carrying a :DynamicCallOverride, and a stale folded edge there is dropped (delete-folded-overridden-dynamic-callnat) before apply-manual-dynamic-callnat runs, so the pinned target replaces it. DELETE .../dynamic-calls/overrides?originFile=&lineNo= resets one site (omit both = all), clearing the manual edges and un-hiding the placeholder inline (no refresh). A target that is not a real MODULE is rejected 400 UNKNOWN_TARGET. Bug B fix: the callees items now carry unresolved (mirroring what graph already exposed), so an unresolved dynamic target is machine-distinguishable from a resolved one without inspecting sourceFile.

Dispatch guards: read guards — it is the only complete condition (item 72). A dispatch-table row's guardField/guardValue/guardValues describe the innermost DECIDE only. Natural nests value-DECIDEs inside each other (824 of upms's 7,729 blocks, 178 modules), and then the innermost guard is just one conjunct: in VMULTMN4, the row for YTABLMA0.TX-TABLA reports #FIELD-NAME = 'TX-TABLA', but the assignment also requires #SHORT-VIEW = 'TABL'. guards is the full chain — [{field, values}], outermost first, joined by AND, each link's values joined by OR. Reading only the legacy fields over-generalises: port that to Java and you get a branch firing where Natural never would. For an unnested DECIDE the chain has one link and says the same as the legacy fields.

  • Still incomplete for NONE/ANY branches (item 73): an assignment in a NONE branch is reported under its enclosing chain alone, but its real condition is "enclosing guard AND NOT any sibling VALUE" — a negation a chain of equalities cannot express. guards is strictly better than the legacy fields, not a total answer.

Within one guard: prefer guardValues over guardValue (item 64). dispatch-table rows carry both. guardValue is lossy and kept only for compatibility: it comma-joins the branch's VALUE literals, which drops blank alternatives entirely and, for a multi-alternative branch, yields a synthetic string the guarded field never equals ("A1, A2"). guardValues is the faithful list — every alternative in source order, blanks included — so VALUE 'GENAGREE-WOUT-SP', ' ' reports ["GENAGREE-WOUT-SP", " "], recording that a blank guard field also routes into that branch. When reasoning about routing (or porting a DECIDE to Java), read guardValues; guardValue will silently under-report the branch's conditions.

Text that is not code never yields a call (items 61 & 63). The CALLNAT/CALLNAT_DYNAMIC patterns are unanchored (a CALLNAT may legally appear mid-line), so both ingest tiers first neutralise non-code text: a full-line * comment is skipped, a trailing /* … is stripped (item 61), and a match whose keyword falls inside a quoted string literal is rejected (item 63). A real CALLNAT 'MOD' is unaffected — its keyword sits outside the quotes. This matters for trusting callers/callees/ call-tree: before item 63, prose such as WRITE(#MSG) 'NACH CALLNAT ISINGEAG:' or #ERR-TYPE := 'Callnat USIA008N' fabricated a CALLNAT_DYNAMIC edge to the real module of that name, so a mere log message appeared as a genuine call — and, because a real module existed, it was not flagged unresolved and could not be reaped by the item-62 cleanup. If you query a graph ingested before 2026-07-16, re-ingest (ac refresh) before trusting call-graph edges into modules that are also mentioned in log/error text.

LoC / SLoC metrics (item 46)

Every file-level node (a MODULE program/class, or a DATA_STRUCTURE for a Natural .lda/.pda data area) is stamped at ingest with two deterministic line metrics:

  • loc — physical lines of the file (language-independent; a trailing newline adds no phantom line).
  • sloc — source lines of code: non-blank, non-comment lines, computed per language. Natural drops full-line */**//* and inline /* comments; Java drops // and /* … */ blocks while keeping those tokens when they appear inside string literals.

Both the cheap Tier-1 coarse scan and the deep Tier-2 parse use the same per-language counter, so a module's loc/sloc are identical at any ingest depth — you can sum them to get exact, reproducible project totals.

Where to read them:

  • list_modules (GET /modules) — loc/sloc on each row.
  • module_context (GET /modules/{name}/context) — loc/sloc on the module.
  • inspect_node (GET /nodes/{id}) — loc/sloc in the node's raw properties.
  • project_loc (GET /loc, ac loc) — the rollup: a per-language breakdown (fileCount, loc, sloc) plus a project-wide total, optionally narrowed by ?language= / ?sourceFile=. Each source file is counted once even when it yields several nodes (Java inner classes, Natural inline groups).

null metrics mean the node predates item 46 — re-ingest (refresh) to backfill.

Generated vs. user-exit split (item 47)

A project can be created with a source language (required at creation; an attribute only — ingest still classifies files by extension) and a generatedDir/userExitDir pair (directory names, matched as path components like excludeDirs; both or neither). Generated modules already contain their hand-written user-exit twin inline, so at ingest a module under generatedDir whose name also occurs under userExitDir is annotated with that twin's LoC/SLoC (userExitLoc/userExitSloc). User-exit files are not ingested as standalone modules (they would collide by name) — the walk skips userExitDir.

Consequence for all non-LoC analysis: the generatedDir copy is the canonical, sole module for every structural query (call graph, DB access, functions, data structures, identifiers, dataflow, dispatch table). userExitDir exists only to compute the generated-vs-manually-written LoC split below; it never contributes nodes/edges. So when verifying an API response against source for a Natural module, always read the generatedDir file (e.g. generated_src/subprogram/WGEAGB0S.nat), not the user_exit fragment.

project_loc (GET /loc, ac loc) then reports, per language row and in the project total:

  • loc/sloc — the total (generated, which already includes the user exits).
  • userExitLoc/userExitSloc — the sum of the annotated user-exit twins (the hand-written part).
  • generatedExclusiveLoc/generatedExclusiveSloc — total − user-exit, clamped ≥0 per file (the purely generated part).

All three are 0 for projects without a generated/user-exit split. Create with ac project create <name> <root> -l natural -g generated_src -u user_exit, or add the split to an existing project via ac project update <name> -g generated_src -u user_exit.

?depth= means module hops (item 65)

On db-accesses / sql-statements (and the ?module= scope of variables/{name}/reads|writes), depth=N means N module calls away — the same unit ego_graph?depth= and call-tree neighbours use. depth=1 = the modules this one directly CALLNATs, regardless of how deeply the calling statement sits inside subroutines.

Before item 65 these endpoints bounded the traversal on raw CALLS edges. A CALLS edge starts at the statement making the call, not at the MODULE node, so the traversal also stepped through internal PERFORM jumps and depth measured statement nesting, not dependency distance. Concretely: WGEAGB0S reached YGEAGBNH's tables through two module calls, but the raw path is 5 edges (WGEAGB0S → GET-DATA → GET-MAIN-DATA → BGEAGFN0 → READ-FILE → YGEAGBNH), so db-accesses?depth=2 returned [] — reading as "no DB access" — and only depth=5 was truthful.

If you scripted a depth workaround (a deliberately large depth to compensate), drop it: depth is now the value you'd naturally expect, and inflated values just widen the result set.

call-tree's depth column is still raw-hop based and mixes internal subroutines into the tree: a direct dependency called from the main body shows depth=1 while one called two subroutines deep shows depth=3. Use ego_graph when you need module-level distance. Tracked as an open roadmap item.

Framework-mediated DB access via INCLUDE macros (item 44)

Natural's generic table-access framework hides a CALLNAT inside a copycode member, invoked with a statement-level macro:

INCLUDE YFRAMGC0 'YFRAML01.C-MOD-GET' '"CO-TABLA-CO-ELEMENTO-ALT"' '"YELEMGN0"' 'YELEMKEY' 'YELEMREC'

The CALLNAT to the generic accessor (YELEMGN0) lives in the copycode, not in the including module, so before item 44 both callees and db-accesses were empty for such modules. The parser now recognises the framework macro and emits a CALLS edge to the accessor named in the macro arguments (de-quoted; e.g. '"YELEMGN0"' → YELEMGN0), tagged with edgeKind = INCLUDE_MACRO on callers/callees. Because the edge is a normal CALLS, the accessor's own table access surfaces transitively: GET /modules/{name}/db-accesses?depth=N reports the table with via = the accessor module. The recognised macros and which argument names the accessor are described declaratively in FrameworkMacros (ac-parser-natural). Scope: the targeted recogniser only — general .nsc copycode expansion is still open.

db-accesses?depth=N is a superset of db-accesses (item 93). Besides READS/WRITES it also returns the mode: "DECLARES" rows — a Java entity's own MAPS_TO table and a repository's repositoryEntity table (item 32) — for every module in the closure, with via naming the declaring module. Before item 93 the transitive query carried only the READS/WRITES branch, so asking the same module with depth dropped its declared table and a Java caller's transitive db-accesses came back empty although the entity it persists through maps to a real table.

Natural view aliases are resolved to the underlying table (item 95). A Natural DML statement names a view variable (1 VDB2-VERSIS_LITERALES VIEW OF VERSVW_LITERALES), not the DDM. db-accesses reports the table — FIND VDB2-VERSIS_LITERALES, FIND NUMBER NEXT-VIEW and STORE VDB2-VERSIS_LITERALES in YLITEMN0 all come back as VERSVW_LITERALES, matching the SQL SELECT … FROM rows in the same module. Before item 95 the alias itself was the reported name, which (a) split one table across several names, (b) made the generator's boilerplate alias NEXT-VIEW a single node shared by 11 modules meaning 11 different tables, and (c) hid every VERSVW_LOGFILE write behind 11 VDB2-*-VLOG aliases. Table names are upper-cased (Natural is case-insensitive).

Natural UPDATE(<label>.) / DELETE(<label>.) count as writes (item 96). These act on the current record of the labelled FIND/READ loop, and are reported as WRITES on that loop's table. This is what makes the Y****MN0 access layer's update/delete path visible: YLITEMN0 reports WRITES VERSVW_LITERALES at the STORE and at UPDATE(HOLD-PRIME.) / DELETE(HOLD-PRIME.), where before item 96 it reported only the STORE — reading, wrongly, as an insert-only layer. A reference that resolves to no labelled loop (an unknown label, or the numeric source-line form) records nothing rather than guessing a table.

call-tree/graph agree with callees about overridden dynamic calls (item 97). A manual dynamic-call override hides the placeholder marker rather than deleting it. All read paths now filter it, so a pinned CALLNAT <var> shows the real target and never the variable name. Everything driven by the call-tree BFS — graph, db-accesses?depth=N, sql-statements?depth=N — inherits this.

callers on a dynamically-called module is an over-approximation, and says so. A Natural web-service module is reached by CALLNAT #WIF, resolved by naming pattern: W-LST-N0.nat:362 alone resolves to 29 W****B*S/W****X*S targets, so WGEAGB0S lists W-LST-N0 and W-MNT-N0 as callers. The rows are tagged edgeKind: "CALLNAT_DYNAMIC" — treat those as may-call, not does-call, and check dynamic-calls/overrides / dynamic-calls/unresolved when the distinction matters.

XML payload / interface schema (item 45)

Natural XML wrapper subprograms build a wire payload by mapping data-area fields to XML tags via the ADD-XML-LINE idiom (#W-TAG := '<tag>' / #W-VALUE := <field> / PERFORM ADD-XML-LINE, where the subroutine COMPRESSes '<' #W-TAG '>' #W-VALUE). The deep parser extracts that contract as PAYLOAD_FIELD nodes and exposes it:

  • module_payload (GET /modules/{name}/payload, ac payload <module>) → an array of {tag, field, direction, lineNo, sourceFile} triples. direction is REQUEST for an emitted (outbound) field. field is the unqualified payload field name (WXMLIN.P-COD-USUARIO → P-COD-USUARIO). sourceFile is the file lineNo refers to — the module's own file for source=IDIOM, or the interface PDA's file for source=PDA (so a caller opens the right file at the line, not the module at a stray line).

  • source=IDIOM (item 45): extracted from a static ADD-XML-LINE emit sequence with literal tags.

  • source=PDA (item 46b): the module is a generic, runtime-driven serializer (it calls the YFRAMN07 tag-builder or has an ADD-XML-LINE/ADD-XML-ACT subroutine) with no static tag list in its source — real production wrappers like WNAUTD0S are this shape. The contract is then derived from the module's PARAMETER USING interface PDA: each field is a payload field, the wire tag is the field name with the framework's EXAMINE … '#' REPLACE '_' normalisation applied (#→_), direction REQUEST. Idiom fields take precedence when both exist.

The static idiom also handles derived tags (EXAMINE #W-TAG FOR '#' REPLACE '_' → #→_) and both directions: an ADD-XML-LINE-style emit sub is REQUEST; a GET-XML-LINE-style parse sub (with the reverse field := #W-VALUE binding) is RESPONSE.

Empty for modules that neither use the idiom nor are a flagged XML wrapper (or are only coarse-ingested).

Copycode (.cpy) expansion (item 46a)

Natural INCLUDE <member> <args> is a compile-time macro: the copycode body is spliced into the including module (with positional &1&… substitution), so a copycode's CALLNAT/PERFORM, DB access and dataflow live in the copycode, not the module. The deep and coarse parsers now expand statement-level copycode includes before parsing, so those constructs surface on the including module — e.g. a READ or CALLNAT that only exists in a .cpy shows up in the host's db-accesses/callees.

  • Copycode-origin nodes/edges report the real .cpy file + line (so navigation lands in the copycode), and carry viaCopycode=<member> + includedAt=<host line>; host statements keep their own file + line (line numbers are remapped after the splice, never shifted).
  • A line number alone is not a location (item 66). Because of the above, one module's calls and field accesses come from more than one file, and host and copycode lines are freely mixed — so always read the line together with the file the endpoint gives it:
    • callers/callees/functions/{f}/callers return sites: [{lineNo, callSiteFileIndex, viaCopycode, includedAt}] (not a bare lineNos array). callSiteFileIndex indexes sourceFiles and is the file the call is written in; the entry's own sourceFileIndex is a different thing — the file the named module/function is defined in. viaCopycode/includedAt are set only for copycode sites.
    • variables/{name}/reads|writes return sourceFile = the file lineNo is in (the .cpy for a copycode access), plus viaCopycode + includedAt. Before item 66 the copycode's line was paired with the host's file: #W-OPTIONS writes in WGEAGB0S were reported at WGEAGB0S.nat:18/20/22, which is its generated comment banner — the writes are really ISICINDI.cpy:18/20/22. Use includedAt when you want the spot in the host module instead.
    • db-accesses / workfile-accesses return sites: [{lineNo, sourceFile, viaCopycode, includedAt}] alongside the (kept, backward-compatible) lineNos array — one entry per statement, each tying its line to the file it truly lives in. sql-statements gains sourceFile + viaCopycode on each statement (its startLine/endLine are lines in sourceFile). Before this, a DB/work-file access written in an INCLUDEd copycode reached the API as a bare copycode-local lineNo with nothing to attribute it to — e.g. the DB2 sequence read SELECT … FROM SYSIBM-SYSDUMMY1 lives in USIX043C.cpy at lines 31/39/45/51/57, but db-accesses for the 9 including modules (YAPRFMN0, YUGRPMN0, …) reported those as bare line numbers that land on the host's own comment/DEFINE DATA lines. The sites file context is the same fix item 66 applied to variables/reads|writes and callees.
  • The same line number can legitimately appear twice (item 69). A host statement on line 10 and a copycode statement on line 10 are two different statements, and both are returned — as separate entries differing only in their file. Until item 69 the graph could not hold both: an edge was identified by (source, target, type, lineNo) with no file, so the second one overwrote the first and a real access was missing from every answer. Treat (file, lineNo) as the identity of a site, never lineNo.
  • call-tree's depth counts module hops (item 67). depth is how many module boundaries the shortest call path crosses, not raw CALLS edges — a CALLNAT made from two subroutines deep is still one hop. The root module's own subroutines are therefore depth 0. Measured on upms: WGEAGB0S's seven direct dependencies used to report depth 1..3 (BGEAGFN0 was 3); all seven now report 1.
    • Results at a given depth are larger than before. A subroutine of a module within depth hops is now inside the bound, because it crosses no further boundary. Previously call-tree?depth=1 could hide a DEFINE SUBROUTINE of the very module you asked about, just because it was PERFORMed from another subroutine (raw depth 2) — that is the same bug seen from the inside.
    • call-tree also returns truncated. true means the intra-module subroutine walk stopped at its raw-hop budget, so some FUNCTION items may be missing — not that your depth was exceeded (that is a normal, complete answer). It is conservative and can be true for a complete result. Tune via agenticcode.call-tree.internal-budget (default 20; the deepest internal chain observed in upms is 9).
    • Since item 94 the budget cannot hide a module. MODULE rows come from the same module-hop BFS that db-accesses/sql-statements use, so a callee one hop away is always listed even when its call site sits behind a long internal PERFORM chain (before item 94 it was dropped, and call-tree then contradicted db-accesses). This also removed the path enumeration that made call-tree?followWiring=true time out on Java projects at depth ≥ 2; followWiring is now usable at full depth.
  • field-flow's depth counts module hops (item 68). Like db-accesses/sql-statements (item 65), variables/{name}/field-flow?depth=N now means "up to N module calls apart", not N raw CALLS edges. Before item 68 a consumer called from inside a subroutine sat several raw hops away and was dropped at depth=1, so the endpoint answered "nothing downstream consumes this field" — read that answer with suspicion on any graph ingested before this change.
  • field-flow no longer fabricates flows between same-named fields (item 77). A bare field reference is resolved against the referencing module's own USING includes. It used to be resolved project-wide: an unresolved bare field is one shared node per (name, project), and the resolver aggregated over all owning modules at once, so (a) two modules including different data areas that both declare the name left both unresolved on the shared node, and (b) a module with no matching include was redirected onto another module's field. Either way the two modules ended up on one node, and field-flow — which pairs a producer with a consumer only when both touch the same node — reported a dataflow between modules that share nothing but a field name. In upms: 38 + 28 placeholders affected (199 module-field pairs). Read any pre-item-77 field-flow result for a common field name with suspicion, and note the answer only changes after a deep re-ingest, since resolution runs there. reads/writes are unaffected — they match every node with the name and report only the accessing side, so they never distinguished the targets in the first place. Residue (item 76): a bare field shared via copycode (132 of 18,539 source nodes in upms) is still one node for several modules; per-module identity is a schema change, not a query fix.
  • Copycode provenance survives field resolution (item 70). viaCopycode/includedAt are now kept for fields addressed by qualified name (MYLDA.Q-FIELD, i.e. a field of a LOCAL USING data area) as well as bare ones. Before item 70 only bare references kept it; qualified ones silently came back with viaCopycode: null and the host file, i.e. they looked exactly like host statements.
  • Excluded from expansion: framework macros (handled by the item-44 targeted recogniser), data-area USING includes, unknown members, and any copycode that declares DEFINE DATA. Recursion is cycle-guarded. Copycodes (.cpy) are not standalone modules — they enter the graph only through the including module.
  • Staleness caveat: the item-41/43 hash check hashes the host file, so auto-invalidation triggers on a change to the host — but a change to an included .cpy alone (host unchanged) is not detected; re-ingest the host (refresh/{host}) to pick it up.

Global Data Areas (.gda) (item 46c)

.gda files are now ingested as DATA_STRUCTUREs like .lda/.pda, and DEFINE DATA GLOBAL USING <gda> resolves to them (the INCLUDE/USING recogniser now accepts GLOBAL, not just PARAMETER/LOCAL).

Deep-ingest: now automatic (lazy Tier-2)

Field-level endpoints (flow-forward, flow-backward, field-flow) and cross-module dynamic CALLNAT resolution need a per-module deep ingest, not just a whole-root refresh. This deep ingest is now triggered automatically on demand: calling a field-level endpoint for a module that is only CALL_GRAPH-ingested runs a scoped deep ingest of that module (and its dependency tree) transparently, then returns the resolved result — no 409, no manual POST /refresh/{name} step. The first such call to a cold module is therefore slower (it walks the root and parses the program tree); subsequent calls hit the already-FULL graph.

The deep ingest is best-effort: if the module cannot be resolved to a source file, the endpoint still falls back to the 409 NOT_DEEPLY_INGESTED / NOT_INGESTED hint with a nextAction rather than a misleading empty result.

Flow path-ingest (auto, cross-module fixpoint). flow-forward, flow-backward, and field-flow go one step further than the single start-module deep ingest: after deep-ingesting the start module they run an ingest-and-re-traverse fixpoint. Each round deep-ingests the frontier — the modules the trace surfaced together with their direct callee modules — in one scope, then re-traverses. This is what lets a dataflow trace cross into a dynamically-dispatched callee (CALLNAT PGM-VAR): that callee is not a static dependency of the start module, so it is only pulled in and linked (caller.arg → callee.param) once a round resolves the dynamic CALLS edge and ingests the target. The loop is bounded by agenticcode.deep-ingest.flow-rounds (default 3) and the per-round agenticcode.deep-ingest.fanout-nodes budget, and stops early (fixpoint) as soon as a round pulls in nothing new — so on an already-deep graph a flow query costs one traversal plus one cheap frontier check, no re-run.

Fan-out warm (auto, on the result set). The fan-out / traversal queries callers, search_identifier, and call-tree also auto-deep-ingest — but on the set of modules their result surfaced, not a single named module. Each runs against the graph as-is, deep-ingests the surfaced modules (blocking, bounded by the fan-out node budget agenticcode.deep-ingest.fanout-nodes, default 50), and — only if that warm actually deepened something — re-runs so the response reflects newly-resolved dynamic dispatch (e.g. a call-tree grows to include a dynamically-dispatched callee once the surfaced program is deep). When everything is already FULL (or the warm resolves nothing) the query returns its first result with no redundant re-run. Note callers warms the already-surfaced callers, so it improves downstream precision but cannot reveal a caller that was invisible at the coarse (call-graph) tier. Other module-level endpoints (context, callees, db-accesses) work regardless of ingest depth and do not trigger a deep ingest.

Bounded fan-out. A by-name deep ingest walks the transitive dependency tree breadth-first, bounded by maxDepth (hops from the named module, default 5, ceiling 20) and maxNodes (files, default 300). When a bound is hit the walk stops early and the ingest response carries a truncation object ({reason: DEPTH|NODES|NODES_AND_DEPTH, maxDepth, maxNodes, hint}) — the modules actually reached are marked FULL, the remainder stays as it was. Raise the limits on an explicit module refresh to pull in more: POST /refresh/{name}?maxDepth=&maxNodes=, MCP refresh(module, maxDepth, maxNodes), or CLI ac refresh <name> --max-depth --max-nodes. Auto-triggered ingests use the server defaults; if a field-level query returns partial data because the target's deep ingest truncated, re-run the explicit refresh with higher limits. (The auto-trigger does not yet accept per-query limit overrides.)

Durable ingest status + coalescing (item 36). Each real MODULE node carries a durable ingestStatus lifecycle — NOT_INGESTED (only its call graph is in the graph) → INGESTING (a deep ingest is in flight) → INGESTED (deeply ingested, ingestDepth = FULL) — separate from ingestDepth. When two calls trigger the same module's deep ingest at once they coalesce rather than both ingesting: within one process an in-process lock serialises them; across processes/restarts a best-effort DB claim marks the module INGESTING and a loser waits for the winner to reach FULL (re-claiming if the claim is released or goes stale after agenticcode.deep-ingest.ingesting-ttl-seconds, default 1800; wait bounded by claim-wait-seconds, default 120). A crash mid-ingest leaves the module re-triggerable (it never reached FULL), and the stale INGESTING is reclaimed on the next call. inspect_node / node_source expose ingestStatus/ingestStatusAt on the module node.

Warm concurrency cap (item 37). All auto deep-ingest/warm work (by-name, fan-out, and flow-frontier) shares a global permit pool (agenticcode.deep-ingest.max-concurrent-warms, default 2), so a burst of queries cannot spawn unbounded parallel parses/Neo4j writes. A permit is acquired only around the actual ingest; if none frees up within agenticcode.deep-ingest.warm-acquire-timeout-seconds (default 10) the warm is skipped and the query returns its Tier-1 (coarse) answer immediately rather than blocking — so under sustained load a query may transiently return shallower data; retry once load subsides, or force it with an explicit POST /refresh/{name}.

OpenAPI contract & CORS (items 48/50)

The server now ships an OpenAPI 3 spec (quarkus-smallrye-openapi): the machine contract the web-UI TypeScript client is generated against. All REST endpoints carry @APIResponse/@Schema annotations, so response bodies are typed in the spec even though the JAX-RS methods return raw Response. Access it at:

  • GET /q/openapi — YAML (or Accept: application/json for JSON)
  • GET /q/swagger-ui — interactive UI (dev)

CORS is enabled (quarkus.http.cors.enabled=true) and restricted to the UI's dev origins (http://localhost:5173, http://localhost:4173) — extend the quarkus.http.cors.origins list per deployment; never ship a wildcard.

Endpoint quick reference

Endpoint Use for
GET /modules?sourceFile=&moduleKind=&extends= List/filter modules; map a source file to its module name(s). Each row carries loc/sloc (item 46) and ingestStatus/ingestDepth (item 50) for status badges without a per-module round trip
GET /loc?language=&sourceFile= Per-language LoC/SLoC rollup (fileCount/loc/sloc) + project total; each file counted once (item 46). For a generated/user_exit project also userExitLoc/userExitSloc + generatedExclusiveLoc/generatedExclusiveSloc (item 47)
GET /modules/{name}/digest Tiny triage view before deciding which modules to expand
GET /modules/{name}/context One-shot overview: functions, callers, callees, DB accesses, SQL/variable summaries (?include= for full lists)
GET /modules/{name}/callers | /callees Direct callers/callees incl. EXTENDS/IMPLEMENTS/INJECTS/REFERENCES. callers scope: external (default) = modules that call this one (CALLNAT/inheritance), rolled up to the calling MODULE: a call made from inside a subroutine/method is attributed to its owning module (never the calling FUNCTION node), and repeated call sites from one caller collapse to a single row whose sites list every line — symmetric with how callees anchors its source side. internal = the module's own subroutines' PERFORM wiring (function-level). The default is external-only, module-typed only, and never lists the module as its own caller (no MODULE→MODULE self-loop); use scope=internal or /functions/{fn}/callers for intra-module / function-level wiring. callees is unchanged (default lists both external CALLNAT and internal PERFORM targets)
GET /modules/{name}/functions/{function}/callers FUNCTION-level callers (item 52): who PERFORMs (Natural) or calls (Java cross-class) a specific subroutine/method, with call-site lineNos. Finer-grained than the module-level /callers (which is module→module). Same CallRefResponse shape. MCP function_callers, CLI ac function-callers <module> <function>
GET /modules/{name}/call-tree?depth= Transitive call graph to scope a feature
GET /dynamic-calls/unresolved | /overrides · POST/DELETE /overrides Manual dynamic-CALLNAT overrides (item 82). unresolved lists open CALLNAT <var> sites {module, originFile, lineNo, variable}; POST /overrides {originFile, lineNo, targets[], variable?, note?} pins a site to real module(s) (applied at once, persisted across refreshes, 400 UNKNOWN_TARGET for a non-module); DELETE /overrides?originFile=&lineNo= resets one site (omit both = all) and restores the placeholder inline; GET /overrides lists them with an obsolete flag. MCP list_unresolved_dynamic_calls/list_dynamic_call_overrides/set_dynamic_call_override/reset_dynamic_call_override, CLI ac dynamic-calls unresolved|overrides|set|reset
GET /modules/{name}/graph?direction=&depth=&limit= Ego graph (item 49): bounded module-level call neighbourhood as nodes + edges (unlike call-tree). direction = out/in/both; limit caps nodes (BFS order) and sets truncated; unresolved targets carry unresolved=true + empty sourceFile. MCP ego_graph, CLI ac ego-graph
GET /modules/{name}/db-accesses | /sql-statements DB tables + mode, raw statement text (pass ?depth= for Natural). db-accesses items carry sites: [{lineNo, sourceFile, viaCopycode, includedAt}] (+ kept lineNos); sql-statements items carry sourceFile + viaCopycode — so a copycode-sourced access (e.g. SELECT … FROM SYSIBM-SYSDUMMY1 in USIX043C.cpy) reports the .cpy line, not a bare number that reads as a host-file line
GET /modules/{name}/workfile-accesses Natural work files (sequential/flat-file I/O — READ/WRITE WORK FILE n), the work-file analogue of db-accesses (item 84): [{workFile, physicalName, mode: READS|WRITES, recordBuffers, lineNos, sites}], aggregated per work-file number + mode. sites: [{lineNo, sourceFile, viaCopycode, includedAt}] gives each access its file context (copycode-aware), like db-accesses. physicalName comes from a DEFINE WORK FILE n '<name>', else null. Kept separate from db-accesses — a work file is not an ADABAS/SQL table (fixes a former bug where READ WORK FILE created a phantom DB_TABLE 'WORK'). MCP workfile_accesses, CLI ac workfile-accesses <module>
GET /modules/{name}/data-structures Which copybooks/inline groups a module uses
GET /modules/{name}/payload Natural XML wire-payload contract: {tag, field, direction, source, lineNo, sourceFile} — static ADD-XML-LINE idiom (source=IDIOM, item 45) or derived from the wrapper's interface PDA (source=PDA, item 46b). sourceFile is the file lineNo refers to (module for IDIOM, PDA for PDA)
GET /modules/{name}/dispatch-table Natural DECIDE ON VALUE OF routing table
GET /modules/{name}/functions?kind= | /functions/{fn}/overrides | /functions/overrides Method list, modifier filter (Java), subclass overrides (single/bulk). Each item carries sourceFile + viaCopycode (item 84): a Natural subroutine pulled in via INCLUDE reports the copycode file and viaCopycode:true, so its startLine/endLine are read as offsets into that copycode — not into the including module's own file (which is shorter). viaCopycode:false = declared inline. Always false for Java
GET /data-structures/{name}/fields | /db-tables/{name}/columns | /modules/{name}/columns Field/column schemas for DTO/entity generation
GET /variables/{name}/reads | /writes | /flow-forward | /flow-backward | /field-flow Impact analysis and dataflow tracing
GET /search/identifier | /search/value | /search/annotation Cross-project lookup by name / literal value / annotation. search/identifier matches the exact declared name but is sigil-insensitive: a leading Natural sigil (# user, & AIV, + GDA) is ignored on both sides, so name=K-OUT-MAX finds the declared #K-OUT-MAX (and vice-versa). Optional scope filters sourceFile=<relpath> and module=<name> (item 53) narrow the match to one file / one module — use them to pinpoint a module-local declaration when a name recurs across dozens of modules (the result is otherwise paginated and the local one may fall off the page). To keep the full cross-project list yet still guarantee a given module's own declaration is on the first page, pass priorityModule=<name> instead of module=: it does not filter, but pins that module's matches to the front (ahead of the otherwise sourceFile-ordered rest) so they survive the limit. This is what the web UI's click-to-identify sends for the open module. MCP search_identifier / CLI ac search-identifier --module --priority-module --source-file --type accept the same filters
GET /search/source?regex=&limit=&ignoreCase= (MCP search_source, ac search-source) Regex grep over module source text (item 54): {module, sourceFile, lineNo, line} hits + truncated. Case-insensitive by default. Complements search_identifier (declared names) — use for code patterns (statements, table names, literals)
GET /nodes/{id} Every property of one node (when a curated DTO is missing something)
GET /nodes/{id}/source | /modules/{name}/source | /source?file= Source text — only when you have no other access to the source (you always do in this repo, see "Reading source in this repo" above). module_source returns the whole file when the line range is omitted (M1), or a [startLine,endLine] slice when both are given. /source?file=<relpath> (MCP file_source, CLI ac file-source) serves a file by relative path rather than module name — for files that aren't standalone modules, e.g. a Natural data area (PDA/LDA) USING'd by a module, whose field line numbers refer to that file. Same whole-file/range + stale-source semantics; the client-supplied path is rejected (400 INVALID_SOURCE_FILE) if it escapes the project root

Full endpoint list, request params, and response field details: x-docs/agent-api-system-prompt.md.

Missing capability?

If the API/MCP/CLI genuinely cannot answer a question (not just unreachable — the capability doesn't exist), finish the task via grep/Explore as a fallback, then use AskUserQuestion to flag the gap and ask whether it should become a roadmap item in x-docs/roadmap.md. Don't silently fall back and move on.