81 KiB
Using AgenticCode on This Repo (Dogfooding)
This repo is ingested as project ac at http://localhost:8787. Per
CLAUDE.md's dogfooding rule, prefer the AgenticCode API over grep/Explore for
call graphs, callers/callees, DB access, dataflow, and module overviews when
working on this repo — it's exactly the tool this project builds.
For full API semantics (params, response shapes, error codes, language
applicability, curl examples), see
x-docs/agent-api-system-prompt.md — that file is the canonical reference and
is not duplicated here. This file only covers what's specific to using the
API as Claude Code, on this checkout.
Tool priority
REST first (GET /api/projects/ac/...) → ac
CLI (ac callers, ac callees, ac call-tree, ac context, ac db-accesses, ...) → grep/Explore. Only fall back past REST/CLI when the
question genuinely isn't answerable by this API at all (see "Missing
capability" below). If the server is unreachable, try
./manage-ac.sh deploy before falling back further.
Re-ingest before trusting results
Query results reflect the last ingest, not the current working tree.
Refresh after code changes before trusting query results:
ac refresh or POST /api/projects/ac/refresh (add --deep / ?deep=true for a
full field-level pass). refresh is the single (re-)ingest surface (item 42) — the
eager ingest-all/ingest-module/ingest-call-graph endpoints were removed.
ac refresh <name> deep-ingests one module + its callees/data areas; add
--neighborhood (POST /refresh/{name}?scope=neighborhood) to also pull in the
module's transitive callers (whole call-graph neighbourhood).
Reconciliation on re-ingest (item 58). A refresh now purges stale nodes:
for every re-parsed file it deletes the nodes the fresh parse no longer produces
(renamed/removed fields, moved statements) rather than leaving them to shadow the
new ones — so identifier counts and /search/identifier results stay clean after a
parser change or an edited source file. Applies to every full-parse path (whole-root
refresh with or without --deep, and per-module refresh/{name}); the coarse
Tier-1 scan run at project creation does not reconcile (it only ever runs on an empty
graph).
Auto-invalidation (item 43). The graph self-heals when files change or are deleted on
disk — you rarely need a manual refresh for staleness:
- Changed file: a field-level query on a module whose source changed re-ingests it
automatically (its stored
sourceHashno longer matches), so results reflect the current file. Granularity is per-module: a changed dependency is picked up when that dependency is itself queried by name. - Deleted file: a whole-project
refreshnow also removes nodes for files deleted from disk (reconciled against the filesystem), closing the earlier gap — no project re-create needed. (A per-modulerefresh/{name}does not sweep; it only touches its own tree.) - Controlled by
agenticcode.auto-invalidate.enabled(defaulttrue).
Reading source in this repo
You have direct filesystem access to this checkout — never call
nodes/{id}/source or modules/{name}/source. Use sourceFile/startLine/
endLine from a graph response (context, digest, search/identifier,
nodes/{id}, ...) and read the file directly. This is always cheaper and gives
full surrounding context; the /source endpoints exist only for API-only
agents with no filesystem access.
Stale-source check (item 41). The /source endpoints compare the file on
disk against the content hash (sourceHash) stored at ingest. If the file
changed since the last ingest they return 409 STALE_SOURCE instead of slicing current text against old line
numbers — re-ingest (refresh) the project to update the graph. Line ranges you
read directly off disk are of course always current; this only guards the API's
own slicing. Copycode/INCLUDE slices are raw pre-expansion file text.
Tier-1 coarse scan on project create (item 36)
Creating a project (POST /api/projects/{p}) now runs a Tier-1 coarse reference scan of the
root before returning, so the project is immediately queryable — no separate ingest call. The scan
is lexer-level (Natural) / a coarse projection of the full parse (Java): it records module/function
/data-structure shells, the identifier index (declared fields, class members), and coarse
CALLS/READS/WRITES/INCLUDES references (Natural PERFORM, CALLNAT '...', dynamic CALLNAT PGM-VAR, PARAMETER/LOCAL USING copybooks; Java resolved calls/type refs) — but no deep
bodies (control flow, statement-level dataflow, arg→param). Scanned modules land
CALL_GRAPH/NOT_INGESTED; field-level detail is filled in by the on-demand deep ingest below. Each
module shell carries a sourceHash. Disable with agenticcode.tier1.scan-on-create=false (creates
an empty project shell to scan later). A missing/unscannable root is non-fatal — the project is
created empty and can be re-scanned.
Unresolved references (item 40)
A reference whose target isn't (yet) ingested — a CALLNAT/PERFORM/USING to a module/copybook
absent from the project, or a dynamic CALLNAT PGM-VAR whose literal can't be recovered — is stored
as a deduped placeholder node (blank sourceFile). Enrichment stamps each with an unresolved
boolean: true while genuinely dangling, false once a real definition of that name is ingested.
/search/identifier returns it as unresolved on each IdentifierMatch, and GET /nodes/{id} carries it in the
node's properties — so an agent can tell a dangling/dynamic
reference apart from a resolved one. In callers/callees such targets already appear as entries
with a blank sourceFile.
Querying a module endpoint: three answers, not one (item 107)
Every GET /api/projects/{p}/modules/{name}/… endpoint used to answer 200 with an all-zeros shell
for a name that exists nowhere in the graph, byte-identical to a real-but-empty module's answer. That
is not cosmetic: call-tree?depth=4 returning 200 with 0 modules reads as "analysed, nothing
found" when the truth is "not analysable" — the module's source was never in the checkout. The
graph knows three states and each now gets its own status:
| state | answer | what it means |
|---|---|---|
no MODULE node for the name |
404 MODULE_NOT_FOUND |
unknown name (typo, or not in this project) |
| placeholder — node exists because something calls it, source never parsed | 409 + {status:"NOT_INGESTED", module, detail, nextAction} |
knowable in principle, not analysed yet |
| ambiguous — several real modules share the name (Java) | 409 + {code:"AMBIGUOUS_NAME", details:{candidates:[...]}} |
the name does not identify a module; pick one |
| real, ingested module | 200 |
the body is the answer — an empty body genuinely means "nothing found" |
Ambiguous names (item 115). A Java simple name is not unique: nested @Nested test classes,
Builder, Config, WorkingStorage. Measured on pur, 163 names covering 385 modules (~8%) were
addressable only ambiguously, and the endpoints used to answer with the union across unrelated
classes — /modules/BrokerHistoryTests/functions returned 627 functions for a class that has 107.
They now refuse and list the candidates. Repeat the request with ?sourceFile=<candidate>:
GET /modules/Shared/functions → 409 AMBIGUOUS_NAME, candidates ["a/Shared.java","b/Shared.java"]
GET /modules/Shared/functions?sourceFile=a/Shared.java → 200
A Java module's name IS its fully-qualified name (item 117) — com.example.OrderService, and
com.example.Outer.Inner for a nested class. That is what makes same-simple-name classes
distinguishable at all. Every module endpoint accepts either form:
GET /modules/com.example.OrderService/digest → 200, always exact
GET /modules/OrderService/digest → 200 when unique, else 409 AMBIGUOUS_NAME
Responses carry simpleName alongside name for display. The ?module= and ?extends= filters and
the ac CLI take either form too; --source-file remains available on every module command.
Which Java types are modules (item 119). Classes, interfaces, enums, records and annotation
types — moduleKind is one of CLASS | INTERFACE | ENUM | RECORD | ANNOTATION (Natural adds
PROGRAM | SUBPROGRAM | …), and ?moduleKind= filters on it. Their content is modelled the way each
kind carries it: a record's components and an annotation type's members are FIELDs (the latter with
defaultValue where declared), an enum's constants are CONSTANTs, and an enum's or record's
implements is a real edge — so a call against an interface fans out to an enum implementing it.
Before 119 these three kinds were not parsed at all: GET /modules/SomeEnum/digest answered 404,
and a record referenced from elsewhere stayed an unresolved placeholder (409 NOT_INGESTED). One
gap remains by design — a record's compact canonical constructor is not a function node, so calls
made in its body are invisible.
Natural is unaffected throughout: its module names are file stems, and colliding identities are
skipped at ingest, so they are unique by construction (upms has zero ambiguous names, pur 163).
Two limits worth knowing. Ambiguous names inside a traversal (a call tree's neighbours) are still
merged — only the request's root module is disambiguated. And a name skipped at ingest as a duplicate
identity still answers 404, not 409 (item 114).
The 409 applies to everything derived from the module's own source: digest, context,
call-tree, callees, db-accesses, workfile-accesses, sql-statements, functions,
functions/overrides, functions/{fn}/overrides, functions/{fn}/callers, data-structures,
dispatch-table, payload, columns.
callers and graph stay 200 for a placeholder — their data comes from the calling modules'
source and is genuine. When you get a 409, fall back to /callers: it is the one honest answer
available for a module whose own source is missing. (digest no longer surfaces those callers, since
its other fields would all be structurally zero.)
/modules/{name}/source is unchanged: it already answered 404 MODULE_NOT_FOUND for both an absent
module and a placeholder, since there is no source to serve either way.
Scope limit — this guards the root module of a request only. A call-tree that traverses into
placeholder targets still reports that subtree as empty without flagging it, so a dispatcher whose
targets are all placeholders still returns 200 with a silently truncated tree. Cross-check the
targets you care about individually (a 409 tells you it is unanalysed) — see item 103.
Data literals are not call targets (item 62). A CALLNAT <bareword> whose target is really a data
value — a browse key reaching the call site through a copycode/macro argument — used to leave a
permanent unresolved placeholder callee. Enrichment now reaps such a placeholder when it has no Natural
sigil (#/&/+), is a real VARIABLE/CONSTANT of the project, matches no real MODULE, and is
only ever reached by inferred (CALLNAT_DYNAMIC/INCLUDE_MACRO) edges. So callees, call-tree, the
ego graph and the unresolved list no longer list data names as modules. Genuine dynamic dispatch
(CALLNAT #PGM-VAR, sigil'd) is still reported as an unresolved target, and a static CALLNAT 'X' is
always trusted even when X collides with a field name.
The ingest summary agrees with the graph (item 64). A refresh/refresh/{name} response's
unresolved list is built during the file walk, independently of the graph — before item 64 it therefore
reported data fields as missing modules (MODULE CO-TABLA, MODULE NAME-DESC-SP) that enrichment had
already reaped, so the summary and the graph disagreed. It now applies the same item-62 gates as the graph
reap, with the "is this a real field?" test asked project-wide (a data literal is often declared outside
the ingested tree). Genuinely missing modules are still reported — including sigil'd dynamic dispatch
(MODULE #GETSHORT-MODUL) and a missing module whose name collides with a field name but is called
statically (the RPC-CNTX class). Treat unresolved as "dependencies that really are absent".
Constant-folded string-assembled targets (item 83). A dispatcher often builds the CALLNAT <var>
name from a base literal plus one or more SUBSTR overlays — e.g. #GETSHORT-MODUL in YGEAGGNH,
assembled by MOVE 'YGEAGKEY' TO #GETSHORT-MODUL then MOVE 'GN0' TO SUBSTR(#GETSHORT-MODUL,6,3) →
YGEAGGN0. The parser records each SUBSTR write as a WRITES carrying substrPos/substrLen
(1-based), and the resolve-dynamic-callnat-fold enrichment step folds the last full-var literal
written before the call site with the intervening overlays (left/substring) and MERGEs a resolved
CALLS edge (callKind=CALLNAT_DYNAMIC, folded=true) to the assembled module when it is a real
MODULE. It runs in every ingest mode (cheap, bounded by dynamic call sites) and over-approximates like
the other dynamic resolvers. This auto-recovers the Y…GNH → Y…GN0 family with no manual override, so
folded sites drop out of dynamic-calls/unresolved and the assembled target appears in
callees/call-tree/graph/ego graph tagged CALLNAT_DYNAMIC.
Pin what the resolvers can't: manual dynamic-CALLNAT overrides (item 82). Some CALLNAT <var>
targets are beyond the auto-resolvers — a name never present as a recoverable literal in the reachable
code (read from a work-file record, supplied by an unlinked caller, …). Such a site stays an unresolved
placeholder. A human or agent resolves it via
POST /api/projects/{p}/dynamic-calls/overrides with the call site's originFile + lineNo (from
GET .../dynamic-calls/unresolved) and the target module name(s) — multiple targets for a genuine
branch. The override is stored as a :DynamicCallOverride node outside the :AstNode graph, so a
refresh never deletes it and an enrichment step (apply-manual-dynamic-callnat, after the auto
dynamic-CALLNAT resolvers, before the placeholder cleanup) re-applies it automatically — MERGEing a
CALLS edge (callKind=CALLNAT_DYNAMIC, resolvedBy='manual') to each target and flagging the
placeholder manualHidden so callees/digest/graph/call-tree show the real target, not the
#var. It only applies while the site is still unresolved: once an auto-resolver catches up, the
override is skipped and listed obsolete — except a constant-fold (item 83), which a manual override
outranks: the fold skips a site carrying a :DynamicCallOverride, and a stale folded edge there is
dropped (delete-folded-overridden-dynamic-callnat) before apply-manual-dynamic-callnat runs, so the
pinned target replaces it. DELETE .../dynamic-calls/overrides?originFile=&lineNo=
resets one site (omit both = all), clearing the manual edges and un-hiding the placeholder inline (no
refresh). A target that is not a real MODULE is rejected 400 UNKNOWN_TARGET. Bug B fix: the
callees items now carry unresolved (mirroring what graph already exposed), so an unresolved
dynamic target is machine-distinguishable from a resolved one without inspecting sourceFile.
Dispatch guards: read guards — it is the only complete condition (item 72). A dispatch-table row's
guardField/guardValue/guardValues describe the innermost DECIDE only. Natural nests
value-DECIDEs inside each other (824 of upms's 7,729 blocks, 178 modules), and then the innermost guard
is just one conjunct: in VMULTMN4, the row for YTABLMA0.TX-TABLA reports #FIELD-NAME = 'TX-TABLA',
but the assignment also requires #SHORT-VIEW = 'TABL'. guards is the full chain — [{field, values}],
outermost first, joined by AND, each link's values joined by OR. Reading only the legacy fields
over-generalises: port that to Java and you get a branch firing where Natural never would. For an
unnested DECIDE the chain has one link and says the same as the legacy fields.
- Still incomplete for
NONE/ANYbranches (item 73): an assignment in aNONEbranch is reported under its enclosing chain alone, but its real condition is "enclosing guard AND NOT any siblingVALUE" — a negation a chain of equalities cannot express.guardsis strictly better than the legacy fields, not a total answer.
Within one guard: prefer guardValues over guardValue (item 64). dispatch-table rows carry both.
guardValue is lossy and kept only for compatibility: it comma-joins the branch's VALUE literals,
which drops blank alternatives entirely and, for a multi-alternative branch, yields a synthetic string the
guarded field never equals ("A1, A2"). guardValues is the faithful list — every alternative in source
order, blanks included — so VALUE 'GENAGREE-WOUT-SP', ' ' reports ["GENAGREE-WOUT-SP", " "], recording
that a blank guard field also routes into that branch. When reasoning about routing (or porting a
DECIDE to Java), read guardValues; guardValue will silently under-report the branch's conditions.
Text that is not code never yields a call (items 61 & 63). The CALLNAT/CALLNAT_DYNAMIC patterns
are unanchored (a CALLNAT may legally appear mid-line), so both ingest tiers first neutralise non-code
text: a full-line * comment is skipped, a trailing /* … is stripped (item 61), and a match whose
keyword falls inside a quoted string literal is rejected (item 63). A real CALLNAT 'MOD' is
unaffected — its keyword sits outside the quotes. This matters for trusting callers/callees/
call-tree: before item 63, prose such as WRITE(#MSG) 'NACH CALLNAT ISINGEAG:' or #ERR-TYPE := 'Callnat USIA008N' fabricated a CALLNAT_DYNAMIC edge to the real module of that name, so a mere log
message appeared as a genuine call — and, because a real module existed, it was not flagged
unresolved and could not be reaped by the item-62 cleanup. If you query a graph ingested before
2026-07-16, re-ingest (ac refresh) before trusting call-graph edges into modules that are also
mentioned in log/error text.
LoC / SLoC metrics (item 46)
Every file-level node (a MODULE program/class, or a DATA_STRUCTURE for a Natural .lda/.pda
data area) is stamped at ingest with two deterministic line metrics:
loc— physical lines of the file (language-independent; a trailing newline adds no phantom line).sloc— source lines of code: non-blank, non-comment lines, computed per language. Natural drops full-line*/**//*and inline/*comments; Java drops//and/* … */blocks while keeping those tokens when they appear inside string literals.
Both the cheap Tier-1 coarse scan and the deep Tier-2 parse use the same per-language counter, so a
module's loc/sloc are identical at any ingest depth — you can sum them to get exact,
reproducible project totals.
Where to read them:
GET /modules—loc/slocon each row.GET /modules/{name}/context—loc/slocon the module.GET /nodes/{id}—loc/slocin the node's raw properties.GET /loc(ac loc) — the rollup: a per-language breakdown (fileCount,loc,sloc) plus a project-wide total, optionally narrowed by?language=/?sourceFile=. Each source file is counted once even when it yields several nodes (Java inner classes, Natural inline groups).
null metrics mean the node predates item 46 — re-ingest (refresh) to backfill.
Generated vs. user-exit split (item 47)
A project can be created with a source language (required at creation; an attribute only — ingest
still classifies files by extension) and a generatedDir/userExitDir pair (directory names,
matched as path components like excludeDirs; both or neither). Generated modules already contain their
hand-written user-exit twin inline, so at ingest a module under generatedDir whose name also occurs
under userExitDir is annotated with that twin's LoC/SLoC (userExitLoc/userExitSloc). User-exit
files are not ingested as standalone modules (they would collide by name) — the walk skips
userExitDir.
Consequence for all non-LoC analysis: the generatedDir copy is the canonical, sole module for
every structural query (call graph, DB access, functions, data structures, identifiers, dataflow,
dispatch table). userExitDir exists only to compute the generated-vs-manually-written LoC split
below; it never contributes nodes/edges. So when verifying an API response against source for a Natural
module, always read the generatedDir file (e.g. generated_src/subprogram/WGEAGB0S.nat), not the
user_exit fragment.
GET /loc (ac loc) then reports, per language row and in the project total:
loc/sloc— the total (generated, which already includes the user exits).userExitLoc/userExitSloc— the sum of the annotated user-exit twins (the hand-written part).generatedExclusiveLoc/generatedExclusiveSloc— total − user-exit, clamped ≥0 per file (the purely generated part).
All three are 0 for projects without a generated/user-exit split. Create with
ac project create <name> <root> -l natural -g generated_src -u user_exit, or add the split to an
existing project via ac project update <name> -g generated_src -u user_exit.
?depth= means module hops (item 65)
On db-accesses / sql-statements (and the ?module= scope of variables/{name}/reads|writes),
depth=N means N module calls away — the same unit /modules/{name}/graph?depth= and call-tree neighbours
use. depth=1 = the modules this one directly CALLNATs, regardless of how deeply the calling
statement sits inside subroutines.
Before item 65 these endpoints bounded the traversal on raw CALLS edges. A CALLS edge starts at the
statement making the call, not at the MODULE node, so the traversal also stepped through internal
PERFORM jumps and depth measured statement nesting, not dependency distance. Concretely:
WGEAGB0S reached YGEAGBNH's tables through two module calls, but the raw path is 5 edges
(WGEAGB0S → GET-DATA → GET-MAIN-DATA → BGEAGFN0 → READ-FILE → YGEAGBNH), so db-accesses?depth=2
returned [] — reading as "no DB access" — and only depth=5 was truthful.
If you scripted a depth workaround (a deliberately large depth to compensate), drop it: depth
is now the value you'd naturally expect, and inflated values just widen the result set.
call-tree'sdepthcolumn is still raw-hop based and mixes internal subroutines into the tree: a direct dependency called from the main body showsdepth=1while one called two subroutines deep showsdepth=3. Use the ego graph (/modules/{name}/graph) when you need module-level distance. Tracked as an open roadmap item.
Framework-mediated DB access via INCLUDE macros (item 44)
Natural's generic table-access framework hides a CALLNAT inside a copycode member, invoked with a
statement-level macro:
INCLUDE YFRAMGC0 'YFRAML01.C-MOD-GET' '"CO-TABLA-CO-ELEMENTO-ALT"' '"YELEMGN0"' 'YELEMKEY' 'YELEMREC'
The CALLNAT to the generic accessor (YELEMGN0) lives in the copycode, not in the including module,
so before item 44 both callees and db-accesses were empty for such modules. The parser now
recognises the framework macro and emits a CALLS edge to the accessor named in the macro arguments
(de-quoted; e.g. '"YELEMGN0"' → YELEMGN0), tagged with edgeKind = INCLUDE_MACRO on
callers/callees. Because the edge is a normal CALLS, the accessor's own table access surfaces
transitively: GET /modules/{name}/db-accesses?depth=N reports the table with via = the accessor
module. The recognised macros and which argument names the accessor are described declaratively in
FrameworkMacros (ac-parser-natural). Scope: the targeted recogniser only — general .nsc copycode
expansion is still open.
db-accesses?depth=N is a superset of db-accesses (item 93). Besides READS/WRITES it also
returns the mode: "DECLARES" rows — a Java entity's own MAPS_TO table and a repository's
repositoryEntity table (item 32) — for every module in the closure, with via naming the declaring
module. Before item 93 the transitive query carried only the READS/WRITES branch, so asking the
same module with depth dropped its declared table and a Java caller's transitive db-accesses came
back empty although the entity it persists through maps to a real table.
Natural view aliases are resolved to the underlying table (item 95). A Natural DML statement names a
view variable (1 VDB2-VERSIS_LITERALES VIEW OF VERSVW_LITERALES), not the DDM. db-accesses reports
the table — FIND VDB2-VERSIS_LITERALES, FIND NUMBER NEXT-VIEW and STORE VDB2-VERSIS_LITERALES
in YLITEMN0 all come back as VERSVW_LITERALES, matching the SQL SELECT … FROM rows in the same
module. Before item 95 the alias itself was the reported name, which (a) split one table across several
names, (b) made the generator's boilerplate alias NEXT-VIEW a single node shared by 11 modules meaning
11 different tables, and (c) hid every VERSVW_LOGFILE write behind 11 VDB2-*-VLOG aliases. Table
names are upper-cased (Natural is case-insensitive).
…including aliases declared in a USING data area (item 98). A view is often declared not in the
module but in a LOCAL USING area, in the data-area export form (V 1VDB2-VERSIS_GENAGREE VERSVW_GENAGREE … — no VIEW OF text). Those resolve too: YGEAGBNH's
FIND (1) VDB2-VERSIS_GENAGREE reports VERSVW_GENAGREE. Resolution is scoped to each module's own
USING set, never by name — alias names are boilerplate, and NEXT-VIEW alone is declared over 100
different tables in upms. A module whose USING areas give two different tables for one alias is left
unresolved rather than guessed.
Natural UPDATE(<label>.) / DELETE(<label>.) count as writes (item 96). These act on the current
record of the labelled FIND/READ loop, and are reported as WRITES on that loop's table. This is
what makes the Y****MN0 access layer's update/delete path visible: YLITEMN0 reports WRITES VERSVW_LITERALES at the STORE and at UPDATE(HOLD-PRIME.) / DELETE(HOLD-PRIME.), where before
item 96 it reported only the STORE — reading, wrongly, as an insert-only layer. A reference that
resolves to no labelled loop (an unknown label, or the numeric source-line form) records nothing rather
than guessing a table.
call-tree/graph agree with callees about overridden dynamic calls (item 97). A manual
dynamic-call override hides the placeholder marker rather than deleting it. All read paths now filter it,
so a pinned CALLNAT <var> shows the real target and never the variable name. Everything driven by the
call-tree BFS — graph, db-accesses?depth=N, sql-statements?depth=N — inherits this.
callers on a dynamically-called module is an over-approximation, and says so. A Natural web-service
module is reached by CALLNAT #WIF, resolved by naming pattern: W-LST-N0.nat:362 alone resolves to 29
W****B*S/W****X*S targets, so WGEAGB0S lists W-LST-N0 and W-MNT-N0 as callers. The rows are
tagged edgeKind: "CALLNAT_DYNAMIC" — treat those as may-call, not does-call, and check
dynamic-calls/overrides / dynamic-calls/unresolved when the distinction matters.
XML payload / interface schema (item 45)
Natural XML wrapper subprograms build a wire payload by mapping data-area fields to XML tags via the
ADD-XML-LINE idiom (#W-TAG := '<tag>' / #W-VALUE := <field> / PERFORM ADD-XML-LINE, where the
subroutine COMPRESSes '<' #W-TAG '>' #W-VALUE). The deep parser extracts that contract as
PAYLOAD_FIELD nodes and exposes it:
-
GET /modules/{name}/payload(ac payload <module>) → an array of{tag, field, direction, lineNo, sourceFile}triples.directionisREQUESTfor an emitted (outbound) field.fieldis the unqualified payload field name (WXMLIN.P-COD-USUARIO→P-COD-USUARIO).sourceFileis the filelineNorefers to — the module's own file forsource=IDIOM, or the interface PDA's file forsource=PDA(so a caller opens the right file at the line, not the module at a stray line). -
source=IDIOM(item 45): extracted from a staticADD-XML-LINEemit sequence with literal tags. -
source=PDA(item 46b): the module is a generic, runtime-driven serializer (it calls theYFRAMN07tag-builder or has anADD-XML-LINE/ADD-XML-ACTsubroutine) with no static tag list in its source — real production wrappers likeWNAUTD0Sare this shape. The contract is then derived from the module'sPARAMETER USINGinterface PDA: each field is a payload field, the wire tag is the field name with the framework'sEXAMINE … '#' REPLACE '_'normalisation applied (#→_), directionREQUEST. Idiom fields take precedence when both exist.
The static idiom also handles derived tags (EXAMINE #W-TAG FOR '#' REPLACE '_' → #→_) and
both directions: an ADD-XML-LINE-style emit sub is REQUEST; a GET-XML-LINE-style parse sub
(with the reverse field := #W-VALUE binding) is RESPONSE.
Empty for modules that neither use the idiom nor are a flagged XML wrapper (or are only coarse-ingested).
Copycode (.cpy) expansion (item 46a)
Natural INCLUDE <member> <args> is a compile-time macro: the copycode body is spliced into the
including module (with positional &1&… substitution), so a copycode's CALLNAT/PERFORM, DB access
and dataflow live in the copycode, not the module. The deep and coarse parsers now expand statement-level
copycode includes before parsing, so those constructs surface on the including module — e.g. a READ
or CALLNAT that only exists in a .cpy shows up in the host's db-accesses/callees.
- Copycode-origin nodes/edges report the real
.cpyfile + line (so navigation lands in the copycode), and carryviaCopycode=<member>+includedAt=<host line>; host statements keep their own file + line (line numbers are remapped after the splice, never shifted). - A line number alone is not a location (item 66). Because of the above, one module's calls and field
accesses come from more than one file, and host and copycode lines are freely mixed — so always read the
line together with the file the endpoint gives it:
callers/callees/functions/{f}/callersreturnsites: [{lineNo, callSiteFileIndex, viaCopycode, includedAt, includePath}](not a barelineNosarray).callSiteFileIndexindexessourceFilesand is the file the call is written in; the entry's ownsourceFileIndexis a different thing — the file the named module/function is defined in.viaCopycode/includedAt/includePathare set only for copycode sites.variables/{name}/reads|writesreturnsourceFile= the filelineNois in (the.cpyfor a copycode access), plusviaCopycode+includedAt. Before item 66 the copycode's line was paired with the host's file:#W-OPTIONSwrites inWGEAGB0Swere reported atWGEAGB0S.nat:18/20/22, which is its generated comment banner — the writes are reallyISICINDI.cpy:18/20/22. UseincludedAtwhen you want the spot in the host module instead.includedAtis always a line in the module's own file, andincludePathshows the whole chain (item 104). NaturalINCLUDEnests, often through a positional argument (INCLUDE USIX050C 'YFRAMMC1'→INCLUDE &1&→INCLUDE YFRAMC01), and only the innermost member is named byviaCopycode.includedAtused to be theINCLUDEline in the enclosing.cpy— a file the response never named — soISI173N0 → YFRAMN04reported line 27, which is a comment inISI173N0.natand in truth line 27 ofYFRAMMC1.cpy. NowincludedAtis the host-module line (232), andincludePath: [{sourceFile, lineNo}, …]lists every hop, host first, innermost last (empty for a direct statement, one entry for a one-level include).db-accesses/workfile-accessesreturnsites: [{lineNo, sourceFile, viaCopycode, includedAt, includePath}]alongside the (kept, backward-compatible)lineNosarray — one entry per statement, each tying its line to the file it truly lives in.sql-statementsgainssourceFile+viaCopycodeon each statement (itsstartLine/endLineare lines insourceFile). Before this, a DB/work-file access written in anINCLUDEd copycode reached the API as a bare copycode-locallineNowith nothing to attribute it to — e.g. the DB2 sequence readSELECT … FROM SYSIBM-SYSDUMMY1lives inUSIX043C.cpyat lines 31/39/45/51/57, butdb-accessesfor the 9 including modules (YAPRFMN0, YUGRPMN0, …) reported those as bare line numbers that land on the host's own comment/DEFINE DATAlines. Thesitesfile context is the same fix item 66 applied tovariables/reads|writesandcallees.
- The same line number can legitimately appear twice (item 69). A host statement on line 10 and a
copycode statement on line 10 are two different statements, and both are returned — as separate entries
differing only in their file. Until item 69 the graph could not hold both: an edge was identified by
(source, target, type, lineNo)with no file, so the second one overwrote the first and a real access was missing from every answer. Treat(file, lineNo)as the identity of a site, neverlineNo. call-tree'sdepthcounts module hops (item 67).depthis how many module boundaries the shortest call path crosses, not rawCALLSedges — aCALLNATmade from two subroutines deep is still one hop. The root module's own subroutines are therefore depth 0. Measured onupms:WGEAGB0S's seven direct dependencies used to report depth 1..3 (BGEAGFN0was 3); all seven now report 1.- Results at a given
depthare larger than before. A subroutine of a module withindepthhops is now inside the bound, because it crosses no further boundary. Previouslycall-tree?depth=1could hide aDEFINE SUBROUTINEof the very module you asked about, just because it wasPERFORMed from another subroutine (raw depth 2) — that is the same bug seen from the inside. call-treealso returnstruncated.truemeans the intra-module subroutine walk stopped at its raw-hop budget, so someFUNCTIONitems may be missing — not that yourdepthwas exceeded (that is a normal, complete answer). It is conservative and can betruefor a complete result. Tune viaagenticcode.call-tree.internal-budget(default 20; the deepest internal chain observed inupmsis 9).- Since item 94 the budget cannot hide a module.
MODULErows come from the same module-hop BFS thatdb-accesses/sql-statementsuse, so a callee one hop away is always listed even when its call site sits behind a long internalPERFORMchain (before item 94 it was dropped, andcall-treethen contradicteddb-accesses). This also removed the path enumeration that madecall-tree?followWiring=truetime out on Java projects atdepth ≥ 2;followWiringis now usable at full depth.
- Results at a given
field-flow'sdepthcounts module hops (item 68). Likedb-accesses/sql-statements(item 65),variables/{name}/field-flow?depth=Nnow means "up to N module calls apart", not N rawCALLSedges. Before item 68 a consumer called from inside a subroutine sat several raw hops away and was dropped atdepth=1, so the endpoint answered "nothing downstream consumes this field" — read that answer with suspicion on any graph ingested before this change.field-flowno longer fabricates flows between same-named fields (item 77). A bare field reference is resolved against the referencing module's ownUSINGincludes. It used to be resolved project-wide: an unresolved bare field is one shared node per(name, project), and the resolver aggregated over all owning modules at once, so (a) two modules including different data areas that both declare the name left both unresolved on the shared node, and (b) a module with no matching include was redirected onto another module's field. Either way the two modules ended up on one node, andfield-flow— which pairs a producer with a consumer only when both touch the same node — reported a dataflow between modules that share nothing but a field name. Inupms: 38 + 28 placeholders affected (199 module-field pairs). Read any pre-item-77field-flowresult for a common field name with suspicion, and note the answer only changes after a deep re-ingest, since resolution runs there.reads/writesare unaffected — they match every node with the name and report only the accessing side, so they never distinguished the targets in the first place. Residue (item 76): a bare field shared via copycode (132 of 18,539 source nodes inupms) is still one node for several modules; per-module identity is a schema change, not a query fix.- Copycode provenance survives field resolution (item 70).
viaCopycode/includedAtare now kept for fields addressed by qualified name (MYLDA.Q-FIELD, i.e. a field of aLOCAL USINGdata area) as well as bare ones. Before item 70 only bare references kept it; qualified ones silently came back withviaCopycode: nulland the host file, i.e. they looked exactly like host statements. - Excluded from expansion: framework macros (handled by the item-44 targeted recogniser), data-area
USINGincludes, unknown members, and any copycode that declaresDEFINE DATA. Recursion is cycle-guarded. Copycodes (.cpy) are not standalone modules — they enter the graph only through the including module. - Staleness caveat: the item-41/43 hash check hashes the host file, so auto-invalidation triggers
on a change to the host — but a change to an included
.cpyalone (host unchanged) is not detected; re-ingest the host (refresh/{host}) to pick it up.
Global Data Areas (.gda) (item 46c)
.gda files are now ingested as DATA_STRUCTUREs like .lda/.pda, and DEFINE DATA GLOBAL USING <gda> resolves to them (the INCLUDE/USING recogniser now accepts GLOBAL, not just
PARAMETER/LOCAL).
Deep-ingest: now automatic (lazy Tier-2)
Field-level endpoints (flow-forward, flow-backward, field-flow) and
cross-module dynamic CALLNAT resolution need a per-module deep ingest, not
just a whole-root refresh. This deep ingest is now triggered
automatically on demand: calling a field-level endpoint for a module that is
only CALL_GRAPH-ingested runs a scoped deep ingest of that module (and its
dependency tree) transparently, then returns the resolved result — no 409,
no manual POST /refresh/{name} step. The first such call to a cold module is
therefore slower (it walks the root and parses the program tree); subsequent
calls hit the already-FULL graph.
The deep ingest is best-effort: if the module cannot be resolved to a source
file, the endpoint still falls back to the 409 NOT_DEEPLY_INGESTED /
NOT_INGESTED hint with a nextAction rather than a misleading empty result.
Flow path-ingest (auto, cross-module fixpoint). flow-forward,
flow-backward, and field-flow go one step further than the single start-module
deep ingest: after deep-ingesting the start module they run an ingest-and-re-traverse
fixpoint. Each round deep-ingests the frontier — the modules the trace surfaced
together with their direct callee modules — in one scope, then re-traverses. This is
what lets a dataflow trace cross into a dynamically-dispatched callee (CALLNAT PGM-VAR): that callee is not a static dependency of the start module, so it is only
pulled in and linked (caller.arg → callee.param) once a round resolves the dynamic
CALLS edge and ingests the target. The loop is bounded by
agenticcode.deep-ingest.flow-rounds (default 3) and the per-round
agenticcode.deep-ingest.fanout-nodes budget, and stops early (fixpoint) as soon as a
round pulls in nothing new — so on an already-deep graph a flow query costs one
traversal plus one cheap frontier check, no re-run.
Fan-out warm (auto, on the result set). The fan-out / traversal queries
callers, /search/identifier, and call-tree also auto-deep-ingest — but on
the set of modules their result surfaced, not a single named module. Each
runs against the graph as-is, deep-ingests the surfaced modules (blocking,
bounded by the fan-out node budget agenticcode.deep-ingest.fanout-nodes,
default 50), and — only if that warm actually deepened something — re-runs so
the response reflects newly-resolved dynamic dispatch (e.g. a call-tree grows
to include a dynamically-dispatched callee once the surfaced program is deep).
When everything is already FULL (or the warm resolves nothing) the query
returns its first result with no redundant re-run. Note callers warms the
already-surfaced callers, so it improves downstream precision but cannot
reveal a caller that was invisible at the coarse (call-graph) tier. Other
module-level endpoints (context, callees, db-accesses) work regardless of
ingest depth and do not trigger a deep ingest.
Bounded fan-out. A by-name deep ingest walks the transitive dependency tree
breadth-first, bounded by maxDepth (hops from the named module, default 5,
ceiling 20) and maxNodes (files, default 300). When a bound is hit the walk
stops early and the ingest response carries a truncation object
({reason: DEPTH|NODES|NODES_AND_DEPTH, maxDepth, maxNodes, hint}) — the modules
actually reached are marked FULL, the remainder stays as it was. Raise the
limits on an explicit module refresh to pull in more:
POST /refresh/{name}?maxDepth=&maxNodes=, or CLI ac refresh <name> --max-depth --max-nodes. Auto-triggered ingests
use the
server defaults; if a field-level query returns partial data because the target's
deep ingest truncated, re-run the explicit refresh with higher limits. (The
auto-trigger does not yet accept per-query limit overrides.)
Durable ingest status + coalescing (item 36). Each real MODULE node carries a
durable ingestStatus lifecycle — NOT_INGESTED (only its call graph is in the
graph) → INGESTING (a deep ingest is in flight) → INGESTED (deeply ingested,
ingestDepth = FULL) — separate from ingestDepth. When two calls trigger the same
module's deep ingest at once they coalesce rather than both ingesting: within one
process an in-process lock serialises them; across processes/restarts a best-effort DB
claim marks the module INGESTING and a loser waits for the winner to reach FULL
(re-claiming if the claim is released or goes stale after
agenticcode.deep-ingest.ingesting-ttl-seconds, default 1800; wait bounded by
claim-wait-seconds, default 120). A crash mid-ingest leaves the module re-triggerable
(it never reached FULL), and the stale INGESTING is reclaimed on the next call.
GET /nodes/{id} / /nodes/{id}/source expose ingestStatus/ingestStatusAt on the module node.
Warm concurrency cap (item 37). All auto deep-ingest/warm work (by-name, fan-out,
and flow-frontier) shares a global permit pool
(agenticcode.deep-ingest.max-concurrent-warms, default 2), so a burst of queries
cannot spawn unbounded parallel parses/Neo4j writes. A permit is acquired only around
the actual ingest; if none frees up within
agenticcode.deep-ingest.warm-acquire-timeout-seconds (default 10) the warm is skipped
and the query returns its Tier-1 (coarse) answer immediately rather than blocking —
so under sustained load a query may transiently return shallower data; retry once load
subsides, or force it with an explicit POST /refresh/{name}.
OpenAPI contract & CORS (items 48/50)
The server now ships an OpenAPI 3 spec (quarkus-smallrye-openapi): the machine
contract the web-UI TypeScript client is generated against. All REST endpoints
carry @APIResponse/@Schema annotations, so response bodies are typed in the
spec even though the JAX-RS methods return raw Response. Access it at:
GET /q/openapi— YAML (orAccept: application/jsonfor JSON)GET /q/swagger-ui— interactive UI (dev)
CORS is enabled (quarkus.http.cors.enabled=true) and restricted to the UI's dev
origins (http://localhost:5173, http://localhost:4173) — extend the
quarkus.http.cors.origins list per deployment; never ship a wildcard.
Endpoint quick reference
| Endpoint | Use for |
|---|---|
GET /modules?sourceFile=&moduleKind=&extends= |
List/filter modules; map a source file to its module name(s). Each row carries loc/sloc (item 46) and ingestStatus/ingestDepth (item 50) for status badges without a per-module round trip |
GET /loc?language=&sourceFile= |
Per-language LoC/SLoC rollup (fileCount/loc/sloc) + project total; each file counted once (item 46). For a generated/user_exit project also userExitLoc/userExitSloc + generatedExclusiveLoc/generatedExclusiveSloc (item 47) |
GET /modules/{name}/digest |
Tiny triage view before deciding which modules to expand |
GET /modules/{name}/context |
One-shot overview: functions, callers, callees, DB accesses, SQL/variable summaries (?include= for full lists) |
GET /modules/{name}/callers | /callees |
Direct callers/callees incl. EXTENDS/IMPLEMENTS/INJECTS/REFERENCES. callers scope: external (default) = modules that call this one (CALLNAT/inheritance), rolled up to the calling MODULE: a call made from inside a subroutine/method is attributed to its owning module (never the calling FUNCTION node), and repeated call sites from one caller collapse to a single row whose sites list every line — symmetric with how callees anchors its source side. internal = the module's own subroutines' PERFORM wiring (function-level). The default is external-only, module-typed only, and never lists the module as its own caller (no MODULE→MODULE self-loop); use scope=internal or /functions/{fn}/callers for intra-module / function-level wiring. callees is unchanged (default lists both external CALLNAT and internal PERFORM targets) |
GET /modules/{name}/functions/{function}/callers |
FUNCTION-level callers (item 52): who PERFORMs (Natural) or calls (Java cross-class) a specific subroutine/method, with call-site lineNos. Finer-grained than the module-level /callers (which is module→module). Same CallRefResponse shape. CLI ac function-callers <module> <function> |
GET /modules/{name}/call-tree?depth= |
Transitive call graph to scope a feature |
GET /dynamic-calls/unresolved | /overrides · POST/DELETE /overrides |
Manual dynamic-CALLNAT overrides (item 82). unresolved lists open CALLNAT <var> sites {module, originFile, lineNo, variable}; POST /overrides {originFile, lineNo, targets[], variable?, note?} pins a site to real module(s) (applied at once, persisted across refreshes, 400 UNKNOWN_TARGET for a non-module); DELETE /overrides?originFile=&lineNo= resets one site (omit both = all) and restores the placeholder inline; GET /overrides lists them with an obsolete flag. CLI ac dynamic-calls unresolved|overrides|set|reset |
GET /modules/{name}/graph?direction=&depth=&limit= |
Ego graph (item 49): bounded module-level call neighbourhood as nodes + edges (unlike call-tree). direction = out/in/both; limit caps nodes (BFS order) and sets truncated; unresolved targets carry unresolved=true + empty sourceFile. CLI ac ego-graph |
GET /modules/{name}/db-accesses | /sql-statements |
DB tables + mode, raw statement text (pass ?depth= for Natural). db-accesses/workfile-accesses return every row when no limit is given (item 103) — they used to default to 50, and since the response is a bare array with no total and no truncated flag the cut was invisible: WGEAGB0S?depth=10 returned 50 of 64 rows and hid 7 tables outright. An explicit limit is still honoured exactly. db-accesses items carry sites: [{lineNo, sourceFile, viaCopycode, includedAt}] (+ kept lineNos); sql-statements items carry sourceFile + viaCopycode — so a copycode-sourced access (e.g. SELECT … FROM SYSIBM-SYSDUMMY1 in USIX043C.cpy) reports the .cpy line, not a bare number that reads as a host-file line |
GET /modules/{name}/workfile-accesses |
Natural work files (sequential/flat-file I/O — READ/WRITE WORK FILE n), the work-file analogue of db-accesses (item 84): [{workFile, physicalName, mode: READS|WRITES, recordBuffers, lineNos, sites}], aggregated per work-file number + mode. sites: [{lineNo, sourceFile, viaCopycode, includedAt}] gives each access its file context (copycode-aware), like db-accesses. physicalName comes from a DEFINE WORK FILE n '<name>', else null. Kept separate from db-accesses — a work file is not an ADABAS/SQL table (fixes a former bug where READ WORK FILE created a phantom DB_TABLE 'WORK'). CLI ac workfile-accesses <module> |
GET /modules/{name}/data-structures |
Which copybooks/inline groups a module uses. A USING <member> binds by member (file) name, never by a level-1 record inside the file (item 100) — before that, WGEAGB0S USING W-WIF-A2 reported old/W-WIF-A7.pda (whose level-1 record is a copy-pasted 1W-WIF-A2), and a data area with several level-1 records and none named after the member (VLAYERLA.lda, USIX020L.lda) resolved to nothing at all (sourceFile: null, area: UNKNOWN, fieldCount: 0) although the file was ingested. One row per resolved definition, (name, sourceFile) (item 102) — never one row blending an arbitrary file with another definition's fieldCount |
GET /modules/{name}/payload |
Natural XML wire-payload contract: {tag, field, direction, source, lineNo, sourceFile} — static ADD-XML-LINE idiom (source=IDIOM, item 45) or derived from the wrapper's interface PDA (source=PDA, item 46b). sourceFile is the file lineNo refers to (module for IDIOM, PDA for PDA) |
GET /modules/{name}/dispatch-table |
Natural DECIDE ON VALUE OF routing table |
GET /modules/{name}/functions?kind= | /functions/{fn}/overrides | /functions/overrides |
Method list, modifier filter (Java), subclass overrides (single/bulk). Each item carries sourceFile + viaCopycode (item 84): a Natural subroutine pulled in via INCLUDE reports the copycode file and viaCopycode:true, so its startLine/endLine are read as offsets into that copycode — not into the including module's own file (which is shorter). viaCopycode:false = declared inline. Always false for Java |
GET /data-structures/{name}/fields | /db-tables/{name}/columns | /modules/{name}/columns |
Field/column schemas for DTO/entity generation. Every field carries sourceFile (item 101). When a structure name resolves to several definitions (42 level-1 names recur across upms data areas), the member root — the definition whose file basename equals the name, i.e. what a USING <member> binds to — wins; ?sourceFile= pins a specific one. Before item 101 the definitions were silently unioned: W-WIF-A2 returned 15 fields, the merge of W-WIF-A2.pda (5) and W-WIF-A7.pda (10), a layout that exists nowhere |
GET /variables/{name}/reads | /writes | /flow-forward | /flow-backward | /field-flow |
Impact analysis and dataflow tracing |
GET /search/identifier | /search/value | /search/annotation |
Cross-project lookup by name / literal value / annotation. search/identifier matches the exact declared name but is sigil-insensitive: a leading Natural sigil (# user, & AIV, + GDA) is ignored on both sides, so name=K-OUT-MAX finds the declared #K-OUT-MAX (and vice-versa). Optional scope filters sourceFile=<relpath> and module=<name> (item 53) narrow the match to one file / one module — use them to pinpoint a module-local declaration when a name recurs across dozens of modules (the result is otherwise paginated and the local one may fall off the page). To keep the full cross-project list yet still guarantee a given module's own declaration is on the first page, pass priorityModule=<name> instead of module=: it does not filter, but pins that module's matches to the front (ahead of the otherwise sourceFile-ordered rest) so they survive the limit. This is what the web UI's click-to-identify sends for the open module. CLI ac search-identifier --module --priority-module --source-file --type accept the same filters. Latency (item 105): a lookup whose hits lie in a Natural data area used to take 60-75 s — every fan-out query deep-ingested the surfaced .lda/.pda, which can never reach FULL (a data area yields no MODULE node), so it was re-warmed on every call and each warm dragged a whole-project finalize behind it. Data areas are now excluded from the fan-out warm; they have no deep tier to gain |
GET /search/source?regex=&limit=&ignoreCase= (ac search-source) |
Regex grep over module source text (item 54): {module, sourceFile, lineNo, line} hits + truncated. Case-insensitive by default. Complements /search/identifier (declared names) — use for code patterns (statements, table names, literals) |
GET /nodes/{id} |
Every property of one node (when a curated DTO is missing something) |
GET /nodes/{id}/source | /modules/{name}/source | /source?file= |
Source text — only when you have no other access to the source (you always do in this repo, see "Reading source in this repo" above). /modules/{name}/source returns the whole file when the line range is omitted (M1), or a [startLine,endLine] slice when both are given. /source?file=<relpath> (CLI ac file-source) serves a file by relative path rather than module name — for files that aren't standalone modules, e.g. a Natural data area (PDA/LDA) USING'd by a module, whose field line numbers refer to that file. Same whole-file/range + stale-source semantics; the client-supplied path is rejected (400 INVALID_SOURCE_FILE) if it escapes the project root |
Full endpoint list, request params, and response field details:
x-docs/agent-api-system-prompt.md.
Missing capability?
If the API/CLI genuinely cannot answer a question (not just
unreachable — the capability doesn't exist), finish the task via
grep/Explore as a fallback, then use AskUserQuestion to flag the gap and
ask whether it should become a roadmap item in x-docs/roadmap.md. Don't
silently fall back and move on.