Files
agenticCode/x-docs/features.md
Ingo Schnabel 0dfc85a11b Fixes
2026-07-19 10:34:15 +02:00

171 KiB
Raw Blame History

AgenticCode — Implemented Features

Completed work, moved out of x-docs/roadmap.md (which now tracks only open items). Each entry records what was built; IDs are preserved from the roadmap (some IDs recur across sections — they are kept as-is for traceability).

Metrics

  • 47. Generated vs. user-exit LoC/SLoC split (2026-07-13) — a project can declare a source language (required at creation; attribute only — ingest still classifies by extension) and a generatedDir/userExitDir pair (directory names, matched as path components like excludeDirs; both-or-neither, 400 LANGUAGE_REQUIRED/GENERATED_USEREXIT_PAIR otherwise). A generated module already contains its hand-written user-exit twin inline, so at ingest a module under generatedDir whose name also occurs under userExitDir is annotated with that twin's LoC/SLoC (userExitLoc/userExitSloc); user-exit files are not ingested as standalone modules (the walk skips userExitDir, avoiding name collisions). New UserExitMetrics.scan computes the twins with the same per-language LineCounter; ProjectIngestService.withUserExitMetrics stamps the generated nodes in both the coarse (Tier-1) and deep paths (correct at any depth). GET /loc (project_loc MCP, ac loc) now reports, per language row and in the total: loc/sloc (total, incl. exit), userExitLoc/userExitSloc, and generatedExclusiveLoc/generatedExclusiveSloc (= total − exit, clamped ≥0 per file). Fields added to ProjectInfo/ProjectRequest (REST + CLI project create/update with -l/-g/-u; project create has no MCP tool). Tests: SourceFilesTest, ProjectResourceIT (validation), UserExitLocIT (rollup split against the counter).

  • 46. Deterministic per-file LoC / SLoC, per language (2026-07-12) — every file-level node (MODULE, or DATA_STRUCTURE for a Natural .lda/.pda) is stamped at ingest with loc (physical lines) and sloc (source lines: non-blank, non-comment). SLOC is computed per language by a new LineCounter SPI in ac-parser-core (LocMetrics record + loc/sloc property keys): NaturalLineCounter follows natural-grammar.md §22.6.1 (full-line */**//*, inline /*); JavaLineCounter is a char-state scan handling //, /* … */ and string/char/text-block literals (so // inside a string stays code). Both the Tier-1 coarse scan and the Tier-2 deep parse call the same counter, so a module's metrics are identical at any ingest depth (numbers are reproducible and summable). Stamped in NaturalCoarseScanner/JavaCoarseScanner and, on the deep path, in ProjectIngestService.withShellMetrics (via AstIngestService.count). Surfaced on list_modules and module_context (new loc/sloc fields), on inspect_node (raw properties), and via a new rollup: GET /api/projects/{p}/loc?language=&sourceFile= (PROJECT_LOC cypher → GraphRepository.projectLoc → ProjectLoc DTO) — per-language fileCount/loc/sloc plus a project total, each source file counted once (collapses multi-node files via max per (language, sourceFile) before summing). Delivered as REST + MCP (project_loc) + CLI (ac loc) together. Tests: LocMetricsTest, NaturalLineCounterTest, JavaLineCounterTest, and AnalysisResourceIT (metrics match the counter; rollup totals/per-file consistent).

Natural ingest fidelity (WSUBPX0S findings)

Surfaced 2026-07-12 while re-analysing WSUBPX0S (project upms) via the API vs. the grep-based analyses. Two classes of information the API could not recover; both now fixed. Executable acceptance tests: ac-code-server/.../api/NaturalIncludeMacroAndXmlPayloadIT.java (the _target tests are green; fixtures under src/test/resources/fixtures/natural/framework-gaps/).

  • 44. Framework-mediated DB access via INCLUDE macros (2026-07-12) — table access performed through the generic table-access framework (INCLUDE YFRAMGC0 … '"<accessor>"' …) was invisible: the CALLNAT to the generic accessor lives in the copycode member, not the including module, so callees/db-accesses were empty. NaturalParser and NaturalCoarseScanner now recognise the statement-level framework macro (a new INCLUDE_MACRO matcher, gated on the declarative FrameworkMacros registry that maps a macro name → the positional index of the accessor argument), de-quote the accessor name ('"FRELEMG0"' → FRELEMG0, args collected across continuation lines), and emit a CALLS edge tagged callKind=INCLUDE_MACRO (new CallKind constant; surfaced as edgeKind on callers/callees). Because it's a normal CALLS, the accessor's table access surfaces transitively via the existing db-accesses?depth= query with via=accessor. Scope: targeted recogniser only (roadmap option 1); general .nsc copycode expansion and direct READS/WRITES mode tagging were deliberately left out (unnecessary for the acceptance criteria).

  • 45. XML payload / interface schema (2026-07-12) — XML wrapper subprograms map data-area fields to XML tags via the ADD-XML-LINE emit idiom (#W-TAG := '<tag>' / #W-VALUE := <field> / PERFORM ADD-XML-LINE, whose subroutine COMPRESSes '<' #W-TAG '>' #W-VALUE); that tag↔field↔direction contract was not captured. The deep parser now detects the emit subroutine and its tag/value variable pair, walks the emit sequences, and models each triple as a new PAYLOAD_FIELD node ({tag, field, direction=REQUEST}, qualifier stripped to the field name) contained by the module. Exposed as REST GET /modules/{name}/payload + MCP module_payload + CLI ac payload (PAYLOAD cypher → GraphRepository.payload → PayloadField DTO). Scope: the emit (REQUEST) idiom with literal tags; parse-side (RESPONSE) and derived-tag normalisation remain open.

  • 46a. General copycode (.cpy) expansion (2026-07-13) — statement-level INCLUDE <member> <args> is now expanded (deep + coarse) before parsing, so a copycode's CALLNAT/PERFORM, DB access and dataflow surface on the including module (previously invisible). CopycodePreprocessor splices the resolved copycode body with positional &1&… substitution and records a per-line origin table; after parsing, node/edge line numbers are remapped back to their real file positions — host statements to the host, copycode statements into the .cpy (tagged viaCopycode/includedAt) — so line-based navigation stays correct despite the splice. Copycodes are resolved by a per-ingest CopycodeLibrary (.cpy name→path, read + cached); they are not ingestible as standalone modules. Framework macros (item 44), data-area USING, unknown members and copycodes declaring DEFINE DATA are excluded; recursion is cycle-guarded. Threaded via a new CopycodeResolver into NaturalParser.parse/NaturalCoarseScanner.scan (Java unaffected). Acceptance: CopycodeExpansionIT (a host whose only DB access + call live in a .cpy).

  • 46b. XML payload derived from the interface PDA (2026-07-13) — real production XML wrappers (WNAUTD0S-style) are generic, runtime-driven serializers (YFRAMN07 tag-builder, ADD-XML-LINE/ ADD-XML-ACT) with no static tag list in source, so item-45's idiom scan finds nothing. Such a module is now flagged xmlWrapper with its PARAMETER USING interfacePda, and GET /payload derives the contract from that PDA when no static idiom exists: each field → a payload field, wire tag = field name normalised (#→_), direction REQUEST, source=PDA (vs source=IDIOM). Idiom fields take precedence. Acceptance: NaturalIncludeMacroAndXmlPayloadIT.genericWrapperPayloadIsDerivedFromInterfacePda. The static idiom additionally handles derived tags (EXAMINE #W-TAG FOR '#' REPLACE '_' → #→_) and both directions: ADD-XML-LINE-style emit sub → REQUEST, GET-XML-LINE-style parse sub (reverse field := #W-VALUE binding) → RESPONSE.

  • 46c. Ingest Global Data Areas (.gda) (2026-07-13) — .gda files now classify as Natural DATA_STRUCTUREs (like .lda/.pda) and DEFINE DATA GLOBAL USING <gda> resolves to them (the USING recogniser now accepts GLOBAL). Was: GDAs never ingested, GLOBAL USING unresolved.

Lazy / deferred three-tier ingest

Reworked ingest from eager whole-project parsing into a lazy, on-demand model. Three tiers: Tier 1 = cheap eager reference index (per file: nodes, identifiers, coarse call/DB references — no deep bodies); Tier 2 = lazy deep ingest (control flow, statement-level dataflow, precise reads/writes) triggered on demand; Tier 3 = source served from the filesystem, no longer stored on nodes. Reverse queries (callers, search_identifier, flow_backward) stay answerable because Tier 1 pre-indexes coarse references globally. (Item 43, automatic invalidation, remains open in the roadmap.)

  • 36. Tier 1 reference index + tri-state ingest status (2026-07-11) — on project create, scan every source file and create nodes (no deep edges) plus the identifier index and coarse call/DB references. Each node carries a tri-state status not-ingested / ingesting / ingested with a lock to serialise concurrent deep-ingest of the same node. Coarse references include Natural dynamic call targets (CALLNAT PGM-VAR), copycode/INCLUDE, continuation lines — computed by the lexer, not grep.

    • The Tier-1 coarse reference scan on project create: new CoarseScanner SPI (ac-parser-core); NaturalCoarseScanner is a lexer-level single pass (module /function/data-structure shells, declared-field identifier index, coarse CALLS/READS/WRITES/INCLUDES incl. PERFORM, CALLNAT '...', dynamic CALLNAT PGM-VAR, PARAMETER/LOCAL USING copybooks — attributed to the enclosing subroutine exactly as NaturalParser does, so coarse edges merge cleanly with a later deep ingest); JavaCoarseScanner reuses JavaParser but persists only the coarse projection (class/method shells + call/type edges), since a symbol-free Java scan would mis-resolve. Each shell carries a sourceHash (feeds item 41). AstIngestService.coarseScan + ProjectIngestService.scanTier1 walk the root, persist the coarse graph and run the cheap CALL_GRAPH enrichment (so cross-file callers/callees resolve immediately) — modules land CALL_GRAPH/NOT_INGESTED. Wired into POST /api/projects/{p} create (config agenticcode.tier1.scan-on-create, default true; non-fatal on an unscannable root). Covered by Tier1IndexIT, NaturalCoarseScannerTest, JavaCoarseScannerTest, SourceHashTest.
    • The durable tri-state status + lock: real MODULE nodes carry a durable ingestStatus (NOT_INGESTED/INGESTING/INGESTED, IngestStatus enum) with an ingestStatusAt stamp, kept separate from ingestDepth (enrichment tier). GraphRepository.claimIngesting/clearIngesting (CypherQueries.CLAIM_INGESTING/ CLEAR_INGESTING) implement a best-effort DB claim — Neo4j locks only at SET, so it is not a hard mutex; correctness rests on ingestModule being idempotent, with true intra-JVM exclusion still from the in-process monitor. DeepIngestCoordinator sets INGESTING at the start of a by-name deep ingest (markIngestDepth(FULL) flips the tree to INGESTED; failure rolls back to NOT_INGESTED), and a cross-process loser coalesces by polling moduleIngestState until the winner reaches FULL, re-claiming if the claim is released or goes stale (agenticcode.deep-ingest.ingesting-ttl-seconds, default 1800; poll bound claim-wait-seconds, default 120). Covered by DeepIngestStatusIT and CoalescingIT (concurrent deep ingests converge to a single FULL node).
  • 37. Tier 2 lazy deep-ingest on every call (2026-07-11) — every MCP/API call deep-ingests the nodes it touches, transitioning them to ingested. Decision: this includes fan-out result sets (e.g. search_identifier, callers), not just the explicitly named node — so the node budget (item 39) is the primary cost bound and its default must be a real, tuned number. Traversal queries (call_tree, flow_*) deep-ingest along the path as they walk it.

    • Module-gated field-level queries auto-trigger: DeepIngestCoordinator.ensureDeep runs a scoped ingestModule (with a per-(project,module) coalescing lock) from the withDeepModule choke point in both McpQueryTools and AnalysisResource; flow-forward/backward, field-flow etc. auto-deep-ingest instead of returning 409, falling back to the hint only when the module does not resolve.
    • Fan-out result-set warm for callers, search_identifier, and call_tree: blocking-then-rerun — each query runs against the graph as-is, DeepIngestCoordinator.ensureDeepMany deep-ingests the surfaced result-set files (by path, via the multi-root ProjectIngestService.ingestFiles) under a fan-out node budget (agenticcode.deep-ingest.fanout-nodes, default 50), and the query re-runs only if the warm deepened something. Wired at the withFanoutWarm choke point in both surfaces. Note: callers warms already-surfaced callers (downstream precision) but cannot reveal a caller invisible at the coarse tier.
    • Cross-module flow_* path-ingest (37a) for flow-forward, flow-backward, and field-flow: an ingest-and-re-traverse fixpoint at the withFlowPathWarm choke point — after ensureDeep(startModule) it runs the trace and, each round, deep-ingests the frontier (surfaced modules seeded with the start module, plus their direct callee modules) in one scope (DeepIngestCoordinator.ensureFlowFrontier → GraphRepository.flowFrontierSourceFiles → ProjectIngestService.ingestFiles), then re-traverses. Ingesting the whole frontier (callers included, even if FULL) re-links each caller into the freshly-ingested callee params, letting a trace cross into a dynamically-dispatched callee. Bounded by agenticcode.deep-ingest.flow-rounds (default 3) + the fanout-nodes budget; stops early at the fixpoint. Covered by AnalysisResourceIT.flowForwardPathWarmCrossesIntoDynamicallyDispatchedCallee.
    • Global concurrency cap: all deep-ingest/warm work in DeepIngestCoordinator (by-name ensureDeep, fan-out ensureDeepMany, flow ensureFlowFrontier) is gated by a fair Semaphore (agenticcode.deep-ingest.max-concurrent-warms, default 2), acquired via tryAcquire only around the actual ingest (never while a coalesce-waiter waits); on timeout (warm-acquire-timeout-seconds, default 10) the warm is skipped and the query returns its Tier-1 answer.
  • 38. Deep-ingest depth cap (2026-07-10) — the by-name deep-ingest BFS (ProjectIngestService.ingestModule) is bounded by maxDepth hops from the named module: default 5 (agenticcode.deep-ingest.default-depth), clamped to the ceiling 20 (agenticcode.deep-ingest.max-depth). Hitting the bound truncates + reports (not a hard error) — the reached modules are marked FULL and the response carries an IngestSummary.Truncation. Override via ?maxDepth= (REST, now on refresh/{name}), the refresh MCP tool's maxDepth arg, and ac refresh --max-depth (CLI).

  • 39. Deep-ingest node budget (2026-07-10) — alongside the depth cap, the BFS ingests at most maxNodes files (default 300, agenticcode.deep-ingest.default-nodes); overflow truncates + reports the same way (reason NODES/NODES_AND_DEPTH). Override via ?maxNodes= (REST), the refresh MCP tool's maxNodes arg, and ac refresh --max-nodes (CLI). Defaults are @ConfigProperty and should still be tuned against the real ac project.

  • 40. Unresolved-reference nodes (2026-07-12) — when Tier 1 sees a variable / dynamic call target it cannot resolve, store it as a deduped placeholder node/edge flagged unresolved, so reverse queries (e.g. callers) surface dynamic call sites while still distinguishing them from resolved edges. Resolution pass runs during Tier 2.

    • Placeholder nodes (sourceFile = "") were already deduped (MERGE on type,name,sourceFile,project); this adds the unresolved flag. New enrichment step STAMP_UNRESOLVED_PLACEHOLDERS (runs last in finalizeProject/finalizeProjectScoped, project-wide, cheap, idempotent, every mode) sets unresolved = (no real definition of that (type,name) is ingested) — the resolution pass: a reference reads unresolved=false once its target file is ingested (a real twin exists), true while genuinely dangling. Surfaced explicitly on search_identifier (IdentifierMatch.unresolved) and, for free, on inspect_node (nodes/{id} returns the whole node). Reverse call queries (callers/callees) already surface these targets as entries with a blank sourceFile. Covered by UnresolvedRefIT (dangling → true; ingest target → false).
  • 41. Tier 3: stop storing source on nodes (2026-07-12) — drop stored source text; serve module_source/node_source by opening sourceFile (project-relative path + content hash) and slicing startLine..endLine. Stale check: on read, compare the stored content hash against the current file; on mismatch, return a structured STALE_SOURCE error ("run refresh") instead of slicing — never serve current lines against old line numbers.

    • Source text was already not stored on nodes (served off disk by SourceSnippetService since item 28); this adds the mandatory stale check. Shared SourceHash (SHA-256) is stamped on each file's MODULE/DATA_STRUCTURE shell at ingest — by the Tier-1 scanners and (for robustness) the deep full-parse paths (ProjectIngestService.withSourceHash), surviving re-ingest via SET +=. On a source read, GraphRepository.sourceHash(project, sourceFile) supplies the stored hash and SourceSnippetService.read recomputes the current file's hash; on mismatch it throws StaleSourceException, which both REST (409 STALE_SOURCE) and MCP (STALE_SOURCE tool error) return with a "re-ingest (refresh)" message. A legacy file with no stored hash skips the check (serves unchecked). Covered by StaleSourceIT (fresh→200, edit→409, re-ingest→200). Copycode/INCLUDE slices remain raw pre-expansion file text. Note: a fan-out warm query (e.g. search_identifier) re-ingests and re-hashes a surfaced module, clearing its staleness as a side effect.
  • 42. refresh tool (manual invalidation) + remove ingest_all / ingest_module (2026-07-12) — remove the eager bulk/one-module ingest tools across all three surfaces (MCP McpIngestTools, REST AnalysisResource, ac CLI) and replace with a refresh operation that re-runs Tier 1 + deep ingest. Interim manual invalidation until item 43.

    • refresh is the single (re-)ingest surface across REST + MCP + CLI; ingest_all/ingest_module/ingest-call-graph were removed from all three. REST POST /refresh (whole root, call-graph tier; ?deep=true for field-level) and POST /refresh/{name}?maxDepth=&maxNodes= (deep re-ingest a module tree); MCP one refresh(project, module?, deep?, maxDepth?, maxNodes?) tool; CLI ac refresh [name] [--deep] [--max-depth --max-nodes]. ProjectIngestService.refreshProject/refreshModule delegate to the existing (now internal) ingest primitives — a whole-root refresh runs the full-parse call-graph pass (a superset of the create-time coarse Tier-1 scan, so intra-module dynamic CALLNAT still resolves), and is the remedy for a STALE_SOURCE read (re-stamps sourceHash). It MERGEs current files (does not wipe; deletions await item 43). All ~20 ITs bootstrapping via the old endpoints were migrated (behavior- preserving URL swaps); CLAUDE.md + mcp-api-usage dogfooding steps now say refresh. Covered by RefreshIT (whole-project + module refresh work; removed endpoints 404).
  • 43. Automatic invalidation (hash-based) + deleted-file sweep (2026-07-15) — the graph now self-heals when source files change or are deleted on disk, without a manual refresh. Chosen approach: hash-based, lazy (on-access) — not a filesystem watcher (no background thread; reuses the item-41 sourceHash). Two parts, gated by agenticcode.auto-invalidate.enabled (default true).

    • Part A — change detection → auto re-ingest. DeepIngestCoordinator.ensureDeep/ensureDeepMany (the two deep-ingest choke points every field-level query routes through) no longer short-circuit a FULL module blindly: they compare the file's current SourceHash to the stored one (same compare as the /source stale check). A FULL-but-changed module is invalidated (GraphRepository.invalidateModule → shell reset to CALL_GRAPH/NOT_INGESTED, so isFull() is false and the coalesce loop re-ingests) and re-parsed — with the item-58 reconciliation purging its stale nodes. Conservative false (never thrash a FULL module) when the flag is off, no source file, no stored hash (legacy node), or the file is unreadable/missing (a deleted file is Part B's job). Per-module granularity: only the queried module's own file is checked, not its whole dependency tree (a changed dependency is caught when it is queried by name).
    • Part B — deleted-file sweep. A whole-project refresh (ProjectIngestService.refreshProject) now reconciles against the filesystem: GraphRepository.distinctSourceFiles → any real sourceFile no longer present on disk → deleteNodesForSourceFiles (DETACH DELETE). Closes the item-58 gap (a refresh MERGEd current files but left nodes for deleted files). Reconciles against disk existence, not the walked-candidate list, so on-demand dependency files (PDAs/LDAs/.cpy) that exist but aren't top-level candidates are never wrongly swept.
    • No new REST/MCP/CLI surface — Part A is internal to the query path; Part B changes refresh semantics (now also drops deleted-file orphans). Both are documented behaviour changes, so the three surfaces stay in sync by construction. Covered by AutoInvalidationIT (edit a FULL module on disk → a field query auto-re-ingests, 409 STALE_SOURCE → 200, no manual refresh) and DeletedFileSweepIT (delete a file → refresh removes its module + field nodes, survivors kept). PlaceholderResolveNullLineIT (which seeds phantom nodes for files it never writes) opts out via a @TestProfile disabling the sweep.

Layered / performance-aware ingest

  • P2-c. ingest-all defers field resolution AND dataflow by default (fast whole-root) (2026-06-21) — even with P2-b's scoped per-program resolution, a whole-root ingest-all still ran the global field-placeholder resolution AND the argument→parameter dataflow — measured live on upms: the field WRITES pass took ~26 min and link-args-to-params ~25 min (each one transaction over ~100k+ edges). So /ingest-all now defaults to the fast call-graph pass (EnrichmentLevel.CALL_GRAPH): placeholder resolution (CALLS/INCLUDES/USES_TYPE/EXTENDS/IMPLEMENTS) + intra-module dynamic CALLNAT + polymorphic (CHA) fan-out. It defers the expensive field-placeholder resolution and the dataflow + cross-module dynamic CALLNAT dispatch to a scoped per-program deep ingest (/ingest/{name}, P2-b) — which fans out top-down from the analyzed program (the user's insight: dynamic dispatch is a top-down concern, not needed globally bottom-up). The old full whole-root behaviour is available via POST /ingest-all?deep=true. Implementation: EnrichmentLevel { CALL_GRAPH, FULL } (each carrying dataflow/resolveFields gates), finalizeProject(project, EnrichmentLevel), ingestAll(project, deep), AnalysisResource.ingestAll gains ?deep. The default tags modules CALL_GRAPH, so field-level endpoints return the 409 NOT_DEEPLY_INGESTED deep-ingest hint (P2-a). Covered by AnalysisResourceIT (ingestAllDefaultDefersDataflowAndFieldResolutionToDeepIngest: fast ingest → intra-dynamic resolved, cross-dynamic absent + 409 on field-flow → /ingest/{name} → cross-dynamic resolved

    • 200); the IT's ingestAll() helper uses ?deep=true so existing field/dataflow assertions still exercise the full path. An earlier attempt kept dataflow+cross-callnat in a CALL_GRAPH_DATAFLOW middle tier; live upms testing showed link-args-to-params alone was ~25 min, so it was dropped.
  • P2-b. Scope deep field resolution to the program tree (2026-06-20) — in the accumulating one-project model (call-graph everything, then deep-ingest programs over time), a by-name deep ingest previously ran the unscoped field-placeholder resolution over the whole project (~110k edges → ~28 min once the full call graph is loaded). Added $names-scoped variants of the field-resolution and dataflow queries (resolvePlaceholderFieldTargetsScoped, …ByNameScoped, resolveBareIncludedFieldTargetsScoped, LINK_ARGS_TO_PARAMS_SCOPED, …_JAVA_SCOPED) that start from the just-ingested program tree's modules instead of scanning all placeholder edges. GraphRepository.finalizeProjectScoped(project, moduleNames) runs them (call-graph resolution, placeholder cleanup, CHA fan-out stay project-wide/idempotent); ingestModule now calls it with the BFS-collected module names. So deep ingest stays fast (≈ the small-project 8.5 s case) even when the project already holds the whole call graph. Correctness covered by IngestModuleIT field-flow/dataflow tests (now exercising the scoped path); all ingest/analysis ITs green. Live perf re-validated on upms (2026-06-21): whole-codebase ingest-call-graph = 6309 modules in 67 s (placeholders left intact), then a by-name deep ingest of KDWWIFN0 (56-module tree) with the full call graph already loaded ran in 5.2 s — confirming the scoped resolution avoids the old unscoped whole-graph path (~28 min). Depth guard verified end-to-end: flow-forward returns 200 for the FULL module, 409 NOT_DEEPLY_INGESTED for a call-graph-only module (ACCNPE01), and 409 NOT_INGESTED for an un-ingested name (JE999), each with an actionable nextAction.

  • P2-a. Layered ingest: call-graph mode + ingest-depth guard (2026-06-20) — field-placeholder resolution is volume-bound (whole-codebase ≈110k edges → not viable even cycle-safe/APOC node-global; measured), while per-program deep ingest is fast (~8.5 s for KDWWIFN0's 56-file tree). So ingest is now layered: (1) POST /api/projects/{p}/ingest-call-graph — whole-root, persists everything but runs only the cheap call-graph/module-level enrichment (CALLS/INCLUDES/USES_TYPE/EXTENDS/IMPLEMENTS + CHA fan-out), skips field-placeholder resolution, and leaves placeholders intact for later deep ingest. Tags modules ingestDepth=CALL_GRAPH. (2) POST /ingest/{name} (by-name) and ingest-all run full enrichment and tag the ingested modules ingestDepth=FULL (accumulates over time in one project). (3) Field-level endpoints (flow-forward/flow-backward/field-flow) check the target module via GraphRepository.moduleIngestState and, when not FULL, return an agent-actionable 409 { status: NOT_DEEPLY_INGESTED|NOT_INGESTED, module, detail, nextAction:{method,path} } instead of a misleading empty result. New IngestDepth/ModuleIngestState, finalizeProject(full) splitting enrichmentSteps(full) into call-graph vs field groups, depth-marking Cypher, per-step finalize logging (label/duration/rows/JVM heap). New IT callGraphIngestEnablesCallGraphAndGuardsFieldFlow (409→200 after deep ingest); full AnalysisResourceIT+ingest ITs green.

Parser / model

  • 1. Persisted per-language source/module kind (done 2026-07-07, scoped down after investigation) — add an AstNode property capturing each module's finer kind beyond the coarse NodeType. Delivered: Java moduleKind = CLASS/INTERFACE (exact, from type.isInterface()); Natural moduleKind = a best-effort PROGRAM/SUBPROGRAM guess (has a top-level PARAMETER section → SUBPROGRAM). GET /modules gained a ?moduleKind= filter. Not attempted (found to be a separate, larger feature, or genuinely unrecoverable from source text alone): Java enum/record — JavaParser doesn't parse these as MODULE nodes at all today (only ClassOrInterfaceDeclaration), so this needs a new node-family traversal, not just a property; Natural COPYCODE/MAP/GDA — per natural-grammar.md §22.2, PROGRAM/SUBPROGRAM/SUBROUTINE are real SAG Natural catalog metadata, and main-program/external-subprogram/ subprogram are syntactically identical productions — no reliable source-level signal exists for these without the original catalog. See x-docs/mcp-api-usage-ac-implementation.md step 0 for the heuristic's limits.

Foundation (schema, health, errors, ingest CLI)

  • 1. Neo4j indexes & constraints — index (:AstNode) (project, name) and (project, type, name), unique constraint on (:Project) (name). Created on application startup via SchemaInitializer (idempotent IF NOT EXISTS).
  • 2. Health checks — Neo4jHealthCheck (@Readiness) verifies Neo4j connectivity via Driver.verifyConnectivity().
  • 3. Structured error responses for AnalysisResource — shared ErrorResponse record ({error, code, details}), reused by ProjectResource.
  • 4. Project-existence validation — ingest and all project-scoped query endpoints return 404 PROJECT_NOT_FOUND if the project hasn't been created.
  • P0-a. Uniqueness constraint/index on AstNode.id — mergeEdge matches both endpoints by id (MATCH (a:AstNode {id: $sourceId}), (b:AstNode {id: $targetId})), but id was not indexed (only (project, name) and (project, type, name) were). Every edge merge did a full label scan over all AstNode nodes accumulated so far, so ingest got progressively slower as more files/projects were ingested. Added CREATE CONSTRAINT ast_node_id_unique IF NOT EXISTS FOR (n:AstNode) REQUIRE n.id IS UNIQUE to SCHEMA_STATEMENTS, giving edge merges O(1) lookups regardless of graph size.
  • P0-b. Batch ingest writes with UNWIND — GraphRepository.save() previously issued one tx.run() per node and per edge (N+M round trips per ingested file). Replaced with CypherQueries.MERGE_NODES (UNWIND $nodes AS n MERGE ...), one round trip for all nodes of a file, and CypherQueries.mergeEdgesBatch(EdgeType) (UNWIND $edges AS e MATCH ... MERGE (a)-[:<TYPE>]->(b)), one round trip per distinct EdgeType present in the file (edges grouped via Collectors.groupingBy(AstEdge::type)) — Cypher relationship types can't be parameterized, so true O(1) isn't possible, but this cuts round trips from N+M to 1 + (number of distinct edge types, usually 1-3). Same merge semantics, low risk.
  • P0-c. --exclude-dir option for ingest CLI command — recursive ingest now skips any file whose path (relative to the ingest root) contains a component matching one of the repeatable -x/--exclude-dir <name> values (case-insensitive), e.g. --exclude-dir user_exit --exclude-dir test to skip subprograms already represented in generated_sources. Excluded files are printed as SKIP (excluded) <file> and counted in a new excluded summary counter alongside ingested/skipped/failed.
  • P0-d. ingest-module <folder> <moduleName> — dependency-driven ingest (2026-06-15) — new CLI command that ingests a single module and its transitive CALLNAT/PERFORM-external/EXTENDS/IMPLEMENTS targets and INCLUDE/USING'd data areas, located by filename stem (case-insensitive) within a given folder. POST .../ingest/{java,natural} now returns 202 {"dependencies": [{"name":..., "type": "MODULE"|"DATA_STRUCTURE"}]} (new AstIngestService/AnalysisResource.IngestResponse/DependencyRef), derived from placeholder nodes (sourceFile="") created by the parsers. The CLI does a BFS over these names: found files are ingested and their dependencies enqueued; names not found in the folder are printed as UNRESOLVED <name> and counted separately (not a failure). --exclude-dir applies to the file index, so dependencies under excluded folders are reported as unresolved. Shared file-walk helpers (readSourceFile, endpointFor, isExcluded) extracted from IngestCommand into IngestSupport. New IT naturalIngestReturnsUnresolvedDependencies.
  • P0-e. .lda/.pda (Natural data area) parsing (2026-06-15) — NaturalParser now recognizes .lda/.pda files (via IngestSupport.endpointFor) and parses their wrapper-less field-export format (<TYPE><LENGTH><LEVELDIGIT><NAME> ... CONST<...>|INIT<...>, glued type/level/name prefixes, R <level><name> REDEFINE markers, *-comment lines, multiple top-level (level 1) DATA_STRUCTURE roots per file) into the same DATA_STRUCTURE/VARIABLE/CONSTANT node shape as DEFINE DATA fields, so data-structure-fields works unchanged. Combined with item 6 (placeholder resolution), an ingested LDA/PDA now links up with the DATA_STRUCTURE placeholder created by LOCAL/PARAMETER USING in referencing .nat programs. New unit tests (YFRAML01_SAMPLE.lda, WGEAGL01_SAMPLE.pda) and IT (ingestingLdaResolvesIncludePlaceholder).
  • P0-f. GET /variables/{name}/writes and /reads (2026-06-15) — queries returning every FUNCTION/MODULE that has a WRITES/READS edge to a VARIABLE/CONSTANT with the given name, with sourceFile and containing module (CypherQueries.variableReads(maxDepth)/ variableWrites(maxDepth), GraphRepository.variableWrites/variableReads, VariableAccessLocation record). Optional ?module=...&depth=N query params restrict results to that module or anything in its transitive CALLS tree (up to depth hops, same agenticcode.call-tree.default-depth/max-depth config and clamping as call-tree); the Cypher computes each writer/reader's containing module via CONTAINS*0.. and checks owner.name = $module OR EXISTS { (module)-[:CALLS*1..depth]->(owner) }. New CLI variable-writes/variable-reads <name> [-m/--module <name>] [--depth N].

Ingest performance

  • 24. Index the stale-file sweep — persist was a label scan per file (2026-07-16) — the sibling defect to P1-m below, one statement further on: P1-m indexed the node MERGE key, but DELETE_STALE_FILE_NODES (the item-58 reconcile sweep, which runs in the same mergeResults transaction) matches on (project, sourceFile) — and a Neo4j composite index only applies when every one of its properties is constrained, so the 4-property (project, sourceFile, type, name) index does not cover that 2-property prefix. The planner therefore fell back to NodeByLabelScan + Filter, and because the UNWIND $files batch drives it through an Apply, it scanned once per file: measured 564,451 db hits per file on a 282k-node graph, i.e. ~56M node reads per 200-file batch. The scan covers the whole AstNode label — all projects — which is why per-batch cost grew with total graph size (2m29s early vs 2m48s late). Fix is one line in SCHEMA_STATEMENTS: ast_node_project_sourcefile on (project, sourceFile). The plan becomes a NodeIndexSeek at 137 db hits per file (~4000×). Verified that the new, less selective index does not displace the 4-property one for the MERGE keys (still 1–2 db hits). Also serves SOURCE_HASH (item 43 auto-invalidation) and the sourceFile:"" placeholder sweeps, which had the same blocked shape. Index-only; ensureSchema adds it IF NOT EXISTS, so existing deployments pick it up on the next restart — no migration. Measured live on a full upms call-graph refresh (6311 files, reconcile on): whole refresh 179 s end-to-end — parse ~63 s, persist 104.5 s across 32 batches (median 1.7 s/batch, min 0.2 s, max 10.7 s), finalize ~12 s — against ~2.5–2.8 min per batch logged before the fix, and with no per-batch growth left. A second refresh took 95 s and reproduced the graph exactly (454,255 nodes / 1,495,146 rels, no duplicate MERGE keys), i.e. idempotent. Note the pre-fix graph held only 216,492 upms nodes: the slow run evidently never completed, so the fix also produced the first complete upms graph rather than merely a faster one. Regression test StaleFileSweepIndexIT asserts the cause — it EXPLAINs the real DELETE_STALE_FILE_NODES constant and requires a NodeIndexSeek on AstNode(project, sourceFile) (the planner names indexes by properties, not by index name), after CALL db.awaitIndexes() since a POPULATING index is invisible to the planner and would make the assertion race startup. Vacuity-checked: with the schema line removed the plan falls back to the filtered label scan and the test fails.
  • 25. ingest-all performance: batch persist transactions — already implemented (closed 2026-07-16) — the roadmap item claimed "persist still opens one transaction/session per file"; that was stale. ProjectIngestService chunks files into persistBatch calls sized by agenticcode.ingest.batch-size (default 200), and GraphRepository.mergeResults aggregates each chunk into batched UNWIND MERGEs (one per node class, one per edge type) in a single transaction — delivered earlier as P1-k below. Closed as done; it was item 24 above, not batching, that actually held persist back.
  • P1-m. Index the node MERGE key to kill O(N²) persist (2026-06-20) — on a whole-root ingest of upms (~6309 files), per-batch persist time grew with the graph (batch 1 ≈ 1 s, batch 8 ≈ 4.5 min) — quadratic. MERGE_NODES/ MERGE_POSITIONAL_NODES match on (type, name, sourceFile, project) but the only supporting index was (project, type, name); for low-selectivity names that recur project-wide (CONTROL_FLOW IF/FOR/REPEAT/DECIDE, VARIABLE/DATA_STRUCTURE FILLER/#CODE/SQLERR) that bucket grows with the graph and MERGE filtered sourceFile/startLine in memory over it. Added composite index ast_node_project_sourcefile_type_name on (project, sourceFile, type, name) so MERGE seeks per-file (a handful of nodes). Index-only; ensureSchema adds it IF NOT EXISTS.
  • P1-l. Split enrichment into per-statement transactions (2026-06-20) — finalizeProject ran all ~15 project-wide enrichment statements in one transaction; on the full upms graph that single transaction blew Neo4j's dbms.memory.transaction.total.max (5.4 GiB) and aborted, so placeholders never resolved. finalizeProject now runs each ordered, idempotent statement (enrichmentStatements()) in its own executeWriteWithoutResult, bounding peak per-transaction memory. Same graph result (ITs green); a lone over-large step would still need CALL { … } IN TRANSACTIONS (not yet required).
  • P1-k. Batch persistence in ingest all (2026-06-19) — ingestAll persisted one file per transaction (fresh session + ~8 statements + commit), sequentially — fine for a small slice but slow once the P1-j dedup fix made it ingest thousands of files (each round trip is fixed overhead × N files). New GraphRepository.persistBatch / AstIngestService.persistBatch persist a chunk of files in one transaction, aggregating all files' nodes/edges into batched UNWIND writes (one statement per node label / edge type per chunk instead of per file). ingestAll now chunks the non-conflicting results by agenticcode.ingest.batch-size (default 200) and runs one finalizeProject after. Correctness subtlety: different files emit the same placeholder (CALLNAT/USING target, sourceFile="") with distinct UUIDs but one MERGE key; persisting all nodes before all edges collapses them (last id wins) and would orphan earlier files' edges. mergeResults dedupes nodes by MERGE key (canonical = first id) and remaps every edge endpoint onto the canonical id — the per-file transactions previously masked this. Caught by AnalysisResourceIT (interface fan-out + shared-PDA field-flow) during implementation; full ingest/analysis ITs green.
  • P1-j. Fix over-aggressive duplicate detection in ingest all (2026-06-19) — ProjectIngestService.ingestAll flagged cross-file duplicates by scanning every real MODULE/DATA_STRUCTURE node, so two unrelated Natural programs that each contain an inline DEFINE DATA group with a common name (FILLER, SQLERR, #ERROR-GROUP, …) were treated as duplicates and both whole files skipped. On the upms codebase this silently dropped 4379 files (3894 of 3895 "duplicate" identities were inline DATA_STRUCTURE names; only 1 was a real MODULE), so e.g. W-MNT-N0/KDWWIFN0 never ingested and their callers/callees couldn't resolve. New fileIdentities(Parsed) keys duplicate detection on the file's own identity only: real MODULE nodes (Java FQN-aware) for program/class files, and a single (DATA_STRUCTURE, stem) for .lda/.pda data-area files — inline group nodes are excluded. Real MODULE and data-area-file duplicates are still reported. In-memory only (no schema change/re-ingest of node shape). New IT sharedInlineDataStructureNamesDoNotSkipModules; full AnalysisResourceIT + IngestModuleIT green. A persisted per-language source/module kind is tracked separately as future work.
  • P1-i. Defer enrichment in by-name ingest (2026-06-19) — POST /ingest/{name} (ProjectIngestService.ingestModule, the CLI ingest <module> path) called save() per file in its dependency BFS, which re-ran the full ~15-statement project-wide enrichment (placeholder resolution + dataflow + polymorphic fan-out) after every file — ≈O(N²) over a growing graph, the dominant cost for Natural modules that fan out through CALLNAT/PERFORM/ USING (observed on ingest KDWWIFN0). Now mirrors ingestAll: persist() per file (merge only), then one finalizeProject() after the BFS. Enrichment is idempotent/order-independent, so the end graph is identical (verified by IngestModuleIT, which exercises cross-module dataflow that depends on enrichment). Removed the now-unused AstIngestService.save() / GraphRepository.save().
  • P1-h. Run DB endpoints off the event loop (2026-06-19) — the Neo4j blocking driver is wrapped in Uni.createFrom().item(...) in every GraphRepository method, whose supplier runs on the subscribing thread. The Uni-returning JAX-RS query methods carried no @Blocking, so for those the driver ran on (and blocked) the Vert.x event loop — latent until a slow query tripped the 2 s BlockedThreadChecker (observed on DELETE /api/projects = CLEAR_ALL full-graph DETACH DELETE, blocked ~3.3 s). Fixed by moving @Blocking to class level on AnalysisResource and ProjectResource so all endpoints dispatch to a worker thread (the two now-redundant method-level @Blocking on the ingest endpoints were removed). Verified via access log (executor-thread-* instead of vert.x-eventloop-thread-*); ProjectResourceIT + AnalysisResourceIT green.
  • 23. ingest-all performance: enrich once, not per file (2026-06-19) — ingest-all ran the project-wide enrichment block (placeholder resolution for all RESOLVABLE_EDGE_TYPES, field-target resolution, bare-field redirection, placeholder deletes, LINK_ARGS_TO_PARAMS/_JAVA, LINK_CALLS_TO_IMPLEMENTATIONS) inside every GraphRepository.save, i.e. once per file — each pass scans the whole project graph, so cost grew ~O(N²) in the file count and dominated large Java scans. Split persistence from enrichment: GraphRepository refactored into mergeResult/runEnrichment helpers, exposing persist(project, result) (MERGE only) and finalizeProject(project) (enrichment once), with save retained as merge + enrich in one tx for single-file/by-name ingest. ProjectIngestService.ingestAll now persists each survivor then calls finalizeProject once (skipped when nothing persisted). The enrichment is idempotent and order-independent, so the final graph is identical; verified by the existing AnalysisResourceIT (ingest-all-based) suite. ingestModule unchanged (still save per file).

Call graph, edges & provenance

  • P1-g. Language-aware edgeKind (2026-06-19) — edgeKind in callers/callees was always derived from the target node type as Natural's CALLNAT/PERFORM, mislabelling Java calls (e.g. a new Foo() constructor showed as CALLNAT). Each parser now stamps a callKind property on the CALLS edge using the new CallKind enum (ac-parser-core): Natural → CALLNAT/PERFORM; Java → METHOD_CALL (intra- and cross-class invocations) / CONSTRUCTOR (new). Stamping at parse time lets Java distinguish constructors from method calls, which the previous target-type-only logic could not. CypherQueries.callers/callees now return coalesce(r.callKind, <legacy CASE>) so graphs ingested before the stamp fall back to the old labels (no forced re-ingest; re-ingest needed for Java precision). Synthetic inheritance edges (LINK_CALLS_TO_IMPLEMENTATIONS) copy s.callKind = r.callKind. New IT assertions for Java (METHOD_CALL/CONSTRUCTOR) and Natural (PERFORM/CALLNAT).
  • 6. Cross-file placeholder resolution (2026-06-15) — decided against a generic EnrichmentPipeline.run() (no second use case identified yet; enrichment/ stays empty). Instead, GraphRepository.save() runs CypherQueries.resolvePlaceholderTargets(EdgeType) for each of CALLS/INCLUDES/USES_TYPE/EXTENDS/IMPLEMENTS, then DELETE_RESOLVED_PLACEHOLDERS, in the same transaction as the node/edge merges. This redirects edges pointing at an unresolved placeholder (MODULE/DATA_STRUCTURE, sourceFile="") onto a real node sharing (type, name, project) once one exists, and deletes the now-orphaned placeholder — works regardless of ingest order.

Natural parser — statements, dataflow, dispatch

  • P1-a. SQL/ADABAS statement extraction for agentic Panache generation — new NodeType.DB_ACCESS node per SQL/ADABAS statement occurrence (table, access mode, raw statement text, line span), CONTAINS edge from the containing function/module, USES_TYPE edge to the accessed DB_TABLE and (for SELECT ... INTO VIEW) the associated DATA_STRUCTURE. New GET /api/projects/{project}/modules/{name}/sql-statements endpoint + CLI command, so an agent can read table/view/mode/statement text and the call-graph context and write the equivalent Panache query itself. Existing function -> DB_TABLE READS/WRITES edges and /db-accesses stay unchanged.
  • P1-b. MOVE/ASSIGN/COMPUTE variable read/write edges — Natural construct #7 (CLAUDE.md priority list). Declared VARIABLE/CONSTANT names from DEFINE DATA are tracked case-insensitively; for MOVE <src> TO <target>{...} and (COMPUTE|ASSIGN) <target> (=|:=) <expr>, READS/WRITES edges are added from the current function to known variables/constants referenced as source/target/expression operands. Exotic MOVE variants (BY NAME/BY POSITION, EDITED, SUBSTRING(...), ALL, ENCODED, NORMALIZED, *JUSTIFIED) are skipped. No new placeholder nodes for unknown identifiers/literals.
  • P1-c. Bare := assignment and qualified-target field lookup (2026-06-15) — NaturalParser previously required a COMPUTE/ASSIGN keyword to recognize an assignment; Natural's implicit-assignment form <target> := <expr> (no keyword) is now matched via a new BARE_ASSIGN pattern (fallback after ASSIGN_COMPUTE, restricted to := to avoid colliding with IF/comparison =). lookupVariable now also strips a <QUALIFIER>. prefix (e.g. CDBRPDA.SORT-KEY -> SORT-KEY) when the fully qualified name isn't a known local variable, so qualified field references resolve to the declared field name. New unit tests bareAssignWithoutKeywordIsRecognized / qualifiedAssignTargetMatchesUnqualifiedDeclaredField. New IngestModuleIT ingests the WGEAGB0S fixture tree (100 files).
  • P1-d. ADD/SUBTRACT/MULTIPLY/DIVIDE read/write edges (2026-06-15) — Natural's arithmetic statements read and write their operands but were not recognized at all by NaturalParser. New patterns ADD_STATEMENT, SUBTRACT_STATEMENT, MULTIPLY_STATEMENT, DIVIDE_STATEMENT add READS edges for every identifier in the operand-list/expression operands (via new addExpressionReads helper) and READS/WRITES edges for the accumulator/target (addOperandAccess helper): ADD ... TO x → reads+writes x; ADD ... GIVING x → writes only x; SUBTRACT ... FROM x [GIVING y]; MULTIPLY x BY y [GIVING z]; DIVIDE x INTO y [GIVING z] [REMAINDER r]. New unit tests covering all four statements and their GIVING/no-GIVING variants.
  • P1-e. lineNo on AstEdge (2026-06-15) — AstEdge gained an int lineNo field (the source line of the relationship/statement), threaded through both parsers' edge() helpers and all call sites. Persisted as a relationship property: buildMergeEdgeQueries() now does MERGE (a)-[:%s {lineNo: e.lineNo}]->(b) (so the same edge type between the same nodes at different lines becomes distinct relationships), and buildResolvePlaceholderTargetQueries() preserves r.lineNo when redirecting placeholder edges. Exposed via lineNo in VariableAccessLocation and CallReference (callers/callees now return one row per call site).
  • P1-f. DECIDE FOR/DECIDE ON as CONTROL_FLOW (2026-06-15) — per the construct-coverage review, DECIDE FOR/DECIDE ON (Natural's switch/case) appears in 8 of 26 WGEAGB0S fixtures but was not recognized. New DECIDE_STATEMENT/END_DECIDE patterns add a CONTROL_FLOW node (dataType="DECIDE", value = the full statement text) spanning to END-DECIDE, mirroring IF/FOR/REPEAT. Statements inside the block still get normal READS/WRITES edges. New unit test decideForIsRecognizedAsControlFlowBlock.
  • P1-g. Resolve qualified PARAMETER/LOCAL USING field references (2026-06-15) — MOVE/ASSIGN/etc. targets of the form STRUCT.FIELD (e.g. CDBRPDA.SORT-KEY) where STRUCT is a PARAMETER USING/LOCAL USING include now resolve to a placeholder VARIABLE node FIELD, CONTAINS-child of the placeholder STRUCT DATA_STRUCTURE. Previously such references either fell back to an unrelated same-named local variable or produced no edge at all. New Neo4j post-ingest step (resolvePlaceholderFieldTargets for READS/WRITES, DELETE_RESOLVED_PLACEHOLDER_FIELD_CONTAINS) redirects these placeholder fields onto the real field of the same name in the resolved structure. New unit test qualifiedFieldOfIncludedDataAreaResolvesToPlaceholderUnderThatStructure. Known limitation: two different included structures defining a same-named field converge (merge key has no parent reference); not an issue in current fixtures.
  • 18. Cross-file bare-field resolution (MOVE/ASSIGN/etc.) (2026-06-17) — closes the remaining half of the P1-c/P1-g gap: an unqualified reference (MOVE #X TO SORT-KEY where SORT-KEY is declared only in a LOCAL/PARAMETER USING area). lookupVariable gained an allowIncludePlaceholder flag (true only at explicit MOVE/ASSIGN/COMPUTE/ arithmetic operand sites). When set and the bare name is not a same-file variable, is identifier-shaped, and the module has USING includes, a module-level placeholder VARIABLE (sourceFile="") is created. New resolveBareIncludedFieldTargets redirects it onto a real included field only on a unique single match across m-[:INCLUDES]->(:DATA_STRUCTURE{sourceFile<>""})- [:CONTAINS*1..]->; orphan cleanup reaps unresolved placeholders so no spurious edge survives. New NaturalParserTest cases + IT bareReferenceToIncludedFieldResolvesToRealField.
  • 19. Field-level dataflow for shared PDAs (field-flow) (2026-06-17) — a shared PDA's fields are single nodes, so a field written in the caller and read in a transitively-called module is already connected through that one field node plus the CALLS edge — no new edge type needed. New GET /api/projects/{project}/variables/{field}/field-flow?module=X&depth=N correlates writers and downstream readers of a shared field across the call graph (wmod <> rmod, reachable via CALLS*1..depth, clamped 10), returning "field F is produced in MOD-A:120 and consumed in downstream MOD-B:45". New FieldFlow record + CLI field-flow. Caveat: reachability-based, not order-precise. Fixtures FF_SHARED.lda + FF_PROD.nat
    • FF_CONS.nat; IT fieldFlowTracesSharedPdaFieldAcrossCall.
  • 15. Dataflow analysis through the call hierarchy (2026-06-16, scoped) — captures which variables are passed at each call site. Parser (Phase 1): CALLNAT 'MOD' ARG1 ARG2 captures the positional argument list onto the CALLS edge args property; top-level PARAMETER fields tagged with paramPosition. Enricher (Phase 2): LINK_ARGS_TO_PARAMS runs after placeholder resolution, positionally joining split(r.args,',')[i] to the callee param with paramPosition = toString(i) and MERGE-ing an ARG_TO_PARAM {callSite, position} edge (new EdgeType.ARG_TO_PARAM). Query (Phase 3): GET /variables/{name}/flow-forward|flow-backward?module=&depth= follow ARG_TO_PARAM*1..depth, returning {variable, variableType, module, depth} (DataflowStep); CLI flow-forward/flow-backward. Placeholder-target resolution carries edge properties via SET r2 += properties(r) so args survive redirection. Deferred: PERFORM USING, whole-area PARAMETER USING position mapping, Java method-argument dataflow. Fixtures DF_CALLER.nat
    • DF_CALLEE.nat.

Java parser

  • J6. Ingest noise: JDK/framework filter + target/ exclusion (2026-07-07) — a by-name ingest (POST /ingest/{name}) follows every referenced type; JDK/stdlib/framework types never resolve to a project file, so they dominated the unresolved list and ballooned the dependency fan-out (one job pulled ~1158 files). New ExternalTypes denylist (uppercased simple names: java.lang/util/time/io/nio/math/concurrent/stream

    • CDI/JPA/Panache/Mutiny types) is consulted in the dependency BFS — matching MODULE refs are neither chased nor reported as unresolved, so only genuine gaps remain. The file walk now excludes target/ by default (merged with the project's excludeDirs) so build-output generated sources don't create duplicate module/entity definitions. Documented heuristic (a project class named like a JDK type would also be skipped — vanishingly rare). New JavaIngestNoiseIT (JDK types filtered while a genuine missing type is still reported; target/ copy not flagged as a duplicate); all 80 server ITs green.
  • J5. Cross-class Java dataflow (precise arg→param) (2026-07-07) — flow-forward/flow-backward across classes. Investigation found the generic LINK_ARGS_TO_PARAMS linker already matched Java cross-class (MODULE→MODULE) calls but imprecisely — (callee)-[:CONTAINS*1..]->(param at position i) linked an argument to the position-i parameter of every method in the callee class (no method disambiguation). J5 makes it precise: JavaParser stamps callerFn (calling method) and calleeMethod (invoked method) on cross-class METHOD_CALL edges; new dataflow-gated steps LINK_ARGS_TO_PARAMS_JAVA_CROSS (+ _SCOPED) build ARG_TO_PARAM from each caller-scope argument to the named callee method's parameter at that position; and the generic linker (+ its scoped variant) is guarded with r.calleeMethod IS NULL so it no longer cross-matches Java. flow-forward/flow-backward traverse ARG_TO_PARAM unchanged, so they now span classes correctly (deep-ingest only, like Natural). Overloads still match by name (over-approx); constructor-argument dataflow deferred. New fixtures fixtures/java/flow/* (CheckoutFlow.checkout(amount) → PriceCalculator.applyDiscount(basePrice)) and JavaCrossClassFlowIT (forward + backward); all 78 server ITs + 11 parser tests green.

  • J4. Virtual/override (template-method) dispatch (Java) (2026-07-07) — a template method in an abstract base (doProcessItem → clearTable/writeEntities) dead-ended at the base's abstract method. New EdgeType OVERRIDDEN_BY (base-class FUNCTION → same-named overriding FUNCTION in a subclass), materialized by the always-run, idempotent enrichment step BUILD_OVERRIDDEN_BY (name-based, transitive over EXTENDS; constructors excluded naturally since a super/sub pair never shares a name). Exposed via a new GET /modules/{name}/functions/{fn}/overrides endpoint (GraphRepository.functionOverrides → FUNCTION_OVERRIDES, returning FunctionOverride {module, name, sourceFile, startLine, endLine}) and the function_overrides MCP tool — kept function-level rather than threaded into the module-level call-tree. Name-based matching is heuristic (overloads over-match); documented. New fixtures fixtures/java/override/* (abstract template base + two concrete steps) and JavaOverrideIT (2 tests: all overrides for an overridden method, empty for a non-overridden one); all 76 server ITs + 11 parser tests green.

  • J3. Interface → implementation resolution (Java) (2026-07-07) — new EdgeType IMPLEMENTED_BY (interface → concrete impl), materialized by the always-run, idempotent enrichment step BUILD_IMPLEMENTED_BY (the derived inverse of a real IMPLEMENTS edge). Java MODULE nodes now carry an isInterface property (a small slice of the future module-kind item). New ?resolveInterfaces=true option on callees and call-tree (REST + the MCP tools): callees hops an interface callee to its IMPLEMENTED_BY implementation(s) via coalesce(impl, callee) — a single impl is a clean deterministic hop, multiple impls expand, the interface is dropped; call-tree drops interface nodes that have a known implementation (their impls are already reached via the CHA synthetic CALLS edges). Flag-gated so the default query is byte-unchanged. New fixtures fixtures/java/iface/* (single-impl Notifier/EmailNotifier) reusing the multi-impl gateway fixtures; JavaInterfaceResolutionIT (3 tests); all 74 server ITs + 11 parser tests green. Note: the interface-traversal dead-end J3 originally described was already fixed by the earlier polymorphic CHA step — this item adds the explicit edge and the caller-facing hop/suppress option.

  • J2. DI + class-literal wiring edges (Java) (2026-07-07) — call-tree/ callees on a job class returned empty because steps are wired by CDI injection and class literals, not method calls. Two new EdgeTypes: INJECTS (a class → an injected bean type: @Inject fields, and constructor params of an injection-point constructor — @Inject-annotated, or the sole constructor of a CDI-scoped class) and REFERENCES (a class → a type used as X.class in argument position, e.g. super(AccountKeyInitStep.class, …) / batchlet(refName(X.class))). Emitted by JavaParser.addWiringEdges as class-level (MODULE→MODULE) placeholder edges; added to RESOLVABLE_EDGE_TYPES so the existing resolve-placeholder finalize step resolves them cross-file. Surfaced in callers/callees (query match widened to CALLS|EXTENDS|IMPLEMENTS|INJECTS|REFERENCES) as edgeKind = INJECTS/REFERENCES, mirroring how EXTENDS/IMPLEMENTS already appear; deliberately excluded from call-tree by default (kept CALLS-only; see J9 below for the opt-in traversal). New fixtures fixtures/java/wiring/* (a job with a class-literal step + @Inject collaborator, a constructor-injected bean) and JavaWiringIT (3 tests); all 68 prior ITs green (shared callers/callees query change verified non-regressive), 11 parser tests green. Re-test 2026-07-07 confirms this works: all 3 probed PUR job digests now list their steps via callees.REFERENCES (RiskImportJob → RiskInitStep/RiskProcessingStep/ RiskEndStep; same for MultiTableImportJob, KeyTableExportJob), and the step's callers.REFERENCES names the job. The former job-root dead-end is fixed at digest level. Follow-up usability gap fixed by J9 (below).

  • J9. Opt-in traversal that follows REFERENCES/INJECTS (Java) (2026-07-07) — J2 wired the INJECTS/REFERENCES edges into callees/callers but kept them out of call-tree (CALLS-only), so there was no automatic transitive tree from a job to its steps to their repositories — an agent had to chain digest/callees calls by hand. New followWiring boolean, mirroring the J3 resolveInterfaces pattern end-to-end (REST call-tree query param, MCP call_tree tool arg, GraphRepository.callTree, CypherQueries.callTree): when true, the transitive-traversal relationship pattern becomes CALLS|INJECTS|REFERENCES instead of CALLS (both INJECTS/REFERENCES are MODULE→MODULE, same as the CALLS edges call-tree already traverses, so mixing them into one variable-length path is schema-compatible). Composes with resolveInterfaces. callees/callers unchanged (already surfaced these edges at one hop since J2) — only call-tree's transitive traversal needed the flag. Off by default. New test callTreeFollowsWiringOnlyWhenRequested in JavaWiringIT (default call-tree excludes INJECTS/REFERENCES targets; followWiring=true includes them), reusing the existing fixtures/java/wiring/* fixtures.

  • J1a. JPA/Panache repository calls as Java DB accesses (2026-07-07) — Java db-accesses/sql-statements were always empty; now repository/EntityManager/ Panache active-record calls resolve to READS/WRITES on the entity's DB_TABLE. Parser (JavaParser): tags a repository class with its managed entity (repositoryEntity, from a PanacheRepository<E>/JpaRepository<E,Id>/… generic supertype); treats a PanacheEntity[Base] subclass as an entity; and emits a DB_ACCESS candidate node (contained under the calling function, dataType = READ/WRITE/DELETE by method-name prefix, value = call text, plus javaReceiverType/javaMethod/javaArgType) for a persistence-shaped call — a write/delete verb on any non-JDK receiver, or a read verb on a repository-named / static-entity / EntityManager receiver. A small JDK denylist keeps candidate volume down (J6 will generalize it). Enrichment (RESOLVE_JAVA_DB_ACCESS, a new always-run, project-wide, idempotent finalize step): derives the entity — argument type for em.persist/merge/remove(x), the repository's repositoryEntity, else the receiver itself — follows the entity's MAPS_TO to the DB_TABLE, then MERGEs DB_ACCESS-[:USES_TYPE]->DB_TABLE (for sql-statements) and fn-[:READS|WRITES]->DB_TABLE (for db-accesses). No new EdgeType: DELETE shows as WRITES in db-accesses (edge mode) but DELETE in sql-statements (node mode), mirroring Natural. Unresolved candidates (entity not in graph) stay unlinked — never polluting db-accesses, incremental-ingest-safe (no deletion). New fixtures fixtures/java/jpa/* (Panache repo + entity, EntityManager service, active-record entity) and JavaDbAccessIT (3 tests); all 68 server ITs + 11 parser tests green. J1b deferred (roadmap): @Query/JPQL/native-SQL string parsing, derived-name filters, and no-generic custom repositories (need symbol resolution). Re-test 2026-07-07 fixed by J7 + J8 (below): JavaDbAccessInheritedRepoIT (repo → project base → Panache base, plus a constant-valued @Entity(name=…)) now passes end-to-end — db-accesses/sql-statements resolve to RISK. A regression check against the actual 9 PUR jobs is still worth doing as a follow-up acceptance pass, but the two traced root causes are fixed.

  • J7. Panache-ness inherited through a project base class (2026-07-07) — J1a only recovered repositoryEntity when a repository directly extended a Panache/JPA base type; a repository extending a project-specific abstract base (e.g. RiskRepository extends AbstractPurRepository<Risk, String>, where AbstractPurRepository<Entity, Id> implements PanacheRepositoryBase<Entity, Id>) got nothing, since the entity generic sits one inheritance hop away from the Panache marker. Parser (JavaParser): repositoryEntityType now returns null (not a bogus name) when the repository base type's first type argument is one of its own type parameters rather than a concrete class; a new panacheEntityTypeParam tags such a project base class with the name of that type parameter, alongside its own ordered typeParams; a new typeArgs edge property on EXTENDS records the concrete arguments a subclass supplies (e.g. Risk,String). Enrichment (RESOLVE_PANACHE_INHERITED_ENTITY, a new project-wide, idempotent finalize step run before resolve-java-db-access): finds the base class's panacheEntityTypeParam position in its typeParams, reads the subclass's EXTENDS typeArgs at that position, and SETs repositoryEntity on the subclass — after which RESOLVE_JAVA_DB_ACCESS (J1a, unchanged) resolves its repository calls exactly as for a directly-Panache repository. Resolves one level of indirection (the observed real-world shape). New fixtures fixtures/java/jpa/inherited/* and JavaDbAccessInheritedRepoIT (2 tests); full ac-code-server unit + integration suite green.

  • J8. @Entity(name=CONST) / @Table with a constant table name (verified 2026-07-07, no change needed) — the physical table name is not always a string literal on the annotation; RiskEntity-shaped entities use a same-class constant reference instead (@Entity(name = Risk.TABLE_NAME) with public static final String TABLE_NAME = "RISK";). Turns out resolveTableName already resolved this via existing same-class constant-value resolution (collectConstants/resolveAnnotationString, predating J1a) — confirmed by parsing the JavaDbAccessInheritedRepoIT fixture in isolation before any code change. No fix required; J7 (above) was the actual blocker for that fixture's db-accesses.

  • 12. Java parser — constructors, parameters, field READS/WRITES (2026-06-16) — constructors → FUNCTION nodes named after the class; method/constructor parameters → VARIABLE nodes (CONTAINS from the function); field READS/WRITES from this.field and unshadowed bare-name references (AssignExpr target → WRITES; compound-assign / ++/-- → READS+WRITES; else READS). Parameter/local names shadow field references. Intra-class CALLS now also scans constructor bodies. /variables/{name}/reads|writes gained FIELD to the matched node types. Known limitations (deferred to item 15): same-named params in one file and overloaded constructors converge on one node. New fixtures OrderService.java + BaseService.java.

  • 16. Java parser — constant value resolution (2026-06-16) — JavaParser.literalValue does type-aware extraction (StringLiteralExpr.asString(), Integer/LongLiteralExpr.asNumber() — strips L, Boolean/CharLiteralExpr); non-literal initializers fall through to a null value. collectConstants resolves in-class references (bare NAME and ThisClass.NAME) transitively with cycle guard. resolveAnnotationString resolves an annotation member to a string via the same-class constants (used by item 13 for @Entity(name = TABLE_NAME)). Deferred: cross-class constant resolution / ConstantRefEnricher.

  • 13. Java parser — Hibernate/JPA entity recognition (2026-06-16) — detects @Entity/@Table (table name resolved through item 16, fallback to class name), @MappedSuperclass (no MAPS_TO), and per-@Column field metadata (columnName, columnDefinition, nullable, converterType from @Convert, hibernateType from @Type, isId, declaredIn). AstNode gained an optional Map<String,String> properties (persisted via SET node += n.properties). New EdgeType.MAPS_TO. Inherited columns are collected at query time by walking (:MODULE)-[:EXTENDS*0..]->(:MODULE)-[:CONTAINS]->(:FIELD). New EntityColumn record + ENTITY_COLUMNS query + GET /api/projects/{project}/modules/{name}/columns + CLI entity-columns; DB_TABLE_COLUMNS gained a 3rd UNION branch. Fixtures SampleLegacyEntity + AbstractSampleHistorized + AbstractSampleBase.

  • 14. Java parser — general inheritance enrichment (all classes) (2026-06-16) — implemented query-time (no new edge types). New GET /api/projects/{project}/modules/{name}/functions?includeInherited= endpoint + CLI functions [--include-inherited]: walks (:MODULE)-[:EXTENDS|IMPLEMENTS*0..]->(:MODULE)-[:CONTAINS]->(:FUNCTION) (MODULE_FUNCTIONS_INHERITED) when inherited is requested, else own only (MODULE_FUNCTIONS_OWN); each row tagged with declaredIn. New InheritedFunction record. Deferred: explicit INHERITS/OVERRIDES edge types. Fixtures OrderService.java (overrides describe()) + BaseService.java.

  • 17. Java dataflow (extends item 15 to Java) (2026-06-16) — JavaParser tags each method/constructor parameter VARIABLE node with paramPosition and captures the positional argument list of each intra-class MethodCallExpr onto the CALLS edge args property. New LINK_ARGS_TO_PARAMS_JAVA enricher maps args[i] to the callee function's parameter with paramPosition = i, where the caller-side variable is the caller function's own parameter or a field of its enclosing class. The flow-forward/flow-backward endpoints are language-agnostic. Deferred: cross-class Java calls (resolved in item 20).

  • 20. Cross-class Java call graph (2026-06-17) — Java CALLS edges were intra-class only. Now resolved at the class (MODULE) level, like Natural CALLNAT: JavaParser resolves a method-call receiver to a target class — a typed field/parameter/local (svc.method()), a capitalized bare name (static Foo.bar()), or a new Foo(...) constructor — and emits a MODULE-[:CALLS]->MODULE edge to a placeholder for that class. The placeholder resolves via resolvePlaceholderTargets(CALLS) and surfaces as a DependencyRef. Scope: class-level resolution only. Fixture OrderController.java.

  • 22. Polymorphic Java call resolution + implementation ingest (2026-06-18) — (1) Fan-out — LINK_CALLS_TO_IMPLEMENTATIONS, run last in the save post-processing: for every (MODULE)-[:CALLS]->(base) where base has incoming IMPLEMENTS/EXTENDS, it MERGEs a CALLS edge from the caller to each subtype reachable via an IMPLEMENTS|EXTENDS*1.. chain (Class-Hierarchy- Analysis over-approximation). Synthetic edges carry resolvedVia:'INHERITANCE'; dataflow is intentionally not routed through them. (2) Implementation ingest — ingestModule builds a reverse index (buildImplementorIndex) so when the BFS ingests an interface/base, its implementations are enqueued too. Fixtures fixtures/java/gateway/*; IT javaInterfaceCallsFanOutToImplementations + IngestModuleJavaIT.

  • J1b. Parse @Query/JPQL/native-SQL strings + derived-name filters (Java) (found 2026-07-06, done 2026-07-07) — motivated by a batch-job analysis pass over the PUR pur-batch module (2026-07-06, 9 concrete AbstractPurBatchJob subclasses). J1a resolves calls whose entity is syntactically recoverable (repository generic param, em.persist/merge/remove(arg), Panache active-record receiver). Added: @Query JPQL/native-SQL string parsing (leading DML verb → mode, FROM/UPDATE/INTO target → entity/table, attached at the method declaration since abstract repository methods have no call site to scan); derived query-method name parsing (findAllByClientAndStatus → derivedFilter property, tagged on the existing call-site candidate); and a no-generic-entity fallback (repository interfaces with no resolvable generic type anywhere, e.g. a custom IRiskRepository, guess the entity from a declared method's return type). See x-docs/mcp-api-usage-ac-implementation.md for semantics/limits.

Project model & ingest endpoints

  • 21. Per-project root folder + server-side scan ingest (2026-06-17) — every project now carries a required root (server-side path its sources live under) and optional excludeDirs, stored on the Project node. The server walks the root and parses files: new ProjectIngestService + SourceFiles helper, with AstIngestService split into parse/save. Two new endpoints (@Blocking): POST .../ingest-all (scan the whole root) and POST .../ingest/{name} (ingest one module + its transitive deps via a server-side BFS). Both return an IngestSummary ({ingested, unresolved, duplicates, failed}). Strict duplicate detection (Java FQN; Natural (name, kind) from extension). sourceFile stored relative to the root. Old content-upload path removed. New CLI: project create, project update, project clear, ingest <name>, ingest all.
  • 12. GET /api/projects/{project}/modules — list modules in a project (2026-06-15) — new CypherQueries.LIST_MODULES/GraphRepository.listModules/ AnalysisResource.modules, returns MODULE nodes (name, sourceFile), excluding unresolved placeholders. Optional ?sourceFile=... filters to the module(s) defined in that file. Also added type=<NodeType> filter to search/identifier (400 INVALID_TYPE for an unrecognized type). New IT tests.

Module analysis endpoints (P2 reengineering support)

  • 72. A dispatch row reports its whole guard chain, not just the innermost DECIDE (2026-07-17) — found by a manual VMULTMN4 audit against the Natural source, the second such audit to pay for itself. DispatchEntry's contract is "guardField = guardValue routes to assignedField := assignedValue" — a sufficient condition. NaturalParser.guardProps walked the control-flow stack and returned on the first (innermost) value-DECIDE with an active branch (its javadoc said so outright), so for a nested DECIDE the outer guard was discarded and a conjunction was served as one condition:

    128| DECIDE ON FIRST VALUE OF #SHORT-VIEW
    165|   VALUE 'TABL'
    170|     DECIDE ON FIRST VALUE OF #FIELD-NAME
    171|       VALUE 'TX-TABLA'
    174|         MOVE P-DESCRIPTION TO YTABLMA0.TX-TABLA
    reported: #FIELD-NAME = 'TX-TABLA'
    truth:    #SHORT-VIEW = 'TABL' AND #FIELD-NAME = 'TX-TABLA'
    

    12 of that module's 69 rows (17 %) carried an incomplete condition. It over-generalises, which is the dangerous direction: this table exists to port DECIDEs to Java, and a condition that is too weak ports into a branch firing where the Natural code never would. Corpus scale, measured on disk: 824 of 7,729 value-DECIDE blocks (10.7 %) are nested, across 178 modules (YFRAMN10.nat: 199). Additive, not breaking (correcting the roadmap's own note): whenField/whenValue/whenValues keep their exact meaning — the innermost guard — so existing consumers see what they always saw; the chain travels in new whenChainFields/whenChainValues and surfaces as DispatchEntry.guards ([{field, values}], outermost first, AND-joined; each link's values OR-joined). Encoding needs two levels, so RS (U+001E) separates links above item 64's US (U+001F) for alternatives — neither can occur in Natural source. Decoding is in Java, not Cypher: zipping two nested splits back together needs index arithmetic that Cypher makes unreadable. The ac-ui migration dossier rendered the worst possible form — guardField + the lossy guardValue (item 64 already said "prefer guardValues"), i.e. the innermost guard in its ambiguous encoding, in the one screen meant for porting. It now renders the full conjunction. Live proof after refresh: VMULTMN4:174 → #SHORT-VIEW = TABL AND #FIELD-NAME = TX-TABLA, legacy fields unchanged. Vacuity-checked: reverting guardProps fails 3 of the 4 tests (guards comes back []) while the legacy-field test stays green — which is the additive claim, proven rather than asserted. The chain's order is asserted explicitly: ArrayDeque iterates innermost-first, so a missing reversal would silently yield an inside-out chain that a "contains both" assertion would accept. Left behind: #73 (a NONE branch's condition is a negation no chain can express — guards is strictly better, not total) and #74 (a duplication in the graph that item 72 exposed, not caused).

  • 67. call-tree's depth counts module hops, and says when it truncates (2026-07-17) — the last of the raw-hop siblings. callTree bounded on raw CALLS edges and returned min(length(p)) as the depth column, so the number measured how deeply the calling statement happened to be nested rather than dependency distance. Measured live on upms: WGEAGB0S's seven direct module dependencies came back at depths 1..3 (raw path lengths 3,4,5,6,7,7,8) — BGEAGFN0 reported 3 — although all seven are one module hop away. All seven now report 1. Depth comes from the path, not from ownership: min(size([n IN nodes(p) WHERE n.type='MODULE']) - 1). The decided semantics were "a FUNCTION inherits the hop count of the module that owns it", but measurement killed the literal reading: a subroutine defined in a copycode is one shared node CONTAINSed by every including module — L4N-ENTER (L4NCOPY.cpy) has 137 owners, and 156 upms functions have more than one. "Its module" has no single answer. Counting module nodes on the traversed path yields one answer per route, sidesteps ownership entirely, and gives the intended result (a root's own subroutines are depth 0). The bound was the hard part, and the first design was wrong. A raw *1..N cannot express a module-hop bound. The initial plan — a raw budget of depth × (1 + budget) — was killed in validation: at depth=1 a budget of 7 means 8 raw hops, and 8 raw hops reach 8 module levels when internal chains are short, so a depth=1 request would have expanded ~8 module levels and discarded the rest (costing what depth=8 costs today). Worse, it could report a wrong depth: a target reachable at module-depth 1 via a long internal chain and at module-depth 5 via a short one would be found only on the short route. Fix: moduleTree (item 65) supplies the modules within maxDepth module hops, and a QPP with a per-hop predicate prunes the traversal to them while expanding, so it can never wander beyond the requested depth. Verified against live upms before writing any code (QPP had already refuted one of my designs in item 65). truncated (new response field). The raw quantifier survives as a pure safety cap (agenticcode.call-tree.internal-budget, default 20). Because pruning bounds the search space, the quantifier is free — measured on upms, quantifiers 8/20/40 gave identical results at 2.0/2.6/2.1 s — so it is sized far above the observed maximum. Sizing it from a 300-module sample would have been wrong: that sample said "internal chains ≤ 6", but WGEAGB0S alone reaches ISINDATE at 8 raw hops for one module hop, and the true longest path is 10. Since any cap can cut, exhausting it now sets truncated: true rather than returning a short list that reads as complete — the exact failure mode that made items 65 and 68 bugs instead of documented limits. Conservative: true is possible for a complete result. A "known imprecision" I filed here as #71 was retracted 2026-07-17 — it was my measurement error, not a defect. The claim was that Natural external subroutines (module A performs a subroutine owned by module B) do not increment depth, "4,808 such calls in upms". That count applied the single-owner filter to the target but not the source, so every copycode-shared source function (L4N-ENTER: 137 owning modules) was counted once per owner. Source and target single-owner returns 0 — and it must, because resolvePlaceholderTargets never resolves a FUNCTION placeholder across modules, so the cross-module FUNCTION edge the claim assumed cannot exist. Path-based depth has no gap here. Two regressions the existing suite caught, both mine. followWiring (item J9) returned an empty list: the bounding BFS followed only CALLS, so with CALLS|INJECTS|REFERENCES traversal every wiring-only neighbour fell outside the module set and the predicate pruned it away — a bound must follow the same edge types as the traversal it bounds (MODULE_HOP_OUT_WIRING). And AnalysisResourceIT encoded the old semantics; checking the fixture instead of just relaxing the assertion showed the old expectation was the bug: INITIALIZATIONS and R-ADDRESS_SP are both subroutines of YADDRBN0_SAMPLE.nat itself, yet depth=1 omitted R-ADDRESS_SP — a DEFINE SUBROUTINE of the very module being asked about, hidden because it sat two raw edges away. Consequence to expect: a given depth now returns more, since the internals of modules within the bound are inside it. Both halves vacuity-checked: reverting the formula made DEPTHLEAF vanish from a depth=1 answer entirely (expected: <1> but was: null), and the truncation test is paired with an assertion that truncated is false on a normal query, so an always-true flag could not fake it. Dead code removed on the way: four unused callTree overloads (CypherQueries ×2, GraphRepository ×2).

  • 68. field-flow bounds on module hops, via a derived CALLS_MODULE edge (2026-07-17) — the last of the three raw-hop siblings item 65 uncovered. fieldFlow gated producer→consumer pairs on EXISTS { (wmod)-[:CALLS*1..%d]->(rmod) }, so its depth counted a module's internal PERFORM jumps: a consumer called from two subroutines deep sat 3 raw edges away and fell outside depth=1. The endpoint then reported nothing downstream consumes this field — a confident false negative, the same failure mode as item 65's "no DB access". Unlike item 65 this is a pairwise reachability test between two arbitrary modules, and the producer is optional ($module IS NULL), so item 65's per-root moduleTree BFS does not apply — it would have to run once per candidate producer. Fix: an enricher materialises (MODULE)-[:CALLS_MODULE]->(MODULE) (the persisted module-level projection of the call graph, same shape as EGO_NEIGHBORS_OUT), and fieldFlow bounds on that, which makes a plain variable-length bound correct again. (Correction: this was expected to supply item 67's bound too. It did not — item 67 turned out to need a set of module names to prune a QPP with, which is what item 65's moduleTree BFS already returns, so CALLS_MODULE is used by field-flow alone. It earns its keep there: field-flow is a pairwise test between two arbitrary modules, where a per-root BFS does not apply.) DELETE before rebuild is the load-bearing part, not MERGE. The edge is derived, and MERGE is idempotent only for edges that still exist — it never removes ones that shouldn't. The item-58 sweep deletes stale nodes only, so a surviving module's relationships are never reaped. Proven by removing the DELETE step: after deleting the CALLNAT from the producer's source and refreshing, field-flow still answered [STALECONS] — a dependency that no longer exists in the code. Ordering matters too: the projection runs last, after every step that adds CALLS edges (placeholder + dynamic-CALLNAT resolution, link-calls-to-implementations) or removes them (the item-62 data-literal reaping), since a projection is only as correct as the graph at the moment it runs. Both fixes vacuity-checked: reverting the bound made the field-flow test fail with an empty producer list.

  • 69. The origin file is part of an edge's identity (2026-07-16) — data loss, found by item 66's own regression fixture rather than by review. An edge was identified as MERGE (a)-[r:%s {lineNo: e.lineNo}]->(b) — (source, target, type, lineNo), with no file. Copycode expansion (item 46a) made that key ambiguous: a host statement on line 10 and a copycode statement on line 10 share it entirely, so the second SET r += e.properties overwrote the first. The endpoint then returned one row where two statements exist, and a real write was gone from the graph. Proven by reverting the fix: with MOVE 'Y' TO #KEY on COPYHOST.nat:10 and MOVE 'Z' TO &1& on MYTABLECOPY.cpy:10, variables/#KEY/writes returned only [{sourceFile=MYTABLECOPY.cpy, lineNo=11, ...}, {sourceFile=MYTABLECOPY.cpy, lineNo=10, ...}] — the host's own write had vanished, and the surviving edge even claimed viaCopycode=MYTABLECOPY. Fix: MERGE (a)-[r:%s {lineNo: e.lineNo, originFile: coalesce(e.properties.originFile, a.sourceFile)}]->(b). The fallback is a.sourceFile, not '' (as first sketched in the roadmap): for a host statement the edge really does originate in its source node's file, so r.originFile holds the true file on every discriminated edge and item 66's coalesce(r.originFile, <src>.sourceFile) keeps returning exactly what it did before. An '' default would have broken all four of those queries, since coalesce replaces only null. Scope was 20 MERGE sites, not the 1 the roadmap named: the batch merge plus 7 placeholder-resolution re-merges (which re-key the edge) and 6 dynamic-CALLNAT merges. In the dynamic ones the origin had to be threaded explicitly through the WITH DISTINCT caller, dyn, lit, which drops the src call-site node. Deliberately excluded: CONTAINS/INCLUDES/USES_TYPE (unique per node pair by construction — a discriminator cannot prevent a collision there and would add a string property to the most numerous edge type in the graph); the Java {lineNo: a.startLine} merges (copycode is Natural-only); and ARG_TO_PARAM — its collisions are real but benign, since two colliding edges carry identical (callSite, position) and no per-site payload, so discriminating them would only add parallel edges for dataflow to walk. The type test is a switch, not a static Set field: the query maps are static finals that call it from their initialisers, and a field declared after them is still null at that point (this bit — compiled clean, failed at runtime). Corpus incidence remains unmeasurable after the fact: the collision destroys the evidence one would count. Existing graphs are not migrated — the stale sweep deletes only nodes, so old edges would linger beside the new key; delete + re-ingest the project (~179 s for upms).

  • 70. Placeholder field resolution kept the copycode provenance (2026-07-16) — found while fixing item 69, by reading the queries around it. Four of the six resolution queries copied only two properties when redirecting an edge onto its real target (SET r2.value = r.value; SET r2.lineNo = r.lineNo), while the other two used SET r2 += properties(r). The four discarded viaCopycode, includedAt and originFile — the very provenance item 66 had just added. This made item 66 only half-effective, which its upms verification could not reveal: a bare field reference (#W-OPTIONS) resolves through resolveBareIncludedFieldTargets, which copies everything — and that is the field item 66 was verified on. A qualified reference (MYLDA.Q-FIELD, a field of a LOCAL USING data area) takes resolvePlaceholderFieldTargets instead and silently came back with viaCopycode: null, i.e. indistinguishable from a host statement. The split is in NaturalParser.lookupVariable: bare names go to scope.bareFields(), qualified ones to scope.placeholderFields() under a placeholder DATA_STRUCTURE. Fix: SET r2 += properties(r) in all four. Proven by reverting it — the new fixture returned expected: <[QUALCOPY]> but was: <[null]>. Note the two fixes are complementary: item 69's key incidentally carries originFile through resolution, but viaCopycode/includedAt need the +=.

  • 66. Call sites and variable accesses name the file their line is in (2026-07-16) — found by the same manual WGEAGB0S audit as item 65. The graph was already right; the API threw the answer away. The parser stamps copycode-origin edges (item 46a) with viaCopycode + includedAt and gives copycode-origin nodes the real .cpy file — exactly what this doc promised — but viaCopycode appeared nowhere in ac-neo4j-store or ac-code-server, so a line belonging to a .cpy was served as if it were a host line:

    before: #W-OPTIONS writes → sourceFile WGEAGB0S.nat, lineNo 18/20/22
    truth:  WGEAGB0S.nat:18/20/22 = the generated comment banner ("* System : Versis", ...)
            ISICINDI.cpy:18/20/22 = &1& := '0' / '1' / ' '   (&1& = '#W-OPTIONS', INCLUDE at host line 676)
    

    Host and copycode lines sat indistinguishably in one list (lineNo: 670 in the same response was a real host line). Severity differed per endpoint: variables/{name}/reads|writes made an explicit false claim (sourceFile = host + lineNo = copycode line); callers/callees gave unattributed lineNos (their sourceFileIndex is the callee's file, so they never claimed a call-site file). Scale in upms: 11,008 CALLS, 1,325 WRITES, 996 READS edges carry viaCopycode. Root cause was one missing line: CopycodePreprocessor.withVia already held origin.sourceFile() where it stamped viaCopycode, but AstEdge has no sourceFile field, so the path was dropped — edges now carry it as originFile. On top of that: AggregatedCallRef.lineNos: [int] → sites: [CallSite] (lineNo, callSiteFileIndex, viaCopycode, includedAt) — per site, because one target can be called from both the host and a copycode, which entry-level provenance cannot express. Deliberately named callSiteFileIndex, not sourceFileIndex: at entry level that word means the callee's file, and reusing it would have baked in the next misreading. variables/{name}/reads|writes return coalesce(r.originFile, f.sourceFile) plus viaCopycode / includedAt. payload already did this correctly (per-entry sourceFile + lineNo) and was the model. Breaking change: REST/MCP/CLI stayed in sync for free (MCP returns the record, the CLI prints the raw JSON); ac-ui's IdentifierPopover was carrying the same confusion — keying PERFORM lines to the caller's definition file — and now reads each site's own file. Live: ADLML02 in WGEAGB0S's callees → lineNo 26, callSiteFile src/manual/copycode/ISIYESNO.cpy, viaCopycode ISIYESNO, includedAt 673, while BGEAGFN0 → lineNo 497, WGEAGB0S.nat, no copycode; the #W-OPTIONS writes above now name ISICINDI.cpy with includedAt 676, line 670 unchanged. Tests in CopycodeExpansionIT (call site + a variable written from both host and copycode), both vacuity-checked: without originFile they fail with expected: <MYTABLECOPY.cpy> but was: <COPYHOST.nat> — the defect verbatim. Building that fixture also turned up roadmap item 69 (two edges on the same line number from different files collapse in the MERGE).

  • 65. ?depth= counts module hops, not raw CALLS edges (2026-07-16) — found by a manual endpoint audit of WGEAGB0S (upms) against the Natural source. db-accesses?depth=1..4 reported zero DB accesses for a module that reaches YGEAGBNH's tables through exactly two module calls. An empty list reads as "this module touches no database" — a confident false negative, the worst failure shape for an agent. Cause: a CALLS edge starts at the statement that makes the call, not at the enclosing MODULE node, so (m)-[:CALLS*1..N]->(hop:MODULE) also steps through internal PERFORM jumps. The real path is MODULE:WGEAGB0S → FUNCTION:GET-DATA → FUNCTION:GET-MAIN-DATA → MODULE:BGEAGFN0 → FUNCTION:READ-FILE → MODULE:YGEAGBNH — 5 raw edges for 2 module hops, so the answer only appeared at depth=5, and the required value depends unpredictably on the callee's subroutine nesting. EGO_NEIGHBORS_OUT (item 49) already defined a module hop correctly as (a)-[:CONTAINS*0..]->()-[:CALLS]->(b:MODULE) and BFS'd one hop at a time — so ego_graph?depth=2 and db-accesses?depth=2 disagreed about the same graph. Fixed by giving the transitive endpoints the same definition: new MODULE_HOP_OUT (batched over a whole BFS frontier — one query per hop, not per module) + GraphRepository.moduleTree, consumed by DB_ACCESSES_FOR_MODULES, SQL_STATEMENTS_FOR_MODULES and the ?module= scope of variables/{name}/reads|writes (whose EXISTS { (root)-[:CALLS*1..N]->(owner) } had the same flaw). Done in Java because Cypher cannot express the hop: a quantified path pattern rejects both the variable-length CONTAINS*0.. inside it ("Variable length relationships cannot be part of a quantified path pattern") and nesting. Live: WGEAGB0S db-accesses now depth=1 → 0 (correct — none of its 7 direct callees touches a table), depth=2 → 7 incl. VDB2-VERSIS_GENAGREE via YGEAGBNH at line 2617 (matches the source's FIND (1) VDB2-VERSIS_GENAGREE), consistent with ego_graph. Regression test ModuleHopDepthIT (fixtures DEPTHROOT/DEPTHLEAF: one module hop, but the CALLNAT two PERFORM levels deep); vacuity-checked — dropping the CONTAINS*0.. step fails 2 of its 3 tests, while the depth=0 test correctly stays green.

  • P2-a. Control-flow parsing (IF/ELSE/END-IF, FOR/END-FOR, REPEAT/END-REPEAT) — Natural construct #6. New NodeType.CONTROL_FLOW node per block (dataType = IF/FOR/REPEAT, value = condition/loop-spec text, line span = full block incl. matching END-*), nested via CONTAINS edges.

  • P2-b. Module "context bundle" endpoint (GET /modules/{name}/context) — aggregates functions, callers/callees, db-accesses/sql-statements, and a variable read/write summary for a module in a single response. New GraphRepository.moduleContext() + CLI context command.

  • P2-c. DATA_STRUCTURE field schema — new GET /api/projects/{project}/data-structures/{name}/fields returns the flattened field list (name, type, dataType/length, const value, immediate parent) of a DEFINE DATA/DDM structure. New CLI data-structure-fields.

  • P2-d. DB_TABLE column schema — new GET /api/projects/{project}/db-tables/{name}/columns derives column info for a DB_TABLE from the DATA_STRUCTURE used as the INTO VIEW target of SELECTs against that table. New CLI db-table-columns.

  • P2-e. Transitive call graph — new GET /modules/{name}/call-tree?depth=N endpoint returns all modules/functions transitively reachable via CALLS edges from the module (each with its minimum hop-depth). depth defaults to 3, clamped to 10. New CLI call-tree --depth.

Correctness gaps (wrong/empty results)

  • 77. Bare-field resolution ignored which module it was resolving for (2026-07-17) — found while root-causing item 74, which it does not fix (see the note there). resolveBareIncludedFieldTargets documented itself as redirecting a bare reference "only when exactly one field of that name is reachable through the module's resolved INCLUDES", but resolved project-wide across all owners at once. An unresolved bare field is one node per (name, project) (item 76 — MERGE_NODES keys on sourceFile, "" here), so (m)-[:CONTAINS]->(ph) matched it once per owning module; WITH ph, collect(DISTINCT realv) then dropped m, and the redirect's MATCH (src)-[r]->(ph) never bound src to m. Two bugs in opposite directions, both measured in upms:

    • under-resolution — two owners disagreeing on the target made size(matches) = 1 fail for everyone, including owners for whom the name was unambiguous (38 placeholders);
    • misattribution — an owner with no matching include had its edges redirected onto another module's field anyway (28 placeholders, 199 module-field pairs). Live example: BMTABBP0 was recorded writing ##MSG-NR of CDPDA-M.pda, a data area it does not include, while the other 63 modules touching that node genuinely include it.

    Both surface as a fabricated dataflow: field-flow pairs producer and consumer only when both touch the same node, so two modules sharing nothing but a field name were reported as passing data between them — a dependency an agent would act on. variables/{name}/reads|writes cannot see any of it (it matches every node with the name and reports only the accessing side), which is why the defect survived this long.

    Fixed by carrying m into the aggregation and binding src with (m)-[:CONTAINS*0..1]->(src). The *0..1 bound is exact, not a guess: a placeholder edge's src is only ever the MODULE (35,572 edges) or a FUNCTION directly under it (157,616; all 16,231 such functions are direct children) — never a CONTROL_FLOW node, so the walk also stays clear of item 75's CONTAINS cycles. Also aligned the scoped and unscoped twins, which had silently disagreed (*1..10 vs unbounded → the same module could resolve differently depending on which pass ran); the shared bound is now INCLUDE_FIELD_DEPTH, and it is required, not an optimisation, because CONTAINS is not acyclic. DELETE_RESOLVED_BARE_PLACEHOLDER_CONTAINS became per-module in the same change — its global "no edges left" test would otherwise have kept a resolved module's CONTAINS alive on the strength of another module's unresolved edges.

    Cost: the resolver is ~4.4× slower on upms (34 s → 2 m 30 s for WRITES), paid on deep ingest only. Known residue (item 76): 132 of 18,539 source nodes are copycode FUNCTIONs shared by several modules; their edge to the placeholder is a single edge, so per-module binding cannot separate what is one node. Needs per-module placeholder identity, not a better query. Verified by BareFieldModuleScopeIT (2 of 3 tests fail without the fix; the third is a control proving resolution still works, since resolution failing everywhere would also yield no flow).

  • 75. Dynamic-CALLNAT resolvers wedged on the cyclic CONTAINS graph (code + ITs done 2026-07-17; corpus re-verify pending). A whole-root deep refresh of upms hung finalize step 17 (resolve-dynamic-callnat-intra-indirect) for ~2 h without completing, blocking every later step (including item 77's bare-field resolution — so item 77 could not be corpus-verified). Cause: that step joins three unbounded (caller:MODULE)-[:CONTAINS*0..]-> anchors, and per-caller path enumeration over item 75's cyclic copycode containment is cubic (a read-only probe of the exact query for the single caller DAGNTFN0 did not finish in 60 s). Fix: the six RESOLVE_DYNAMIC_CALLNAT_* queries (INTRA / INTRA_INDIRECT / CROSS + _SCOPED twins) find a module's own statements by sourceFile equality (a MODULE is 1:1 with its sourceFile; its CALLNAT sites and WRITES statements share it) instead of descending CONTAINS — an index hash-join that cannot cycle. The cross-module resolver's dispatch variable may live in an included PDA, so it is scoped to the caller's own file or a data structure the caller INCLUDES (name-only match would pull 20958 unrelated same-named vars → the scope filter keeps 307). Proven equivalent on upms: resolve-dynamic-callnat-intra returns the identical 31 (caller, target, lineNo) triples project-wide; the rewritten indirect step runs project-wide in ~6 s instead of wedging. Existing dynamic-dispatch ITs (intra / indirect / cross / scoped / unresolved- survival, 8 tests) stay green. Does not remove the underlying CONTAINS cycles (still open in roadmap.md #75); it removes this family's dependence on traversing them. Corpus-verified on upms v68 (2026-07-18): 251 resolved dynamic CALLS (107 indirect), cross-module 113 callers / 165 edges.

  • 75b. link-args-to-params (dataflow step 27) rewritten off CONTAINS* (2026-07-18) — the same item-75 blow-up: three unbounded CONTAINS* anchors (src, cv, pv) over the cyclic copycode graph made it the single slowest finalize step on upms (~25 min). All three now resolve through the (project, sourceFile, …) index instead of descent: the call site src by sourceFile equality (module ↔ file is 1:1), and the argument variable cv / callee parameter pv over the file-set {module's own file} ∪ {files it INCLUDES}, with cv resolved before pv is unwound so the file-lists are not cross-producted. Corpus-verified: ARG_TO_PARAM still built (436 edges on the v68 recreate) and the step no longer registers as a running query across 45 s poll windows (≈ seconds, not 25 min). After this, a full deep upms finalize (~28 min) is dominated entirely by step 19 (resolve-field- placeholder WRITES), which is an item-76 cardinality problem (shared placeholder nodes), not CONTAINS* — tracked separately.

  • 76. Per-module identity for field placeholders (2026-07-18) — an unresolved field (VARIABLE/CONSTANT, sourceFile="") used to be ONE shared node per (name, project), referenced by up to 916 modules. That sharing made resolve-field-placeholder WRITES (finalize step 19) join (ph)-[:CONTAINS]->(phv) × (src)-[:WRITES]->(phv) into a ~28.7M-row cartesian → ~55 min on upms. Fix: GraphRepository.placeholderOwner stamps an ownerModule (the referencing file) onto field placeholders only; it joins both the in-memory dedup key (mergeKey) and the DB merge key (MERGE_NODES/MERGE_POSITIONAL_NODES). MODULE/DB_TABLE/DATA_STRUCTURE placeholders keep ownerModule="" and stay shared (a CALLNAT/USING target still resolves once). A companion sweep DELETE_STALE_PLACEHOLDER_NODES reaps a re-ingested file's obsolete placeholders (the file sweep skips sourceFile=""). Corpus (v69): step 19 55 min → 55 s, whole deep finalize ~59 min → ~6 min, item-77 misattribution stays 0, graph grows only +2,207 nodes (resolved placeholders are deleted in step 26), ARG_TO_PARAM well-formed (30,583 edges, 0 position mismatch). Field/dynamic ITs green.

  • P1-h. Transitive DB access resolution (2026-06-16) — /db-accesses and /sql-statements returned [] for modules whose SQL lives behind one or more CALLNAT hops (e.g. WGEAGB0S → BGEAGFN0 → YGEAGBNH). Added ?depth=N param; when depth>0, new dbAccessesTransitive/sqlStatementsTransitive queries walk up to depth CALLS hops and annotate each result with via (the intermediate module name). Clamped to 10.

  • P1-i. /db-tables/{name}/columns is effectively broken (2026-06-16) — returned only one column for tables whose column schema is encoded in .pda files. parseDataArea now detects VIEW OF <table> and creates a USES_TYPE edge from the DATA_STRUCTURE to a DB_TABLE placeholder. DB_TABLE_COLUMNS upgraded to a UNION query covering both the PDA path and the SELECT INTO VIEW path. New unit test pdaViewOfCreatesUsesTypeEdgeToDbTable.

  • P1-j. SQL statement text is truncated (2026-06-16) — every sqlStatements entry had statement cut off at ...WHERE. The DB_READ handler now reads ahead through subsequent lines until an empty line or END-FIND/END-READ, accumulating all lines into DB_ACCESS.value. New unit test multiLineFindStatementTextIsCapturedFully.

  • P1-k. FIND (1) misparsed as a table named (1) (2026-06-16) — FIND (1) <view> produced a phantom DB-access entry. Fixed by adding (?:\\(\\d+\\)\\s+)? to the DB_READ pattern. New unit test findWithRecordLimitIsNotMisparsedAsTableName.

  • P1-m. Dynamic CALLNAT <var> calls are invisible in the call graph (2026-06-21) — the parser/enricher only recorded literal CALLNAT 'NAME' edges, so variable-target calls (CALLNAT #WIF ...) produced no CALLS edge. Real case: W-MNT-N0 dispatches ~90 Wxxxx*S XML-interface programs via a single CALLNAT #WIF; the target is resolved by KDWWIFN0. Done — and v1 went further, doing both intra- and cross-module. Parser detects CALLNAT <var> (new CallKind.CALLNAT_DYNAMIC), emits a variable-named placeholder + marker edge, captures multi-line argument lists; PARAMETER USING advances paramPosition. Enrichment adds RESOLVE_DYNAMIC_CALLNAT_INTRA and RESOLVE_DYNAMIC_CALLNAT_CROSS (follows ARG_TO_PARAM into the callee — recovers the W-MNT-N0 → KDWWIFN0 → WGEAGB0S chain), plus scoped variants and placeholder cleanup. Required fixing LINK_ARGS_TO_PARAMS to source calls via (:MODULE)-[:CONTAINS*0..]->(src)-[:CALLS] so subroutine-nested CALLNATs match. Live-validated on upms: deep-ingesting W-MNT-N0 (8.4 s) yields callers/WGEAGB0S = W-MNT-N0 (edgeKind=CALLNAT_DYNAMIC) and callees/W-MNT-N0 = 124 resolved dynamic targets.

  • P1-n. Data-structure fields lack scope and often have null dataType (2026-06-21) — the parser now stamps every DEFINE DATA field/variable with a scope property (PARAMETER|LOCAL|GLOBAL|INDEPENDENT), surfaced as a scope field on both DataStructureField and IdentifierMatch (/search/identifier) — the latter exposes the motivating inline param #P-CALLED-PROG. Covered by NaturalParserTest + AnalysisResourceIT.

  • P1-o. No search by literal value (2026-06-21) — a program name like WGEAGB0S can exist in the graph only as a string literal assigned to a field. Added GET /search/value?value={v} (SEARCH_BY_VALUE, ValueMatch): quote-insensitive, returning both nodes carrying the value as a constant (kind=NODE) and literal assignments to a variable (kind=ASSIGNMENT). Missing value → 400 MISSING_VALUE.

  • P1-p. Edge provenance on call results (2026-06-21) — done via the edgeKind=CALLNAT_DYNAMIC provenance carrier: stamped on every inferred edge and surfaced through the existing edgeKind = coalesce(r.callKind, …) projection, so callers/callees distinguish dynamic (inferred) from literal CALLNAT/PERFORM (static) with zero DTO changes.

  • P1-q. /data-structures/{name}/fields over-returns across modules (field pollution) (2026-06-21) — for a structure name shared by many modules (e.g. W-WIF-A1, used by ~80 Wxxxx0S subprograms), the endpoint returned a polluted list (269 rows spanning ~80 unrelated modules). DATA_STRUCTURE_FIELDS now selects the canonical definition(s) — those with a non-empty sourceFile — falling back to empty-sourceFile stubs only when no real def was ingested, and constrains each field's parent to lie within the chosen structure (p = s OR (s)-[:CONTAINS*1..]->(p)). Live: W-WIF-A1 266→115 rows with 0 module-parented leak; YGEAGROW 430→36. Covered by dataStructureFieldsAreScopedToCanonicalDefinition.

  • P1-r. Module→data-structures endpoint (2026-06-21) — added GET /modules/{name}/data-structures (MODULE_DATA_STRUCTURES, ModuleDataStructure) returning one row per referenced structure: name, relationship (USING copybook via INCLUDES / INLINE group via CONTAINS), area (PDA/LDA/GDA/INLINE/UNKNOWN), fieldCount, sourceFile. Live WGEAGB0S now exposes its whole interface (W-WIF-A1 PDA/118, W-WIF-A2 PDA/10, W-WIF-A4 PDA/5, the LDAs) and flags unresolved CDPDA-M/CDPDA-P as UNKNOWN/0. Note: precise PARAMETER/LOCAL/GLOBAL USING scope is not stored on the INCLUDES edge (area is the proxy). Covered by moduleDataStructuresListsReferencedAreas.

  • P1-s. Referenced PDAs/copybooks not fully ingested (empty field lists) (2026-06-21) — WGEAGB0S declares PARAMETER USING CDPDA-M, but GET /data-structures/CDPDA-M/fields returned []. Root cause: the file is ingested, but CDPDA-M.pda's sole top-level group is named MSG-INFO, while a module references it by file/DDM name. Fix: when a data area has exactly one top-level group whose name differs from the file name, parseDataArea wraps it in a root DATA_STRUCTURE named after the file. Multi-top-group areas and matching-name areas keep their existing shape. Covered by parsesPdaWithRedefineGroup + usingResolvesCopybookWhoseGroupNameDiffersFromFileName (fixtures MSGAREA.pda/MSGUSER.nat). Re-ingest required.

  • P1-u. Surface a module purpose/description (2026-06-21) — parsers stamp a description property on the MODULE node: Natural scans the leading comment banner (priority **SAG TITLE: → * Title : → **SAG DESCS(n): → * Function :); Java uses the class Javadoc's first line. Surfaced as ModuleContext.description (/context) via MODULE_SOURCE_FILE (fetched as a ModuleHeader), null when no banner. Covered by NaturalParserTest, JavaParserTest, AnalysisResourceIT. Re-ingest required.

  • P1-x. /callers silently omits incoming EXTENDS/IMPLEMENTS edges (2026-06-21) — "who extends/implements this class?" returned nothing (live on pur, GET /modules/AbstractUPMFESvc/callers was [] despite 22 controllers). Root cause: callers(scope) matched only [r:CALLS], whereas callees(scope) already traversed [r:CALLS|EXTENDS|IMPLEMENTS]. callers(scope) now mirrors callees: matches CALLS|EXTENDS|IMPLEMENTS and tags inheritance callers edgeKind=EXTENDS/IMPLEMENTS. scope=internal still excludes inheritance; scope=external includes it. Covered by callersSurfaceIncomingInheritanceEdges.

  • P1-y. Double-hash (##) field names truncated to # (2026-06-22) — the data-area field regex DATA_AREA_FIELD accepted only a single leading #, so a Natural variable named ##MSG was parsed with name = "#". Widened DATA_AREA_FIELD's name class to [#A-Za-z][#\w-]* so a leading run of # is captured whole. LEVEL_FIELD and IDENTIFIER_TOKEN already handled ##. Covered by doubleHashFieldNamesAreNotTruncated. Re-ingest required for ## names.

  • P1-z. Dispatch table queryable — DECIDE/IF value→assignment correlation (2026-06-22) — analyzing the router KDWWIFN0 could recover the set of 130 dispatched programs but not the mapping (#P-OBJECT-TYPE = 'genagree' → #P-CALLED-PROG := 'WGEAGB0S'). The NaturalParser now tracks the active DECIDE ON [FIRST] VALUE OF <subject> branch: it records the subject per DECIDE block and the current VALUE '<literal>' (handling NONE/ANY VALUE resets and comma-separated value lists), and stamps every literal assignment made inside that branch — including ones nested in an inner IF — with whenField/whenValue properties on the WRITES edge. New endpoint GET /modules/{name}/dispatch-table (DISPATCH_TABLE, DispatchEntry) returns rows {guardField, guardValue, assignedField, assignedValue, lineNo}. Covered by decideValueAssignmentsCarryDispatchGuard + dispatchTableRecoversObjectTypeToProgramMapping (fixture KDDISP.nat). Live-validated on upms (2026-06-22): GET /modules/KDWWIFN0/dispatch-table returned ~140 rows reproducing the hand-built kdwwifn0-prog-routing.csv mapping (genagree → WGEAGB0S/WGEAGX0S, etc.). Re-ingest required. Known limitations: the inner IF #L-LIST (list-vs-detail) sub-guard is not separately captured (both branch assignments share the DECIDE's whenValue); the field-placeholder resolution does not copy edge properties, so a guarded write to a copybook field would lose its guard on redirect.

Token efficiency (payload shape)

  • P1-l. Repeated absolute sourceFile paths inflate payloads (2026-06-16) — callers/callees and call-tree now return a wrapper with a deduplicated sourceFiles: [...] index and items that carry sourceFileIndex: int. ModuleContext gains a top-level sourceFile field. New DTO records: CallRefResponse, AggregatedCallRef, CallTreeResponse, CallTreeItem.
  • P1-m. Call-site aggregation (2026-06-16) — callers/callees queries use collect(r.lineNo) AS lineNos with a GROUP BY (name, type, sourceFile, edgeKind); db-accesses groups by (name, mode). lineNos sorted in Java.
  • P1-n. Separate PERFORM (intra) from CALLNAT (inter) (2026-06-16) — added edgeKind field (CALLNAT/PERFORM/EXTENDS/IMPLEMENTS) to AggregatedCallRef, and ?scope=external / ?scope=internal filter on /callers and /callees. CLI gains --scope.
  • P1-o. Field projection + sub-array pagination on /context (2026-06-16) — added ?include=functions,dbAccesses,... projection and ?limit=N&offset=N pagination applied per sub-array. CLI context gains --include, --limit, --offset.
  • 7. Pagination (2026-06-17) — added ?limit=N&offset=N to callers/callees/db-accesses/search-identifier (and the transitive db-accesses variant), default limit=50, offset=0. CLI gains --limit/--offset. New IT calleesPaginationLimitsAndOffsets.
  • P1-t. /context is heavy by default — make sub-arrays opt-in, return counts (2026-06-22) (found 2026-06-21) — the live WGEAGB0S /context was ~50 KB, ~70% of it the variableAccesses array (347 entries) which the analysis never used. ?include= projection existed but the default still returned every section fully expanded. Flipped the default to lean: heavy sub-arrays (variableAccesses, large sqlStatements) return a summary by default — e.g. variableAccesses: { count, byFunction: {...}, byMode: { READS, WRITES } } — and the full list is emitted only via ?include=. A summary line replaces 347 objects; the biggest single token lever found in the WGEAGB0S analysis.
  • P1-v. Tiny /modules/{name}/digest triage endpoint (2026-06-22) (found 2026-06-21) — /context is the "expand" call (tens of KB); there was no "should I dig deeper" call. Added GET /modules/{name}/digest with a deliberately small contract: description (P1-u), function count, callers/callees names only grouped by edgeKind, DB table names, referenced data-structure names + field counts (P1-r) — designed to a token budget, not a full dump.
  • P1-w. Names-only / field-projection mode on list endpoints (2026-06-22) (found 2026-06-21) — call-tree, callers, callees, search/identifier are often used only to enumerate names, yet each row still carried sourceFile(Index), line ranges, dataType, etc. Added a ?fields=name variant that drops per-row detail to just the name (+ type) when the agent is enumerating, on top of the existing sourceFiles-index dedup (P1-l). Complements P1-t/P1-v.
  • P1-x. Dynamic CALLNAT via lookup array + keep unresolved dynamic calls visible (2026-06-22) — a CALLNAT <var> whose dispatch variable is loaded from a lookup array (e.g. ASSIGN #TBL(1) = 'WPARTD2S', #W-ACT-PROG := #TBL(#I), CALLNAT #W-ACT-PROG — WPARTX2S L994) was dropped entirely: (1) the parser missed the subscripted literal write (ASSIGN #TBL (1) = …, space before (), and (2) the unresolved dynamic marker was deleted, erasing the call site. Fix: parser now captures subscripted assignment targets; a new intra-module indirect resolver follows the dispatch var's same-line READS to the source array and resolves its literals (tagged indirect); unresolved dynamic markers are now kept (only resolved-site markers are reaped) so an agent can still see/investigate the dynamic call in callees/context.

Agent API / MCP tooling gaps

  • 78. project recreate — re-initialise a project from its own stored config (2026-07-17) — POST /api/projects/{p}/recreate (+ ?deep=true) and ac project recreate <name> [--deep]. There was no way to say "wipe upms and set it up exactly as it was": recreating meant reading project list, writing down root/language/excludeDirs/generatedDir/userExitDir, deleting, and re-supplying them by hand. That is not hypothetical — on 2026-07-17 the ac project was deleted, the re-create failed with 400 LANGUAGE_REQUIRED (mandatory since item 47), and the project was gone for a while. A typo in generatedDir would not even error; the item-47 LoC split would just report nonsense. The (:Project) shell is not deleted, only its AstNodes. The shell holds nothing but config (verified: keys(p) = name, root, language, excludeDirs, generatedDir, userExitDir), so keeping it is observably identical to delete-then-create while removing the window in which the config can be lost. The root is resolved before anything is deleted (withResolvedRoot, which already existed for refresh): unlike create — where an unreachable root costs nothing — recreate would otherwise wipe a good graph and then find nothing to rebuild from. A moved root now fails 400 ROOT_NOT_FOUND with the graph untouched. Not atomic and not claimed to be: a failure after the delete leaves the project with an empty graph (re-run to finish), and the root could vanish between check and scan. It turns the common silent failure loud; it does not make the operation transactional. No MCP tool — MCP is deliberately read-only + ingest, and this is destructive. Verified by ProjectRecreateIT (3 tests); the missing-root guard was vacuity-checked by moving the delete ahead of the check, which fails exactly that test and no other.

  • 79. One delete verb, with the blast radius spelled out (2026-07-17) — clear used to be bound twice, at two levels, with drastically different reach: ac clear emptied the entire database while ac project clear <name> deleted one project — a forgotten word apart, and only the first prompted for confirmation. ProjectCommand.ClearCommand was also a byte-for-byte duplicate of DeleteCommand (same body, same endpoint), so it was dead weight on top of an overloaded name. Both are gone; deleting is now:

    ac project delete <name>   one project
    ac project delete --all    every project (prompts; -y skips)
    ac project delete          usage error — never "delete everything"
    

    --all and a name are mutually exclusive and one is required, enforced by picocli's ArgGroup(multiplicity = "1") rather than a hand-rolled check. The confirmation prompt is inherited verbatim from the old ac clear. Covered by ProjectDeleteArgsTest (6 tests, parse-level: apiClient() builds its client inside the call, so there is no seam to inject a fake and executing would fire real HTTP; what a valid parse does is covered by ProjectRecreateIT at the API level). Breaking change: ac clear and ac project clear no longer exist.

  • Batched deletes (2026-07-17, with item 78) — DELETE_PROJECT/CLEAR_ALL were a single DETACH DELETE in one transaction. At the corpus's real size — upms alone is 454,300 AstNodes joined by ~1.3M relationships, 536,210 nodes across all projects — that is one transaction the heap must hold at once, with no server.memory.heap.max_size configured (only a 512M pagecache). Deleting was rare enough to get away with; item 78 makes it routine. Both now use CALL { ... } IN TRANSACTIONS OF $batchSize ROWS (agenticcode.delete.batch-size, default 10000). This forced the delete path off session.executeWrite onto implicit transactions — Neo4j rejects CALL { } IN TRANSACTIONS inside an explicit one ("can only be executed in an implicit transaction") — which is why the existence check is now a separate statement. Trade accepted deliberately: an interrupted delete leaves a partially emptied project that re-running finishes, versus an unbatched delete that risks not completing at all.

  • deploy.sh → manage-ac.sh (2026-07-17) — renamed (via git mv, history preserved) and given a command list. manage-ac.sh deploy is the old up, unchanged. No argument now prints help instead of deploying — a full deploy bumps agenticcode.version and rebuilds everything, which a bare invocation should not trigger by accident. No up alias: keeping two names for one action is the exact mistake item 79 removed. Added restart (restart the server without a build — also the way to abort a running server-side deep refresh, whose HTTP client can be killed without stopping the job), status (containers, server version and projects; every probe guarded so it reports a down stack instead of dying on it under set -e), and down (whole stack incl. neo4j — deliberately not -v, which would delete the graph volume). All in-repo references updated, including the two user-facing CLI strings that told people to run ./deploy.sh cli; features.md's historical entries below are left as they were written.

Found while dogfooding AgenticCode on its own codebase for the J1b task (2026-07-07): friction points where the agent-facing tools didn't answer a question the graph already had the data for.

  • 27. Generic "inspect node" tool (done 2026-07-07) — added GET /nodes/{id} (+ MCP inspect_node): every property of a node, via Node.asMap() so new parser properties show up automatically. search/identifier and search/annotation now also return id so there's a way to get one — without that they'd have been unreachable in practice. Node ids are regenerated on every re-ingest (merge key is (type, name, sourceFile, project); id is unconditionally overwritten) — documented as a limit, not fixed.

  • 28. Source-snippet-by-node tool (done 2026-07-07) — added GET /nodes/{id}/source and GET /modules/{name}/source?startLine=&endLine= (+ MCP node_source/module_source), reusing the ingest-time root resolution and the same UTF-8/ISO-8859-1 fallback decoding the parsers use.

  • 29. Text/string-literal search (done 2026-07-07) — search_identifier only matches identifier names; there was no way to search annotation names or string-literal node values (e.g. "does anything reference @Query", "does any SQL string contain orders"). Added: search/value gained a contains parameter (case-insensitive substring, vs. the existing exact match) for literal/assigned string values; a new search/annotation endpoint (+ MCP tool search_annotation) finds Java classes/methods/constructors/fields carrying a matching annotation, backed by a new generic annotations property captured at parse time on every such node (independent of any annotation's own specific interpretation elsewhere, e.g. @Entity/@Query). See x-docs/mcp-api-usage-ac-implementation.md section 7-8 for usage.

  • 31. Resolve inherited REFERENCES/wiring edges on concrete subclasses (found 2026-07-07/08, PUR pur-batch re-evaluation, done 2026-07-09) — call_tree/ module_digest only show a REFERENCES edge on the class that syntactically declares it. Where the wiring (e.g. JBeret step construction) lives in a shared abstract base class rather than the concrete subclass — AbstractKeyTableImportJob.jobSteps() wires KeyTableImport{Init,Processing, End,PassInit,Logging}Step — the concrete subclasses (KeyTableImportValidationJob/KeyTableImportCommitJob) show no REFERENCES at all and call_tree(..., followWiring=true) on them returns empty, even though call_tree(AbstractKeyTableImportJob, followWiring=true) reaches the full step tree. For 7 of 9 sibling job classes that declare their own wiring directly, followWiring=true already worked well — this item was specifically about the case where the wiring is declared on an ancestor. Fixed by also following REFERENCES edges declared on EXTENDS ancestors (transitively) when traversing followWiring from a concrete class, the same way method calls already resolve inherited members. Materialized as a synthetic edge (resolvedVia: 'INHERITANCE'), same class-level over-approximation as the rest of the CHA-style resolution (doesn't know whether the subclass overrides the specific method that declares the wiring). Covered by JavaInheritedWiringIT.subclassInheritingWiringAlsoResolvesIt (classDeclaringItsOwnWiringResolvesToday is the control test).

  • 32. Resolve DB-table access on the Repository/Entity itself, not only on the calling Logic class (found 2026-07-07/08, PUR pur-batch re-evaluation, done 2026-07-09) — db_accesses/sql_statements correctly resolve a table when a Logic class calls a repository method (via: "RiskLogic" etc., per item J1a/J1b), but db_accesses(RiskRepository)/db_accesses(RiskEntity) and dbTables in their own module_digest stayed empty — even though the repository interface statically carries its generic entity type (AbstractPurRepository<RiskEntity>/IRiskRepository) and the entity carries its @Entity(name = TABLE_NAME). Fixed by attaching the resolved table directly to the Repository/Entity module's own dbTables/ db_accesses, tagged mode: "DECLARES" (not a read/write, just "this module maps to this table") — besides the existing READS/WRITES from called functions. Lets a reference implementation's Repository/Entity pair reveal its table directly instead of requiring a call chain through whichever Logic class happens to use it first. Covered by JavaRepositoryOwnTableIT.repositoryOwnDigestListsItsTableWithoutAnyCaller and .entityOwnDigestListsItsTableWithoutAnyCaller.

  • 33. Hook-contract query for a base class (found 2026-07-07/08, PUR batch-job family comparison, done 2026-07-09) — when comparing/porting a template-method-style family (e.g. AbstractSteuertabellenProcessingStep with hooks like validate()/writeEntities()/clearTable()), there was no way to ask the graph which of a base class's methods are abstract (subclass must implement), final (fixed, subclass must not override), or a plain overridable default — this had to be read from the base class source by hand every time. Added GET /modules/{name}/functions?kind=abstract| final|overridable, sourced from the modifiers already visible to the parser (each function also carries a kind field in the unfiltered response, null for a Natural subroutine or a constructor). Turns "what must a new sibling implement" into one call instead of reading the whole abstract class. Covered by JavaFunctionKindIT's kindAbstractReturnsOnlyValidate/kindFinalReturnsOnlyCommit/ kindOverridableReturnsOnlyLog (unfilteredListsAllThreeMethodsToday is the control test).

  • 34. Bulk function_overrides for all abstract methods of a base class (found 2026-07-07/08, PUR batch-job family comparison, done 2026-07-09) — function_overrides(base, function) took one method name per call; profiling how an entire family (5+ hook methods × 5+ sibling classes) implements its contract needed one call per hook method. Added an optional function param — when omitted, GET /modules/{name}/functions/overrides returns overrides for every abstract method of base at once, grouped by method then by declaring subclass. Complements item 33 (which methods exist) with "how does each sibling implement them". Covered by JavaBulkFunctionOverridesIT.bulkEndpointGroupsOverridesByHookMethod (perMethodOverridesAlreadyWorkForEachHook is the control test).

  • 35. modules?extends={name} filter (found 2026-07-07/08, PUR batch-job family comparison, done 2026-07-09) — listing every concrete subclass of a base class previously required reading callers(base, ...) and filtering for the EXTENDS edge kind by hand. Added a direct filter on the existing GET /modules listing (?extends=AbstractSteuertabellenProcessingStep), making "show me every existing implementation of this pattern" a one-line query, consistent with the existing ?moduleKind=/?sourceFile= filters. Restricts to direct EXTENDS subclasses only (one hop, not transitive). Covered by JavaModulesExtendsFilterIT.extendsFilterListsOnlyDirectSubclasses (unfilteredListingContainsAllFourClassesToday is the control test).

Integration & ops

  • 9. docker-compose.yml committed + wired into the README (2026-07-12) — the root docker-compose.yml (Neo4j 5 + the ac-code-server container, built from src/main/docker/Dockerfile.jvm) is tracked in git and driven by the deploy.sh wrapper (up/stop/logs/cli). The README's "Getting Started" now leads with the Compose/deploy.sh full-stack path and documents a dev-mode variant that runs only the Neo4j service from Compose (docker compose up -d neo4j) alongside mvn quarkus:dev, replacing the previous manual docker run neo4j:5 command.

  • 5. MCP endpoint (HTTP/SSE) (2026-07-07) — implemented the previously empty ac-code-server/.../mcp package as a Quarkus MCP server (io.quarkiverse.mcp:quarkus-mcp-server-sse 1.9.1, augments cleanly against Quarkus 3.36.2), exposing the API to agents over HTTP/SSE at /mcp/sse (server name agenticcode). Two @ApplicationScoped tool beans mirror the REST surface by delegating to the same GraphRepository/ProjectIngestService — no query logic duplicated. McpQueryTools (22 read tools, all @Blocking, returning the same concise JSON as REST via McpSupport): list_projects, list_modules, module_digest, module_context, module_functions, module_data_structures, module_dispatch_table, module_columns, callers, callees, call_tree, db_accesses, sql_statements, data_structure_fields, db_table_columns, search_identifier, search_value, variable_reads, variable_writes, flow_forward, flow_backward, field_flow. McpIngestTools: ingest_all, ingest_module, ingest_call_graph. Deliberately read-only + ingest — destructive project create/update/delete/clearAll are not exposed. Errors reuse the REST {error,code,details} shape as MCP tool errors (PROJECT_NOT_FOUND, ROOT_NOT_SET, INVALID_TYPE, …); the deep-query tools return the same deep-ingest hint (pointing at ingest_module) when a module isn't deep-ingested. Extracted the project-root guard shared with REST into ProjectRootResolver so the validation can't drift between transports. New McpToolsIT drives the whole pipeline through tools only (connect → ingest_all → list_modules/list_projects

    • a PROJECT_NOT_FOUND error case); all 58 REST ITs still green after the resolver refactor. Docs: mcp-api-usage-ac-implementation.md gained an "Access via MCP" section.
  • 30. Version number in startup log / API / MCP (done 2026-07-08) — agenticcode.version is a manually-bumped release counter in application.properties (deliberately independent of the Maven project version, which stays a build/packaging concern). Single source of truth (VersionInfo), exposed three ways: an explicit INFO startup log line (VersionLogger), GET /api/version, and the MCP version tool — plus the MCP protocol's own server-info.version handshake field, which was previously hardcoded to 1.0.0 (already drifted from the real build) and now references agenticcode.version instead of duplicating it.

  • 36. ac-cli parity with the REST/MCP API (done 2026-07-09) — the CLI was missing wrappers for 12 endpoints that already had MCP tools: modules, module-data-structures, dispatch-table, digest, function-overrides (bulk + single), search-value, search-annotation, inspect-node, node-source, module-source, and ingest call-graph. Added all as new ac-cli commands. Also added a hard rule (CLAUDE.md §0) requiring every new/changed REST endpoint to ship with both an MCP tool and an ac-cli command in the same change, so the three stay in sync going forward.

  • 37. ac version command (done 2026-07-09) — closes the last REST/MCP-vs-CLI gap: GET /api/version already had an MCP version tool but no CLI counterpart. deploy.sh now stamps ac-cli's bundled agenticcode.properties (version=) from agenticcode.version (the same release counter VersionInfo uses) at build time (stamp_cli_version, called from both build_all and build_cli_only), so a jar built outside deploy.sh reports dev instead of a stale number. ac version prints both the ac-cli and connected server's version and exits 1 with an error if they differ (skipped when the CLI is a dev build). The interactive shell also prints the ac-cli version on startup and the same mismatch error (VersionCheck, shared with the version command).

  • 38. Log every MCP tool call (name + arguments) (done 2026-07-09) — new @McpLogged/McpLoggingInterceptor CDI interceptor pair (com.agenticcode.codeserver.mcp), applied at class level to McpQueryTools and McpIngestTools. Logs INFO com.agenticcode.mcp: MCP tool call: {toolName}({arg=value, ...}) for every @Tool invocation, reading the tool name off @Tool.name() and real parameter names via reflection (relies on the existing -parameters javac flag). Verified live via McpToolsIT.

  • 54. Source-text regex search (2026-07-14) — GET /api/projects/{p}/search/source?regex=&limit=&ignoreCase= greps the source text of all modules from disk (deduped by file, bounded by a match limit → truncated), returning {module, sourceFile, lineNo, line}. Case-insensitive by default (legacy Natural). SourceSearchService reuses the parsers' UTF-8/ISO-8859-1 decoding; 400 INVALID_REGEX on a bad pattern. REST + MCP (search_source) + CLI (ac search-source <regex> [--limit --case-sensitive]) in sync; Web-UI explorer gains a names / source mode toggle with a results list (each hit opens its module at the line). Complements search_identifier (declared names). Covered by e2e/module-explorer.spec.ts.

  • 56. search_identifier sigil-insensitive name match (2026-07-14) — the name search now ignores a single leading Natural sigil (# user, & AIV, + GDA) on both sides, so name=K-OUT-MAX finds the declared #K-OUT-MAX (and #K-OUT-MAX still works). Implemented in the shared query path (CypherQueries.SEARCH_IDENTIFIER strips left(n.name,1); GraphRepository strips the search term) so REST + MCP (search_identifier) + CLI (ac search-identifier) inherit it by construction; exact matches for sigil-less names (Java, tables) are unchanged. Covered by AnalysisResourceIT.searchIdentifierIsLeadingSigilInsensitive. (Prompted by the Web-UI popover, where Ctrl-click always captures the # but the global search box did not.)

  • 52. Function-level callers ("who PERFORMs this subroutine") (2026-07-15) — new GET /modules/{name}/functions/{function}/callers returns the FUNCTION nodes that CALLS the named subroutine/method — the intra-module PERFORM sites (Natural) or cross-class method callers (Java) — with call-site lineNos, in the same CallRefResponse shape as module-level /callers. CypherQueries.FUNCTION_CALLERS (reuses toCallRefRow/buildCallRefResponse) + MCP function_callers + CLI ac function-callers <module> <function>. Covered by AnalysisResourceIT.functionCallersListsPerformSites (DYNAIDX: INIT-TBL ← DISPATCH; a subroutine called only from the module main body has 0 function callers). UI wiring: the identifier popover's definition-click picker now attributes each PERFORM site to the calling subroutine via function_callers (useFunctionCallers), falling back to "module body" for top-level PERFORMs — the complete site list still comes from callees?scope=internal so body-level PERFORMs (e.g. INITIALIZATION) are not lost. Covered by e2e/identifier-popover.spec.ts (ADD-XML-LINE ← GEN-XML-LINE attribution).

  • 53. search_identifier scoping (sourceFile/module filter) (2026-07-15) — added optional sourceFile=<relpath> and module=<name> filters to search_identifier so a caller can ask "this name, in this module" directly (previously the paginated, project-wide search could push the module-local declaration off the page). module= resolves to the module's source file via an EXISTS { MATCH (:MODULE {name}) WHERE .sourceFile = n.sourceFile } subquery; sourceFile= filters exactly. REST (?sourceFile=&module=) + MCP (search_identifier args) + CLI (ac search-identifier --module --source-file, plus a previously-missing --type) in sync. Covered by AnalysisResourceIT.searchIdentifierCanBeScopedToAModule.

Tests & tooling

  • 8. ac-cli test coverage (2026-06-17) — module had zero tests. Added VariableAccessQueryTest (query-string builder) and CommandResolutionTest (resolveProject() precedence, apiClient() URL normalization, picocli option-parsing). 12 tests.
  • P1-p. Fix pre-existing AnalysisResourceIT failures (2026-06-16) — six IT tests were red on main but unnoticed because *IT is excluded from the default surefire run. (1) Five tests double-encoded the # in the path; switched to RestAssured pathParam so # is encoded exactly once. (2) transitiveDbAccessesAndSqlStatementsIncludeCalleeResults passed inline Natural source as the classpathResource arg; added an ingestInline(...) helper.
  • 11. Natural/Java parser construct coverage review (2026-06-15) — checked NaturalParser against the CLAUDE.md priority construct list and the WGEAGB0S fixture set (26 files): all 7 priority constructs are covered. Found and fixed one gap: DECIDE FOR/DECIDE ON (see P1-f). Other constructs (ESCAPE, RESET, EXAMINE, COMPRESS, SEPARATE, WRITE) have no graph-relevant effects and are intentionally out of scope.

Web UI — code understanding & navigation

A React/TypeScript web UI (ac-ui/) that makes the AgenticCode graph navigable for humans — primary use case: understanding legacy Software AG Natural in order to migrate it to Java (who calls a module, what it calls incl. dynamic CALLNAT, which fields flow where, which DB tables it reads/writes, what breaks on change). Built on the existing REST API; query-driven (never load the whole graph — upms is ~1000–6000+ modules), WebGL graph rendering, OpenAPI-first contract. Full vision/architecture: x-docs/ui-proposal.md. Stack: React + TS + Vite, React Router (URL-driven, shareable deep-links), TanStack Query over the generated OpenAPI client, Sigma.js/graphology for the large call-graph, CodeMirror 6 with a custom Natural language mode + built-in Java. Out of scope: parser/ingest semantic changes (the UI is a consumer) and write access to code (read-only; notes are UI-side). Item 51; M6 (scale & polish) remains open — see roadmap.md.

Backend prerequisites (Milestone M0):

  • 48. OpenAPI spec + generated TS client (backend 2026-07-13) — added quarkus-smallrye-openapi and annotated all REST endpoints (@APIResponse/@Schema, since they return raw Response) so the spec is fully typed; served at /q/openapi (+ Swagger UI in dev). TS client generation (openapi-typescript + openapi-fetch) lands with the frontend unit.
  • 49. Ego-graph subgraph endpoint (2026-07-13) — GET /api/projects/{p}/modules/{name}/graph?depth=&direction=&limit= returns a bounded module-level neighbourhood (nodes + edges, truncated flag, unresolved styling) via a BFS in GraphRepository.egoGraph. REST + MCP (ego_graph) + ac-cli (ac ego-graph) in sync.
  • 50. SPA serving + CORS (2026-07-13) — CORS enabled (quarkus.http.cors.enabled=true), restricted to the Vite dev origins. ingestStatus/ingestDepth now joined into the GET /modules list rows (was only on inspect_node) for the UI's status badges.

Frontend milestones (item 51):

  • M0 Foundation + API contract (2026-07-13) — items 48–50 (backend) + React/Vite frontend (ac-ui/): project picker, virtualized module explorer (filter + search), ingest-status badges, project/module refresh, thin module-detail from context, generated OpenAPI TS client (openapi-typescript
    • openapi-fetch).
  • M1 Navigation (2026-07-13) — whole-file source endpoint (REST + MCP module_source + CLI ac module-source, omit range = whole file); CodeMirror 6 source viewer with a custom Natural StreamLanguage mode + built-in Java, Ctrl/⌘-click name-based identify (popover via search_identifier), caller/callee panels, lazy call-tree (expand-on-demand via callees, cycle-guarded), URL-driven tabs + line deep-links, STALE_SOURCE refresh prompt.
  • M2 Call-graph visualisation (2026-07-14) — interactive ego-graph (Sigma v3 / graphology, WebGL) as a lazy-loaded "Graph" tab: seeds on the current module via the item-49 /graph endpoint, ForceAtlas2 auto-layout, click-to-select, expand-on-demand (one hop, merged), direction/depth/limit controls + truncated badge, unresolved/dispatch (CALLNAT_DYNAMIC)/inheritance edge styling + legend, open-module (disabled for unresolved). No backend change (pure consumer).
  • M3 Migration dossier (2026-07-14) — a lazy per-module Dossier tab (MigrationDossier.tsx) bundling the "porting profile": payload (I/O contract), data structures (expand-on-demand → fields), DB-access matrix (READ/WRITE badges), SQL statements, and the dynamic-CALLNAT dispatch table. Line numbers deep-link into the Source tab (?tab=source&line=). Pure consumer of existing endpoints (payload, data-structures/data-structures/{name}/fields, db-accesses, sql-statements, dispatch-table). Follow-ups: (a) a program-view "ingest +callers/callees" button that deep-ingests the program's whole call-graph neighbourhood (the module + its transitive callers and callees, and via each module's USING/INCLUDE fan-out their data structures) — POST /refresh/{name}?scope=neighborhood (seeds the multi-module BFS ingest with the both-direction ego-graph closure; ProjectIngestService.ingestModules)
    • CLI ac refresh <name> --neighborhood (no MCP tool — refresh is deliberately REST+CLI only); nginx /api proxy timeout raised so the long request survives; (b) source-by-path so a USING data area's field line-links open its own file, not the current module — new GET /api/projects/{p}/source?file=&startLine=&endLine= (reuses SourceSnippetService, rejects root-escaping paths with 400 INVALID_SOURCE_FILE) + MCP file_source + CLI ac file-source; the Source tab gains a ?src=<file> mode with an "included file … back to module" banner.
  • M4 Data-flow & impact (2026-07-14, M4.1–M4.3) — flow-forward/backward / field-flow visualisation + "what breaks?" impact analysis.
    • M4.1a subroutine navigation — Ctrl/⌘-click a subroutine name resolves via the scoped module_functions (reliable despite dozens of same-named subroutines across modules — the global search is capped): clicking a reference (PERFORM) jumps to the definition; clicking the definition goes to its caller(s) — the PERFORM sites from callees?scope=internal lineNos (one → jump, many → a picker). Covered by e2e/identifier-popover.spec.ts (Playwright; 5 cases incl. the upms 60-match reproduction).
    • M4.1 actionable variable matches — identifier-popover non-MODULE matches now expand: show dataType/scope, a jump-to-definition (opens sourceFile:startLine, via the M3 source-by-path infra so a field defined in a USING PDA opens its own file), and lazy reads/writes lists (variable_reads/variable_writes, VariableAccessLocation) with each access site clickable to its sourceFile:lineNo. MODULE matches still navigate.
    • M4.2 flow visualisation — a dedicated Data-flow tab (DataFlowView.tsx) for a selected variable (URL ?tab=flow&flowvar=, seeded on the current module): backward (flow_backward) and forward (flow_forward) DataflowStep chains (indented by depth; variable re-targets the trace, module opens it) and field-flow producer→consumer pairs (field_flow, line numbers deep-link into Source). Reached via a "data-flow →" action in the identifier popover. The flow endpoints auto-deep-ingest the scoped module; a 409 (warm couldn't complete) shows a DEEP_INGEST_REQUIRED prompt wired to the neighbourhood-ingest button.
    • M4.3 impact ("what breaks?") — a per-module Impact tab (ImpactView.tsx): the transitive callers (blast radius if the module changes), via the ego-graph direction=in (useImpactCallers, depth 10 / limit 500), grouped by call distance (direct callers vs. N-hops-away), each dependent clickable to open; unresolved dynamic-dispatch callers shown but not navigable; truncated badge when the node cap is hit.
  • M5 Understanding boosters (2026-07-14, the consumer-feasible parts).
    • M5.1 source + structure side by side — a filterable, collapsible outline (SourceOutline.tsx, module_functions) beside the Source editor; clicking a function scrolls the editor to it (SourceView re-scrolls on line change without a remount — also fixes repeat dossier line-links). Hidden while viewing an included file.
    • M5.2 notes + saved views — per-module migration notes (ModuleNotes.tsx, in Overview) and header saved views (SavedViews.tsx), both persisted to localStorage via useLocalStore (the API is read-only, so annotations stay browser-side).
    • Deferred (need more than a consumer): diff after refresh (backend must expose before/after snapshots) and LLM summaries (no LLM in the stack; the existing module_context.description is shown in Overview as the current summary).
  • M6 — Explorer regex filter (2026-07-14) — the module-list search box accepts a case-insensitive regex over name/sourceFile (falls back to substring while the pattern is incomplete). Covered by e2e/module-explorer.spec.ts. (Rest of M6 — virtualisation/large-graph performance, multi-project, auth, theming, export — remains open in roadmap.md.)

UI-sweep bug fixes (2026-07-14/15)

Found via a UI test sweep + source cross-check.

  • 55. Dispatch table: multi-value DECIDE branch guardValue artifact (2026-07-14) — a DECIDE ON … VALUE 'X', ' ' branch (a real literal + a Natural "also match blank" catch) yielded a dispatch-table guardValue of "X, " (the blank ' ' trimmed to empty, leaving a trailing ", ") instead of X. Fix: NaturalParser.quotedLiterals now drops whitespace-only alternatives (keeping one blank only if a branch has nothing but blanks, so the guard isn't lost); genuine multi-literal branches ('A', 'B') stay comma-joined. Verified live on WGEAGB0S (3 rows @533/537/545 now clean, 0 trailing-comma artifacts across all 11 rows). Covered by NaturalParserTest.multiValueBranchDropsBlankAlternativeFromGuard.
  • Dossier data-structure fields shown out of order (2026-07-14) — the data-structures/{name}/fields API returns fields in graph order ([12,26,40,55,56,41,…]); the UI now sorts by startLine so a data area reads top-to-bottom (this out-of-order display was what made a correct line number look "wrong" earlier).
  • 57. Field format glued to the name (no space) parsed as one token (2026-07-15) — a DEFINE DATA field whose format is attached with no space — e.g. 01 KEY(A1/1:3,1:V), 01 #DELIMITER(A1) (common in NATURAL-CONSTRUCT generated code like CDRANGE) — was stored with the whole token as the identifier name (KEY(A1/1:3,1:V)) and dataType=null, and thus mis-typed as a DATA_STRUCTURE instead of a VARIABLE. 234 of 376 CDRANGE identifiers were affected, making them unsearchable (a search for KEY found nothing) and dataType-less. Root cause: the LEVEL_FIELD name group (\S+) greedily swallowed the attached (…). Fix: name group now stops at ( ([^\s(]+) — Natural identifiers never contain one — in both NaturalParser.LEVEL_FIELD (deep) and NaturalCoarseScanner.LEVEL_FIELD (Tier-1 identifier index), which carried the identical bug. Covered by NaturalParserTest.fieldFormatAttachedWithoutSpaceIsSplitFromName + NaturalCoarseScannerTest.fieldFormatAttachedWithoutSpaceIsIndexedByBareName. Verification note (superseded by #58): the old glued nodes once required a project re-create to clear; since #58 a plain refresh reconciles and purges them.
  • 58. Stale nodes survive refresh — diff-based reconciliation (2026-07-15) — nodes MERGE on (type, name, sourceFile, project), so a re-parse that renames/removes a field (e.g. after the #57 fix) produced a sibling node and the old one lingered — refresh MERGEd but never deleted. On the live ac/legacy graphs this left Tier-1 identifier-index nodes created at project-create time coexisting with the deep parse's nodes (CDRANGE: 460 = 234 stale-glued + 226 clean), inflating counts and returning ghost matches; the only remedy was a project re-create. Fix (roadmap option b): after a batch of fresh nodes is merged, GraphRepository.mergeResults collects each real source file's fresh (canonical) node ids and runs CypherQueries.DELETE_STALE_FILE_NODES — MATCH (n {project, sourceFile}) WHERE NOT n.id IN freshIds DETACH DELETE n. Because MERGE_NODES overwrites node.id to the fresh UUID, a surviving-key node keeps a fresh id (kept) while a renamed/removed field keeps its old id (deleted); moved positional nodes (DB_ACCESS/CONTROL_FLOW) are swept the same way. Gated by a reconcile flag threaded through persist/persistBatch: true for every full parse (whole-root refresh — both deep=true and the default call-graph pass — and the per-module deep refresh/{name}), false for the coarse Tier-1 scan (subset node set; runs only on an empty graph at create, so nothing to delete). sourceFile="" placeholders (shared cross-file targets) are never swept; reconciliation runs at persist-time, before enrichment, so enrichment-derived nodes/edges are unaffected and cross-module edges are rebuilt by finalizeProject. Does not remove nodes for deleted files (still roadmap #43). Covered by RefreshReconciliationIT (rename a declared field on disk → refresh → old identifier gone, new one present); verified live (ac whole-root refresh, 373 files, clean). No REST/MCP/CLI surface change — refresh already exists on all three; the behaviour change is internal.
  • 59. Shared Natural field-declaration tokenizer (2026-07-15) — the LEVEL_FIELD line pattern and its NN [REDEFINE] name (typeSpec) rest decode were duplicated verbatim in NaturalParser (deep) and NaturalCoarseScanner (Tier-1), so #57 had to be fixed twice and the two tiers could drift. Extracted NaturalFieldTokenizer (record NaturalField{level, redefine, name, typeSpec, rest} + static parse(line)), the single owner of the pattern; typeSpec is the raw parenthesised content (what the deep parser stores as dataType, byte-identical to before), with derived baseFormat()/ arrayDims() splitting it and constValue()/initValue() pulling the trailing clause. Both tiers now call the tokenizer instead of holding private copies; node-type/emission policy is unchanged per tier (a pure refactor — the #57 regression tests in both NaturalParserTest and NaturalCoarseScannerTest stay green). New NaturalFieldTokenizerTest (12 cases) covers the edge matrix: format with/without a leading space, A1/1:3,1:V dims, REDEFINE, CONST<>/INIT<>, #/##/&/+ sigils, dimension-only group arrays, and group (no-format) fields. Kills this class of tokenizing bug.
  • 61. Comments parsed as CALLNAT targets (2026-07-16) — found dogfooding WGEAGB0S (upms): the unanchored CALLNAT/CALLNAT_DYNAMIC patterns matched inside comments, so callees/unresolved were polluted by phantom modules. Evidence: tree-wide false unresolved WAS/RESULTED/DOES (from prose like * What the callnat does:, /* …last callnat resulted in end-of-data); WGEAGB0S's ISINDATE call-sites included commented lines 526 & 600 (* CALLNAT 'ISINDATE'); a phantom ADLML02 at copycode lines 18 & 22 (* CALLNAT 'ADLML02'). Root cause: the deep NaturalParser main loop matched on the raw line (no full-line * skip, no inline /* strip), and the coarse NaturalCoarseScanner stripped inline /* but not a leading *. (Anchored patterns — ^\s*PERFORM, ^\s*DECIDE, ^\s*INCLUDE — were already immune; only the unanchored CALLNAT ones leaked.) Fix: both parsers now skip a full-line comment (first non-blank char *, incl. **SAG) and strip inline /* before statement matching. Verified live: WAS/RESULTED/DOES gone (WGEAGB0S tree unresolved 21→17), ISINDATE now [597,626,1267], ADLML02 now [26] (real call only). Covered by NaturalParserTest.commentedAndInlineCommentCallnatsAreNotParsedAsCalls and NaturalCoarseScannerTest.commentedCallnatIsNotIndexedAsACall.
  • 62. Data literals recorded as false MODULE call targets (2026-07-16) — found in a corpus-wide sweep of upms (6311 files): 45 distinct data names (browse keys / codes such as CO-TABLA, COD-EMISOR, AGENT-SP, NAME-DESC-SP) sat in the graph as placeholder MODULE callees, carrying 282 false CALLS edges — polluting callees/call-tree/unresolved across the corpus. Root cause: a CALLNAT <bareword> (a browse key reaching the call site through a copycode/macro argument, e.g. INCLUDE YFRAMGC2 'C-MOD-GET' '"CO-TABLA"' expanding to CALLNAT &3& …) is parsed as a dynamic call and gets a placeholder target. It never resolves — no module of that name exists — and DELETE_DYNAMIC_CALLNAT_PLACEHOLDER_EDGES only drops markers that did resolve, so the false target was kept forever. Fix: a new project-wide enrichment step delete-data-literal-call-placeholders (CypherQueries.DELETE_DATA_LITERAL_CALL_PLACEHOLDERS), running after the dynamic-call resolution and marker cleanup and before stamp-unresolved-placeholders so the unresolved flags see the reaped graph. It reaps a placeholder only when all four hold: (1) the name carries no Natural sigil (#/&/+) — a genuine dispatch variable is always a sigil'd user variable, so CALLNAT #WIF-style markers stay; (2) the name is a real VARIABLE/CONSTANT of the project (DATA_STRUCTURE is deliberately excluded — a module and its interface PDA routinely share a stem name); (3) no real MODULE of that name exists (else it would simply resolve); (4) every incoming CALLS edge is inferred (CALLNAT_DYNAMIC/INCLUDE_MACRO) — a static CALLNAT 'X' is trustworthy, which protects the real external modules RPC-CNTX and USIX081X that collide with field names. Covered by DataLiteralCallCleanupIT (bareword CALLNAT CO-TABLA reaped; sigil'd CALLNAT #DISP marker kept; static CALLNAT 'REALMOD' resolved). Verified live on a clean upms re-create (2026-07-16): delete-data-literal-call-placeholders: 139 ms; rels +0/-274, nodes +0/-42, and the only survivors of the name-based suspect query are RPC-CNTX and USIX081X — both reached exclusively by static CALLNAT edges, i.e. gate 4 protecting them exactly as designed.
  • 63. CALLNAT matched inside a string literal (2026-07-16) — the last of the unanchored-CALLNAT family (#61 = comments, #62 = data literals). The CALLNAT/CALLNAT_DYNAMIC patterns match the keyword anywhere on a line, so prose inside a quoted string fabricated a call. Filed as "low, 1 occurrence"; that was wrong on both counts — a corpus sweep of upms found 3 distinct call sites (each doubled across src/ and generated_src/): PRINT '==> callnat before DREQUFN0' (DREQUDN0) → before; WRITE(#MSG) 'NACH CALLNAT ISINGEAG:' (JE0012N0) → ISINGEAG; #ERR-TYPE := 'Callnat USIA008N' (JA0016N0) → USIA008N. Two of the three are the dangerous class: ISINGEAG and USIA008N are real modules, so the phantom was not placeholder noise but a real → real CALLS edge in the call graph (PROCESS-AGENT-CHANGE → ISINGEAG, MAIN-PART → USIA008N) — and one invisible as a problem, since only before carried unresolved: TRUE while the other two were NULL precisely because a real module of that name exists. Item 62's reaper structurally cannot clean these (its gate 3 requires that no real MODULE share the name), so only a parse-time fix removes them. (The hundreds of 'Start of Callnat'-style lines are harmless: a quote directly follows the keyword, so no identifier matches.) Fix: new shared NaturalLines (the NaturalFieldTokenizer precedent from item 59) holding isInsideStringLiteral/findOutsideStringLiteral plus the formerly duplicated stripInlineComment; both tiers now gate the two CALLNAT matches on the keyword start lying outside a quoted literal — testing the start, not the whole match, is what keeps a real CALLNAT 'MOD' working (keyword outside, argument inside). Natural's '/" delimiters and doubled-delimiter escapes ('IT''S') are honoured. Considered and rejected: handling a /* inside a literal (which stripInlineComment would truncate, leaving an unbalanced quote) — measured 0 such lines that also contain CALLNAT, so it stays out rather than buying speculative complexity. Covered by NaturalParserTest.callnatInsideAStringLiteralIsNotParsedAsACall and NaturalCoarseScannerTest.callnatInsideAStringLiteralIsNotIndexedAsACall, both verified to fail against pre-fix behaviour. Verified live on a clean upms re-create (2026-07-16): the before MODULE node is gone entirely, and the fabricated CALLNAT_DYNAMIC edges into ISINGEAG/USIA008N (2 each) are gone while their genuine static CALLNAT edges (10 and 2) remain — the real → real corruption is removed without touching a single real call.
  • 64. Ingest summary contradicted the graph; dispatch guards lost their alternatives (2026-07-16) — found by dogfooding a deep ingest of WGEAGB0S (upms) and hand-checking every endpoint against the Natural source. Two independent defects: (a) unresolved reported data fields as missing modules. The deep ingest listed MODULE CO-TABLA, MODULE COD-ENTIDAD, MODULE NAME-DESC-SP, MODULE NAME-VALUE-SP — all (A8) fields, no such module anywhere — although the graph was already clean (the scoped enrichment reaped them: rels +0/-7, nodes +0/-5). Root cause: IngestSummary.unresolved is built during the BFS walk from a filename index (ProjectIngestService:581) and never consults the graph, so it applied none of item 62's gates. This made the item-62 fix incomplete: it cleaned the graph surface and missed the summary — an asymmetry item 63 doesn't share, since a parse-time fix stops the ref existing at all. Fix: filter the unresolved refs through item 62's gates — no Natural sigil, reached only by inferred (CALLNAT_DYNAMIC/INCLUDE_MACRO) edges (reconstructed from the parse result's CALLS edges, since DependencyRef carries no provenance), and the name is a real VARIABLE/CONSTANT. That last gate is asked of the graph, project-wide (CypherQueries.DATA_FIELD_NAMES), not of the ingested tree: a first, tree-local attempt fixed only 2 of 4 because NAME-DESC-SP is declared in BMTABBN2.nat, which is outside WGEAGB0S's 164-file tree. The static-CALLNAT gate keeps genuinely-missing modules whose names collide with field names reported (the RPC-CNTX/USIX081X class). Verified live: WGEAGB0S's unresolved list went 18 → 13, all four false positives gone, while the USR* modules, sigil'd #GETSHORT-MODUL, NDBERR and the VDB2-* views all stayed. Covered by DataLiteralUnresolvedSummaryIT (incl. the EXTMOD case: a declared field that is also a statically called missing module must stay reported), verified to fail pre-fix. (b) dispatch guards dropped their VALUE alternatives. guardValue is a single comma-joined string, so VALUE 'GENAGREE-WOUT-SP', ' ' (a Natural "also catch blank") silently lost the blank — quotedLiterals drops whitespace-only alternatives to avoid a trailing ", " — and a multi-alternative branch yields a synthetic "A1, A2" the guarded field never equals. Fix: the parser now also records whenValues, every alternative in source order including blanks, and DISPATCH_TABLE splits it into a real guardValues list. guardValue is untouched, so REST/MCP/CLI/UI keep working (all three surfaces return the record directly, so the field propagates automatically). Encoded with U+001F rather than modelled as a list because AstEdge.properties is Map<String, String>; US cannot occur in Natural source, unlike the comma that made guardValue ambiguous. Verified live: WGEAGB0S now reports 3 branches that also catch blank (0 before), and correctly distinguishes the DECIDE at 531 (blank alternatives) from the one at 602 (none). Covered by NaturalParserTest.multiValueDecideBranchKeepsEveryAlternativeIncludingBlank. Scale, measured honestly: 121 blank-bearing VALUE clauses corpus-wide, 3 real in WGEAGB0S; 0 joined multi-value guards observed in the deep-ingested subset (238 guarded WRITES), so the joined-string flaw is real in the code path but unobserved in practice — a full deep ingest would be needed to settle its true rate.

Qualified-field access & re-ingest reconcile (items 78, 79, 74) — 2026-07-18

Three defects found dogfooding a deep WGEAGB0S audit of upms, all in the Natural field READS/WRITES path, fixed together. All three verified with Testcontainers ITs (independent of the live server); full ac-code-server IT suite green (190/0/0) with no regressions.

  • 78 — group-name-qualified access dropped. A field written qualified by its data area's inner group name rather than the USING name (MSG-INFO.##MSG-PGM := *PROGRAM, where MSG-INFO is the top group of a USING CDPDA-M) produced no edge: NaturalParser.lookupVariable recognised only USING-name qualifiers and fell through to a local-only lookup. Fix: when the qualifier is not a known include and the field is not local, mint a bare included-field placeholder (gated allowIncludePlaceholder && hasIncludes && identifier-shaped) so resolveBareIncludedFieldTargets redirects it to the real PDA field. Test QualifiedFieldResolveIT.groupQualifiedFieldAccessIsCaptured.

  • 79 — qualified read inside an expression dropped. The read side of an assignment/expression/IF tokenised on IDENTIFIER_TOKEN (no dot), so #C := QCPDA.QC-DESC and IF MSG-INFO.##MSG-NR EQ … split into two non-resolving tokens, and IF conditions were not read-scanned at all. Fix: new OPERAND_TOKEN keeps a dotted operand whole; addExpressionReads resolves a qualified token with allowIncludePlaceholder (legitimate explicit field access) but a bare token without (no placeholder for the "expression soup", so item 74's by-name path is not fed); the assign-RHS loop and the IF handler both call it. Test QualifiedFieldResolveIT.qualifiedFieldReadInsideAnExpressionIsCaptured. (The originally-filed "qualifier does not disambiguate a shared name" was disproved by a controlled fixture — the qualifier resolves correctly; the live "unresolved" observation was item-74 stale state.)

  • 74 — USING field misattributed to a same-named local group in another subprogram. WGEAGB0S's BXFRABA4.#L-FORWARD-VIA-PRI was linked not only to the real BXFRABA4.pda but also into unrelated JX0124N6/JX0129N0. (First mis-diagnosed as a stale-edge coexistence surviving re-ingest — a clean recreate disproved that: the misattribution is created fresh in a single pass.) True cause: a Natural USING X references a data area (PDA/LDA/GDA) — always a top-level member of a data-area file — but the placeholder resolver (buildResolvePlaceholderTargetQueries + the two by-name field resolvers) matched USING X to any DATA_STRUCTURE named X, including a 1 BXFRABA4 group those subprograms declare inline, and linked every field under it. Fix: a DATA_STRUCTURE placeholder now resolves only to a real area not owned by a MODULE (NOT EXISTS { (:MODULE)-[:CONTAINS]->(real) }) — a file-level data area, never a program-internal group. Data-area files produce no MODULE node, so real PDAs are kept and inline .nat groups dropped. Test UsingResolvesToDataAreaNotLocalGroupIT.

    • Secondary safeguard (separate scenario, kept): mergeResults also runs DELETE_STALE_RESOLVED_FIELD_EDGES in the reconcile (deep-re-ingest-only) branch — deletes a re-parsed file's prior cross-file READS/WRITES to a VARIABLE/CONSTANT so finalize rebuilds them, so a module that changes which area it includes does not keep the old resolved edge. Scoped to VARIABLE/CONSTANT and cross-file targets; gated on reconcile. Two-phase test QualifiedWriteReconcileIT.reIngestDeletesTheStaleResolvedFieldEdge.

Parallel parse phase (item 24) — 2026-07-18

The ingest parse phase (read file + parse + shell/user-exit metric enrichment) ran as a sequential loop over all candidate files. It is per-file independent — the JavaParser/NaturalParser instances hold no mutable state, and copycodes/userExit are read-only — so ProjectIngestService.ingestRoot now submits one task per file to Executors.newVirtualThreadPerTaskExecutor() (extracted helper parseCandidate). Results are collected in candidate order (the futures list is parallel to candidates), so cross-file duplicate detection and persist order stay deterministic; a per-file parse/read failure is still recorded as an IngestSummary.Failure for that file only. Full IT suite green (192/0/0).

Qualified group-target resolution (item 80) — 2026-07-18

A Natural qualified reference can name a group (a DATA_STRUCTURE), not only a leaf field — e.g. WGEAGB0S writes BGEAGBA0.#P-DESC-NAME, a level-1 group of BGEAGBA0.pda. The qualifier-scoped field resolvers matched a real target of type VARIABLE/CONSTANT only, so such writes/reads stayed unresolved (sourceFile=""). Fix: the two qualifier-scoped (INCLUDES-gated) resolvers (buildResolvePlaceholderFieldTargetQueries + …ScopedQueries) now accept a DATA_STRUCTURE target as well (realv.type IN ['VARIABLE','CONSTANT','DATA_STRUCTURE']). Left the bare-field resolvers (which carry a size(matches)=1 guard) leaf-only, so adding groups cannot turn a previously-unique bare match ambiguous. Test QualifiedGroupTargetResolveIT.qualifiedWriteToAGroupResolvesIntoTheDataArea. (The #MAP-T.* reads that were unresolved in the same WGEAGB0S snapshot are leaves, a separate concern, not covered here.)

Group-qualifier field resolution (item 81) — 2026-07-18

A qualified reference can name a leaf via a group the parser cannot see as a USING member — e.g. WGEAGB0S reads #MAP-T.V-ID, where #MAP-T is a group inside the included BGEAGA01.pda and the leaf V-ID also occurs in dozens of other PDAs. After items 78/79 the read was captured but fell to the bare-field resolver, whose size(matches)=1 guard rightly refused the project-wide-ambiguous leaf, so it stayed sourceFile="". Fix: NaturalParser.lookupVariable now keeps the group qualifier as a qualifierGroup property on the bare placeholder (cached under the compound STRUCT.FIELD key so a group-qualified and a truly-bare reference to the same leaf are distinct); the two bare-included resolvers (buildResolveBareIncludedFieldQueries + …ScopedQueries) then keep only a candidate nested under a group of that name, so the qualifier pins the leaf to the one right PDA. The placeholder is a bare VARIABLE under the module (not a DATA_STRUCTURE area), so the by-name resolvers never see it — no #74-style misattribution risk. Test GroupQualifiedLeafResolveIT.groupQualifierDisambiguatesAnAmbiguousLeaf; full IT suite green (193/0/0).

Read-path CONTAINS bounding (item 75 read path) — 2026-07-19

The runtime read queries traversed a module's statements with an unbounded (m)-[:CONTAINS*0..]->(src). CONTAINS is not acyclic (shared copycode CONTROL_FLOW nodes, parser line-range artefacts), so on a heavy module (ACCNPE01) that expansion blew up and callees/digest/context hung (callees >2 min). Every edge source is a MODULE (depth 0) or a FUNCTION that is a direct CONTAINS child (depth 1) — verified corpus-wide (0 sources deeper; DB_ACCESS parents only FUNCTION/MODULE) — so the traversal is provably equivalent when bounded to *0..1, which cannot walk the cycles. 20 read-side traversals updated (callees, MODULE_HOP_OUT, DISPATCH_TABLE, EGO_NEIGHBORS_*, VARIABLE_ACCESSES, DB_ACCESSES, SQL_STATEMENTS, FUNCTION_CALLERS, SEARCH_BY_VALUE, fieldFlow, BUILD_CALLS_MODULE, and the *_FOR_MODULES batch variants). Query-only change (no recreate). Guard ReadPathBoundedTraversalIT (EXPLAIN asserts no unbounded CONTAINS expand); full IT suite 193/0/0. Live (v76): ACCNPE01 callees

2 min → 0.16 s, digest >10 s → 3.3 s, context >10 s → 3.1 s.

Version bump moved from deploy into the build — 2026-07-19

agenticcode.version (the integer build counter reported by /api/version and used by ac version for stale-CLI detection) used to be incremented by manage-ac.sh deploy (shell bump_version), so a plain mvn clean install never bumped and only a deploy did. It now lives in the Maven build: a new ac-mvn-plugins mojo bump-version, bound to ac-code-server's generate-resources phase, increments the counter and stamps the same number into ac-cli's agenticcode.properties. Binding to generate-resources (before process-resources) means the freshly built jar already reports the bumped number, keeping the jar's baked version and the CLI stamp in lock-step. The mojo only acts when a requested goal is package/install/deploy (read from MavenSession.getGoals()), so mvn test, mvn compile and quarkus:dev do not bump; deploy bumps because it runs a full install. Every such build bumps unconditionally (no source-change gating — high numbers are harmless). manage-ac.sh's bump_version was removed; stamp_cli_version is kept only for the CLI-only rebuild path (cli), which syncs the CLI to the current server version without bumping. Logic unit-tested (BumpVersionMojoTest, 7/0/0); live-verified: mvn generate-resources left the counter unchanged, mvn package bumped 77→78 and stamped ac-cli 78, and target/classes/application.properties carried 78 (timing correct).