Files
agenticCode/x-docs/features.md
Ingo Schnabel 7e0d75cc5e Roadmap
2026-09-23 12:39:15 +02:00

435 KiB
Raw Blame History

AgenticCode — Implemented Features

Completed work, moved out of x-docs/roadmap.md (which now tracks only open items). Each entry records what was built; IDs are preserved from the roadmap (some IDs recur across sections — they are kept as-is for traceability).

The MCP surface no longer exists. It was removed on 2026-08-04 (roadmap item 26) after intermittent session failures that were not fixable from this codebase; all 40 tools had a REST twin, so no capability was lost. REST and the ac CLI are the only access paths.

Entries below still name MCP tools (module_payload, McpQueryTools, …). Those are kept as the historical record of what each feature shipped with — they are not a description of anything callable today. Where an entry reads as present tense, read "REST + CLI".

Metrics

  • 47. Generated vs. user-exit LoC/SLoC split (2026-07-13) — a project can declare a source language (required at creation; attribute only — ingest still classifies by extension) and a generatedDir/userExitDir pair (directory names, matched as path components like excludeDirs; both-or-neither, 400 LANGUAGE_REQUIRED/GENERATED_USEREXIT_PAIR otherwise). A generated module already contains its hand-written user-exit twin inline, so at ingest a module under generatedDir whose name also occurs under userExitDir is annotated with that twin's LoC/SLoC (userExitLoc/userExitSloc); user-exit files are not ingested as standalone modules (the walk skips userExitDir, avoiding name collisions). New UserExitMetrics.scan computes the twins with the same per-language LineCounter; ProjectIngestService.withUserExitMetrics stamps the generated nodes in both the coarse (Tier-1) and deep paths (correct at any depth). GET /loc (project_loc MCP, ac loc) now reports, per language row and in the total: loc/sloc (total, incl. exit), userExitLoc/userExitSloc, and generatedExclusiveLoc/generatedExclusiveSloc (= total − exit, clamped ≥0 per file). Fields added to ProjectInfo/ProjectRequest (REST + CLI project create/update with -l/-g/-u; project create has no MCP tool). Tests: SourceFilesTest, ProjectResourceIT (validation), UserExitLocIT (rollup split against the counter).

  • 46. Deterministic per-file LoC / SLoC, per language (2026-07-12) — every file-level node (MODULE, or DATA_STRUCTURE for a Natural .lda/.pda) is stamped at ingest with loc (physical lines) and sloc (source lines: non-blank, non-comment). SLOC is computed per language by a new LineCounter SPI in ac-parser-core (LocMetrics record + loc/sloc property keys): NaturalLineCounter follows natural-grammar.md §22.6.1 (full-line */**//*, inline /*); JavaLineCounter is a char-state scan handling //, /* … */ and string/char/text-block literals (so // inside a string stays code). Both the Tier-1 coarse scan and the Tier-2 deep parse call the same counter, so a module's metrics are identical at any ingest depth (numbers are reproducible and summable). Stamped in NaturalCoarseScanner/JavaCoarseScanner and, on the deep path, in ProjectIngestService.withShellMetrics (via AstIngestService.count). Surfaced on list_modules and module_context (new loc/sloc fields), on inspect_node (raw properties), and via a new rollup: GET /api/projects/{p}/loc?language=&sourceFile= (PROJECT_LOC cypher → GraphRepository.projectLoc → ProjectLoc DTO) — per-language fileCount/loc/sloc plus a project total, each source file counted once (collapses multi-node files via max per (language, sourceFile) before summing). Delivered as REST + MCP (project_loc) + CLI (ac loc) together. Tests: LocMetricsTest, NaturalLineCounterTest, JavaLineCounterTest, and AnalysisResourceIT (metrics match the counter; rollup totals/per-file consistent).

Natural ingest fidelity (WSUBPX0S findings)

Surfaced 2026-07-12 while re-analysing WSUBPX0S (project upms) via the API vs. the grep-based analyses. Two classes of information the API could not recover; both now fixed. Executable acceptance tests: ac-code-server/.../api/NaturalIncludeMacroAndXmlPayloadIT.java (the _target tests are green; fixtures under src/test/resources/fixtures/natural/framework-gaps/).

  • 44. Framework-mediated DB access via INCLUDE macros (2026-07-12) — table access performed through the generic table-access framework (INCLUDE YFRAMGC0 … '"<accessor>"' …) was invisible: the CALLNAT to the generic accessor lives in the copycode member, not the including module, so callees/db-accesses were empty. NaturalParser and NaturalCoarseScanner now recognise the statement-level framework macro (a new INCLUDE_MACRO matcher, gated on the declarative FrameworkMacros registry that maps a macro name → the positional index of the accessor argument), de-quote the accessor name ('"FRELEMG0"' → FRELEMG0, args collected across continuation lines), and emit a CALLS edge tagged callKind=INCLUDE_MACRO (new CallKind constant; surfaced as edgeKind on callers/callees). Because it's a normal CALLS, the accessor's table access surfaces transitively via the existing db-accesses?depth= query with via=accessor. Scope: targeted recogniser only (roadmap option 1); general .nsc copycode expansion and direct READS/WRITES mode tagging were deliberately left out (unnecessary for the acceptance criteria).

  • 45. XML payload / interface schema (2026-07-12) — XML wrapper subprograms map data-area fields to XML tags via the ADD-XML-LINE emit idiom (#W-TAG := '<tag>' / #W-VALUE := <field> / PERFORM ADD-XML-LINE, whose subroutine COMPRESSes '<' #W-TAG '>' #W-VALUE); that tag↔field↔direction contract was not captured. The deep parser now detects the emit subroutine and its tag/value variable pair, walks the emit sequences, and models each triple as a new PAYLOAD_FIELD node ({tag, field, direction=REQUEST}, qualifier stripped to the field name) contained by the module. Exposed as REST GET /modules/{name}/payload + MCP module_payload + CLI ac payload (PAYLOAD cypher → GraphRepository.payload → PayloadField DTO). Scope: the emit (REQUEST) idiom with literal tags; parse-side (RESPONSE) and derived-tag normalisation remain open.

  • 46a. General copycode (.cpy) expansion (2026-07-13) — statement-level INCLUDE <member> <args> is now expanded (deep + coarse) before parsing, so a copycode's CALLNAT/PERFORM, DB access and dataflow surface on the including module (previously invisible). CopycodePreprocessor splices the resolved copycode body with positional &1&… substitution and records a per-line origin table; after parsing, node/edge line numbers are remapped back to their real file positions — host statements to the host, copycode statements into the .cpy (tagged viaCopycode/includedAt) — so line-based navigation stays correct despite the splice. Copycodes are resolved by a per-ingest CopycodeLibrary (.cpy name→path, read + cached); they are not ingestible as standalone modules. Framework macros (item 44), data-area USING, unknown members and copycodes declaring DEFINE DATA are excluded; recursion is cycle-guarded. Threaded via a new CopycodeResolver into NaturalParser.parse/NaturalCoarseScanner.scan (Java unaffected). Acceptance: CopycodeExpansionIT (a host whose only DB access + call live in a .cpy).

  • 46b. XML payload derived from the interface PDA (2026-07-13) — real production XML wrappers (WNAUTD0S-style) are generic, runtime-driven serializers (YFRAMN07 tag-builder, ADD-XML-LINE/ ADD-XML-ACT) with no static tag list in source, so item-45's idiom scan finds nothing. Such a module is now flagged xmlWrapper with its PARAMETER USING interfacePda, and GET /payload derives the contract from that PDA when no static idiom exists: each field → a payload field, wire tag = field name normalised (#→_), direction REQUEST, source=PDA (vs source=IDIOM). Idiom fields take precedence. Acceptance: NaturalIncludeMacroAndXmlPayloadIT.genericWrapperPayloadIsDerivedFromInterfacePda. The static idiom additionally handles derived tags (EXAMINE #W-TAG FOR '#' REPLACE '_' → #→_) and both directions: ADD-XML-LINE-style emit sub → REQUEST, GET-XML-LINE-style parse sub (reverse field := #W-VALUE binding) → RESPONSE.

  • 46c. Ingest Global Data Areas (.gda) (2026-07-13) — .gda files now classify as Natural DATA_STRUCTUREs (like .lda/.pda) and DEFINE DATA GLOBAL USING <gda> resolves to them (the USING recogniser now accepts GLOBAL). Was: GDAs never ingested, GLOBAL USING unresolved.

Lazy / deferred three-tier ingest

Reworked ingest from eager whole-project parsing into a lazy, on-demand model. Three tiers: Tier 1 = cheap eager reference index (per file: nodes, identifiers, coarse call/DB references — no deep bodies); Tier 2 = lazy deep ingest (control flow, statement-level dataflow, precise reads/writes) triggered on demand; Tier 3 = source served from the filesystem, no longer stored on nodes. Reverse queries (callers, search_identifier, flow_backward) stay answerable because Tier 1 pre-indexes coarse references globally. (Item 43, automatic invalidation, remains open in the roadmap.)

  • 36. Tier 1 reference index + tri-state ingest status (2026-07-11) — on project create, scan every source file and create nodes (no deep edges) plus the identifier index and coarse call/DB references. Each node carries a tri-state status not-ingested / ingesting / ingested with a lock to serialise concurrent deep-ingest of the same node. Coarse references include Natural dynamic call targets (CALLNAT PGM-VAR), copycode/INCLUDE, continuation lines — computed by the lexer, not grep.

    • The Tier-1 coarse reference scan on project create: new CoarseScanner SPI (ac-parser-core); NaturalCoarseScanner is a lexer-level single pass (module /function/data-structure shells, declared-field identifier index, coarse CALLS/READS/WRITES/INCLUDES incl. PERFORM, CALLNAT '...', dynamic CALLNAT PGM-VAR, PARAMETER/LOCAL USING copybooks — attributed to the enclosing subroutine exactly as NaturalParser does, so coarse edges merge cleanly with a later deep ingest); JavaCoarseScanner reuses JavaParser but persists only the coarse projection (class/method shells + call/type edges), since a symbol-free Java scan would mis-resolve. Each shell carries a sourceHash (feeds item 41). AstIngestService.coarseScan + ProjectIngestService.scanTier1 walk the root, persist the coarse graph and run the cheap CALL_GRAPH enrichment (so cross-file callers/callees resolve immediately) — modules land CALL_GRAPH/NOT_INGESTED. Wired into POST /api/projects/{p} create (config agenticcode.tier1.scan-on-create, default true; non-fatal on an unscannable root). Covered by Tier1IndexIT, NaturalCoarseScannerTest, JavaCoarseScannerTest, SourceHashTest.
    • The durable tri-state status + lock: real MODULE nodes carry a durable ingestStatus (NOT_INGESTED/INGESTING/INGESTED, IngestStatus enum) with an ingestStatusAt stamp, kept separate from ingestDepth (enrichment tier). GraphRepository.claimIngesting/clearIngesting (CypherQueries.CLAIM_INGESTING/ CLEAR_INGESTING) implement a best-effort DB claim — Neo4j locks only at SET, so it is not a hard mutex; correctness rests on ingestModule being idempotent, with true intra-JVM exclusion still from the in-process monitor. DeepIngestCoordinator sets INGESTING at the start of a by-name deep ingest (markIngestDepth(FULL) flips the tree to INGESTED; failure rolls back to NOT_INGESTED), and a cross-process loser coalesces by polling moduleIngestState until the winner reaches FULL, re-claiming if the claim is released or goes stale (agenticcode.deep-ingest.ingesting-ttl-seconds, default 1800; poll bound claim-wait-seconds, default 120). Covered by DeepIngestStatusIT and CoalescingIT (concurrent deep ingests converge to a single FULL node).
  • 37. Tier 2 lazy deep-ingest on every call (2026-07-11) — every MCP/API call deep-ingests the nodes it touches, transitioning them to ingested. Decision: this includes fan-out result sets (e.g. search_identifier, callers), not just the explicitly named node — so the node budget (item 39) is the primary cost bound and its default must be a real, tuned number. Traversal queries (call_tree, flow_*) deep-ingest along the path as they walk it.

    • Module-gated field-level queries auto-trigger: DeepIngestCoordinator.ensureDeep runs a scoped ingestModule (with a per-(project,module) coalescing lock) from the withDeepModule choke point in both McpQueryTools and AnalysisResource; flow-forward/backward, field-flow etc. auto-deep-ingest instead of returning 409, falling back to the hint only when the module does not resolve.
    • Fan-out result-set warm for callers, search_identifier, and call_tree: blocking-then-rerun — each query runs against the graph as-is, DeepIngestCoordinator.ensureDeepMany deep-ingests the surfaced result-set files (by path, via the multi-root ProjectIngestService.ingestFiles) under a fan-out node budget (agenticcode.deep-ingest.fanout-nodes, default 50), and the query re-runs only if the warm deepened something. Wired at the withFanoutWarm choke point in both surfaces. Note: callers warms already-surfaced callers (downstream precision) but cannot reveal a caller invisible at the coarse tier.
    • Cross-module flow_* path-ingest (37a) for flow-forward, flow-backward, and field-flow: an ingest-and-re-traverse fixpoint at the withFlowPathWarm choke point — after ensureDeep(startModule) it runs the trace and, each round, deep-ingests the frontier (surfaced modules seeded with the start module, plus their direct callee modules) in one scope (DeepIngestCoordinator.ensureFlowFrontier → GraphRepository.flowFrontierSourceFiles → ProjectIngestService.ingestFiles), then re-traverses. Ingesting the whole frontier (callers included, even if FULL) re-links each caller into the freshly-ingested callee params, letting a trace cross into a dynamically-dispatched callee. Bounded by agenticcode.deep-ingest.flow-rounds (default 3) + the fanout-nodes budget; stops early at the fixpoint. Covered by AnalysisResourceIT.flowForwardPathWarmCrossesIntoDynamicallyDispatchedCallee.
    • Global concurrency cap: all deep-ingest/warm work in DeepIngestCoordinator (by-name ensureDeep, fan-out ensureDeepMany, flow ensureFlowFrontier) is gated by a fair Semaphore (agenticcode.deep-ingest.max-concurrent-warms, default 2), acquired via tryAcquire only around the actual ingest (never while a coalesce-waiter waits); on timeout (warm-acquire-timeout-seconds, default 10) the warm is skipped and the query returns its Tier-1 answer.
  • 38. Deep-ingest depth cap (2026-07-10) — the by-name deep-ingest BFS (ProjectIngestService.ingestModule) is bounded by maxDepth hops from the named module: default 5 (agenticcode.deep-ingest.default-depth), clamped to the ceiling 20 (agenticcode.deep-ingest.max-depth). Hitting the bound truncates + reports (not a hard error) — the reached modules are marked FULL and the response carries an IngestSummary.Truncation. Override via ?maxDepth= (REST, now on refresh/{name}), the refresh MCP tool's maxDepth arg, and ac refresh --max-depth (CLI).

  • 39. Deep-ingest node budget (2026-07-10) — alongside the depth cap, the BFS ingests at most maxNodes files (default 300, agenticcode.deep-ingest.default-nodes); overflow truncates + reports the same way (reason NODES/NODES_AND_DEPTH). Override via ?maxNodes= (REST), the refresh MCP tool's maxNodes arg, and ac refresh --max-nodes (CLI). Defaults are @ConfigProperty and should still be tuned against the real ac project.

  • 40. Unresolved-reference nodes (2026-07-12) — when Tier 1 sees a variable / dynamic call target it cannot resolve, store it as a deduped placeholder node/edge flagged unresolved, so reverse queries (e.g. callers) surface dynamic call sites while still distinguishing them from resolved edges. Resolution pass runs during Tier 2.

    • Placeholder nodes (sourceFile = "") were already deduped (MERGE on type,name,sourceFile,project); this adds the unresolved flag. New enrichment step STAMP_UNRESOLVED_PLACEHOLDERS (runs last in finalizeProject/finalizeProjectScoped, project-wide, cheap, idempotent, every mode) sets unresolved = (no real definition of that (type,name) is ingested) — the resolution pass: a reference reads unresolved=false once its target file is ingested (a real twin exists), true while genuinely dangling. Surfaced explicitly on search_identifier (IdentifierMatch.unresolved) and, for free, on inspect_node (nodes/{id} returns the whole node). Reverse call queries (callers/callees) already surface these targets as entries with a blank sourceFile. Covered by UnresolvedRefIT (dangling → true; ingest target → false).
  • 41. Tier 3: stop storing source on nodes (2026-07-12) — drop stored source text; serve module_source/node_source by opening sourceFile (project-relative path + content hash) and slicing startLine..endLine. Stale check: on read, compare the stored content hash against the current file; on mismatch, return a structured STALE_SOURCE error ("run refresh") instead of slicing — never serve current lines against old line numbers.

    • Source text was already not stored on nodes (served off disk by SourceSnippetService since item 28); this adds the mandatory stale check. Shared SourceHash (SHA-256) is stamped on each file's MODULE/DATA_STRUCTURE shell at ingest — by the Tier-1 scanners and (for robustness) the deep full-parse paths (ProjectIngestService.withSourceHash), surviving re-ingest via SET +=. On a source read, GraphRepository.sourceHash(project, sourceFile) supplies the stored hash and SourceSnippetService.read recomputes the current file's hash; on mismatch it throws StaleSourceException, which both REST (409 STALE_SOURCE) and MCP (STALE_SOURCE tool error) return with a "re-ingest (refresh)" message. A legacy file with no stored hash skips the check (serves unchecked). Covered by StaleSourceIT (fresh→200, edit→409, re-ingest→200). Copycode/INCLUDE slices remain raw pre-expansion file text. Note: a fan-out warm query (e.g. search_identifier) re-ingests and re-hashes a surfaced module, clearing its staleness as a side effect.
  • 42. refresh tool (manual invalidation) + remove ingest_all / ingest_module (2026-07-12) — remove the eager bulk/one-module ingest tools across all three surfaces (MCP McpIngestTools, REST AnalysisResource, ac CLI) and replace with a refresh operation that re-runs Tier 1 + deep ingest. Interim manual invalidation until item 43.

    • refresh is the single (re-)ingest surface across REST + MCP + CLI; ingest_all/ingest_module/ingest-call-graph were removed from all three. REST POST /refresh (whole root, call-graph tier; ?deep=true for field-level) and POST /refresh/{name}?maxDepth=&maxNodes= (deep re-ingest a module tree); MCP one refresh(project, module?, deep?, maxDepth?, maxNodes?) tool; CLI ac refresh [name] [--deep] [--max-depth --max-nodes]. ProjectIngestService.refreshProject/refreshModule delegate to the existing (now internal) ingest primitives — a whole-root refresh runs the full-parse call-graph pass (a superset of the create-time coarse Tier-1 scan, so intra-module dynamic CALLNAT still resolves), and is the remedy for a STALE_SOURCE read (re-stamps sourceHash). It MERGEs current files (does not wipe; deletions await item 43). All ~20 ITs bootstrapping via the old endpoints were migrated (behavior- preserving URL swaps); CLAUDE.md + mcp-api-usage dogfooding steps now say refresh. Covered by RefreshIT (whole-project + module refresh work; removed endpoints 404).
  • 43. Automatic invalidation (hash-based) + deleted-file sweep (2026-07-15) — the graph now self-heals when source files change or are deleted on disk, without a manual refresh. Chosen approach: hash-based, lazy (on-access) — not a filesystem watcher (no background thread; reuses the item-41 sourceHash). Two parts, gated by agenticcode.auto-invalidate.enabled (default true).

    • Part A — change detection → auto re-ingest. DeepIngestCoordinator.ensureDeep/ensureDeepMany (the two deep-ingest choke points every field-level query routes through) no longer short-circuit a FULL module blindly: they compare the file's current SourceHash to the stored one (same compare as the /source stale check). A FULL-but-changed module is invalidated (GraphRepository.invalidateModule → shell reset to CALL_GRAPH/NOT_INGESTED, so isFull() is false and the coalesce loop re-ingests) and re-parsed — with the item-58 reconciliation purging its stale nodes. Conservative false (never thrash a FULL module) when the flag is off, no source file, no stored hash (legacy node), or the file is unreadable/missing (a deleted file is Part B's job). Per-module granularity: only the queried module's own file is checked, not its whole dependency tree (a changed dependency is caught when it is queried by name).
    • Part B — deleted-file sweep. A whole-project refresh (ProjectIngestService.refreshProject) now reconciles against the filesystem: GraphRepository.distinctSourceFiles → any real sourceFile no longer present on disk → deleteNodesForSourceFiles (DETACH DELETE). Closes the item-58 gap (a refresh MERGEd current files but left nodes for deleted files). Reconciles against disk existence, not the walked-candidate list, so on-demand dependency files (PDAs/LDAs/.cpy) that exist but aren't top-level candidates are never wrongly swept.
    • No new REST/MCP/CLI surface — Part A is internal to the query path; Part B changes refresh semantics (now also drops deleted-file orphans). Both are documented behaviour changes, so the three surfaces stay in sync by construction. Covered by AutoInvalidationIT (edit a FULL module on disk → a field query auto-re-ingests, 409 STALE_SOURCE → 200, no manual refresh) and DeletedFileSweepIT (delete a file → refresh removes its module + field nodes, survivors kept). PlaceholderResolveNullLineIT (which seeds phantom nodes for files it never writes) opts out via a @TestProfile disabling the sweep.

Layered / performance-aware ingest

  • P2-c. ingest-all defers field resolution AND dataflow by default (fast whole-root) (2026-06-21) — even with P2-b's scoped per-program resolution, a whole-root ingest-all still ran the global field-placeholder resolution AND the argument→parameter dataflow — measured live on upms: the field WRITES pass took ~26 min and link-args-to-params ~25 min (each one transaction over ~100k+ edges). So /ingest-all now defaults to the fast call-graph pass (EnrichmentLevel.CALL_GRAPH): placeholder resolution (CALLS/INCLUDES/USES_TYPE/EXTENDS/IMPLEMENTS) + intra-module dynamic CALLNAT + polymorphic (CHA) fan-out. It defers the expensive field-placeholder resolution and the dataflow + cross-module dynamic CALLNAT dispatch to a scoped per-program deep ingest (/ingest/{name}, P2-b) — which fans out top-down from the analyzed program (the user's insight: dynamic dispatch is a top-down concern, not needed globally bottom-up). The old full whole-root behaviour is available via POST /ingest-all?deep=true. Implementation: EnrichmentLevel { CALL_GRAPH, FULL } (each carrying dataflow/resolveFields gates), finalizeProject(project, EnrichmentLevel), ingestAll(project, deep), AnalysisResource.ingestAll gains ?deep. The default tags modules CALL_GRAPH, so field-level endpoints return the 409 NOT_DEEPLY_INGESTED deep-ingest hint (P2-a). Covered by AnalysisResourceIT (ingestAllDefaultDefersDataflowAndFieldResolutionToDeepIngest: fast ingest → intra-dynamic resolved, cross-dynamic absent + 409 on field-flow → /ingest/{name} → cross-dynamic resolved

    • 200); the IT's ingestAll() helper uses ?deep=true so existing field/dataflow assertions still exercise the full path. An earlier attempt kept dataflow+cross-callnat in a CALL_GRAPH_DATAFLOW middle tier; live upms testing showed link-args-to-params alone was ~25 min, so it was dropped.
  • P2-b. Scope deep field resolution to the program tree (2026-06-20) — in the accumulating one-project model (call-graph everything, then deep-ingest programs over time), a by-name deep ingest previously ran the unscoped field-placeholder resolution over the whole project (~110k edges → ~28 min once the full call graph is loaded). Added $names-scoped variants of the field-resolution and dataflow queries (resolvePlaceholderFieldTargetsScoped, …ByNameScoped, resolveBareIncludedFieldTargetsScoped, LINK_ARGS_TO_PARAMS_SCOPED, …_JAVA_SCOPED) that start from the just-ingested program tree's modules instead of scanning all placeholder edges. GraphRepository.finalizeProjectScoped(project, moduleNames) runs them (call-graph resolution, placeholder cleanup, CHA fan-out stay project-wide/idempotent); ingestModule now calls it with the BFS-collected module names. So deep ingest stays fast (≈ the small-project 8.5 s case) even when the project already holds the whole call graph. Correctness covered by IngestModuleIT field-flow/dataflow tests (now exercising the scoped path); all ingest/analysis ITs green. Live perf re-validated on upms (2026-06-21): whole-codebase ingest-call-graph = 6309 modules in 67 s (placeholders left intact), then a by-name deep ingest of KDWWIFN0 (56-module tree) with the full call graph already loaded ran in 5.2 s — confirming the scoped resolution avoids the old unscoped whole-graph path (~28 min). Depth guard verified end-to-end: flow-forward returns 200 for the FULL module, 409 NOT_DEEPLY_INGESTED for a call-graph-only module (ACCNPE01), and 409 NOT_INGESTED for an un-ingested name (JE999), each with an actionable nextAction.

  • P2-a. Layered ingest: call-graph mode + ingest-depth guard (2026-06-20) — field-placeholder resolution is volume-bound (whole-codebase ≈110k edges → not viable even cycle-safe/APOC node-global; measured), while per-program deep ingest is fast (~8.5 s for KDWWIFN0's 56-file tree). So ingest is now layered: (1) POST /api/projects/{p}/ingest-call-graph — whole-root, persists everything but runs only the cheap call-graph/module-level enrichment (CALLS/INCLUDES/USES_TYPE/EXTENDS/IMPLEMENTS + CHA fan-out), skips field-placeholder resolution, and leaves placeholders intact for later deep ingest. Tags modules ingestDepth=CALL_GRAPH. (2) POST /ingest/{name} (by-name) and ingest-all run full enrichment and tag the ingested modules ingestDepth=FULL (accumulates over time in one project). (3) Field-level endpoints (flow-forward/flow-backward/field-flow) check the target module via GraphRepository.moduleIngestState and, when not FULL, return an agent-actionable 409 { status: NOT_DEEPLY_INGESTED|NOT_INGESTED, module, detail, nextAction:{method,path} } instead of a misleading empty result. New IngestDepth/ModuleIngestState, finalizeProject(full) splitting enrichmentSteps(full) into call-graph vs field groups, depth-marking Cypher, per-step finalize logging (label/duration/rows/JVM heap). New IT callGraphIngestEnablesCallGraphAndGuardsFieldFlow (409→200 after deep ingest); full AnalysisResourceIT+ingest ITs green.

Parser / model

  • 1. Persisted per-language source/module kind (done 2026-07-07, scoped down after investigation) — add an AstNode property capturing each module's finer kind beyond the coarse NodeType. Delivered: Java moduleKind = CLASS/INTERFACE (exact, from type.isInterface()); Natural moduleKind = a best-effort PROGRAM/SUBPROGRAM guess (has a top-level PARAMETER section → SUBPROGRAM). GET /modules gained a ?moduleKind= filter. Not attempted (found to be a separate, larger feature, or genuinely unrecoverable from source text alone): Java enum/record — JavaParser doesn't parse these as MODULE nodes at all today (only ClassOrInterfaceDeclaration), so this needs a new node-family traversal, not just a property; Natural COPYCODE/MAP/GDA — per natural-grammar.md §22.2, PROGRAM/SUBPROGRAM/SUBROUTINE are real SAG Natural catalog metadata, and main-program/external-subprogram/ subprogram are syntactically identical productions — no reliable source-level signal exists for these without the original catalog. See x-docs/agent-api-usage-ac-implementation.md step 0 for the heuristic's limits.

Foundation (schema, health, errors, ingest CLI)

  • 1. Neo4j indexes & constraints — index (:AstNode) (project, name) and (project, type, name), unique constraint on (:Project) (name). Created on application startup via SchemaInitializer (idempotent IF NOT EXISTS).
  • 2. Health checks — Neo4jHealthCheck (@Readiness) verifies Neo4j connectivity via Driver.verifyConnectivity().
  • 3. Structured error responses for AnalysisResource — shared ErrorResponse record ({error, code, details}), reused by ProjectResource.
  • 4. Project-existence validation — ingest and all project-scoped query endpoints return 404 PROJECT_NOT_FOUND if the project hasn't been created.
  • P0-a. Uniqueness constraint/index on AstNode.id — mergeEdge matches both endpoints by id (MATCH (a:AstNode {id: $sourceId}), (b:AstNode {id: $targetId})), but id was not indexed (only (project, name) and (project, type, name) were). Every edge merge did a full label scan over all AstNode nodes accumulated so far, so ingest got progressively slower as more files/projects were ingested. Added CREATE CONSTRAINT ast_node_id_unique IF NOT EXISTS FOR (n:AstNode) REQUIRE n.id IS UNIQUE to SCHEMA_STATEMENTS, giving edge merges O(1) lookups regardless of graph size.
  • P0-b. Batch ingest writes with UNWIND — GraphRepository.save() previously issued one tx.run() per node and per edge (N+M round trips per ingested file). Replaced with CypherQueries.MERGE_NODES (UNWIND $nodes AS n MERGE ...), one round trip for all nodes of a file, and CypherQueries.mergeEdgesBatch(EdgeType) (UNWIND $edges AS e MATCH ... MERGE (a)-[:<TYPE>]->(b)), one round trip per distinct EdgeType present in the file (edges grouped via Collectors.groupingBy(AstEdge::type)) — Cypher relationship types can't be parameterized, so true O(1) isn't possible, but this cuts round trips from N+M to 1 + (number of distinct edge types, usually 1-3). Same merge semantics, low risk.
  • P0-c. --exclude-dir option for ingest CLI command — recursive ingest now skips any file whose path (relative to the ingest root) contains a component matching one of the repeatable -x/--exclude-dir <name> values (case-insensitive), e.g. --exclude-dir user_exit --exclude-dir test to skip subprograms already represented in generated_sources. Excluded files are printed as SKIP (excluded) <file> and counted in a new excluded summary counter alongside ingested/skipped/failed.
  • P0-d. ingest-module <folder> <moduleName> — dependency-driven ingest (2026-06-15) — new CLI command that ingests a single module and its transitive CALLNAT/PERFORM-external/EXTENDS/IMPLEMENTS targets and INCLUDE/USING'd data areas, located by filename stem (case-insensitive) within a given folder. POST .../ingest/{java,natural} now returns 202 {"dependencies": [{"name":..., "type": "MODULE"|"DATA_STRUCTURE"}]} (new AstIngestService/AnalysisResource.IngestResponse/DependencyRef), derived from placeholder nodes (sourceFile="") created by the parsers. The CLI does a BFS over these names: found files are ingested and their dependencies enqueued; names not found in the folder are printed as UNRESOLVED <name> and counted separately (not a failure). --exclude-dir applies to the file index, so dependencies under excluded folders are reported as unresolved. Shared file-walk helpers (readSourceFile, endpointFor, isExcluded) extracted from IngestCommand into IngestSupport. New IT naturalIngestReturnsUnresolvedDependencies.
  • P0-e. .lda/.pda (Natural data area) parsing (2026-06-15) — NaturalParser now recognizes .lda/.pda files (via IngestSupport.endpointFor) and parses their wrapper-less field-export format (<TYPE><LENGTH><LEVELDIGIT><NAME> ... CONST<...>|INIT<...>, glued type/level/name prefixes, R <level><name> REDEFINE markers, *-comment lines, multiple top-level (level 1) DATA_STRUCTURE roots per file) into the same DATA_STRUCTURE/VARIABLE/CONSTANT node shape as DEFINE DATA fields, so data-structure-fields works unchanged. Combined with item 6 (placeholder resolution), an ingested LDA/PDA now links up with the DATA_STRUCTURE placeholder created by LOCAL/PARAMETER USING in referencing .nat programs. New unit tests (YFRAML01_SAMPLE.lda, WGEAGL01_SAMPLE.pda) and IT (ingestingLdaResolvesIncludePlaceholder).
  • P0-f. GET /variables/{name}/writes and /reads (2026-06-15) — queries returning every FUNCTION/MODULE that has a WRITES/READS edge to a VARIABLE/CONSTANT with the given name, with sourceFile and containing module (CypherQueries.variableReads(maxDepth)/ variableWrites(maxDepth), GraphRepository.variableWrites/variableReads, VariableAccessLocation record). Optional ?module=...&depth=N query params restrict results to that module or anything in its transitive CALLS tree (up to depth hops, same agenticcode.call-tree.default-depth/max-depth config and clamping as call-tree); the Cypher computes each writer/reader's containing module via CONTAINS*0.. and checks owner.name = $module OR EXISTS { (module)-[:CALLS*1..depth]->(owner) }. New CLI variable-writes/variable-reads <name> [-m/--module <name>] [--depth N].

Ingest performance

  • 24. Index the stale-file sweep — persist was a label scan per file (2026-07-16) — the sibling defect to P1-m below, one statement further on: P1-m indexed the node MERGE key, but DELETE_STALE_FILE_NODES (the item-58 reconcile sweep, which runs in the same mergeResults transaction) matches on (project, sourceFile) — and a Neo4j composite index only applies when every one of its properties is constrained, so the 4-property (project, sourceFile, type, name) index does not cover that 2-property prefix. The planner therefore fell back to NodeByLabelScan + Filter, and because the UNWIND $files batch drives it through an Apply, it scanned once per file: measured 564,451 db hits per file on a 282k-node graph, i.e. ~56M node reads per 200-file batch. The scan covers the whole AstNode label — all projects — which is why per-batch cost grew with total graph size (2m29s early vs 2m48s late). Fix is one line in SCHEMA_STATEMENTS: ast_node_project_sourcefile on (project, sourceFile). The plan becomes a NodeIndexSeek at 137 db hits per file (~4000×). Verified that the new, less selective index does not displace the 4-property one for the MERGE keys (still 1–2 db hits). Also serves SOURCE_HASH (item 43 auto-invalidation) and the sourceFile:"" placeholder sweeps, which had the same blocked shape. Index-only; ensureSchema adds it IF NOT EXISTS, so existing deployments pick it up on the next restart — no migration. Measured live on a full upms call-graph refresh (6311 files, reconcile on): whole refresh 179 s end-to-end — parse ~63 s, persist 104.5 s across 32 batches (median 1.7 s/batch, min 0.2 s, max 10.7 s), finalize ~12 s — against ~2.5–2.8 min per batch logged before the fix, and with no per-batch growth left. A second refresh took 95 s and reproduced the graph exactly (454,255 nodes / 1,495,146 rels, no duplicate MERGE keys), i.e. idempotent. Note the pre-fix graph held only 216,492 upms nodes: the slow run evidently never completed, so the fix also produced the first complete upms graph rather than merely a faster one. Regression test StaleFileSweepIndexIT asserts the cause — it EXPLAINs the real DELETE_STALE_FILE_NODES constant and requires a NodeIndexSeek on AstNode(project, sourceFile) (the planner names indexes by properties, not by index name), after CALL db.awaitIndexes() since a POPULATING index is invisible to the planner and would make the assertion race startup. Vacuity-checked: with the schema line removed the plan falls back to the filtered label scan and the test fails.
  • 25. ingest-all performance: batch persist transactions — already implemented (closed 2026-07-16) — the roadmap item claimed "persist still opens one transaction/session per file"; that was stale. ProjectIngestService chunks files into persistBatch calls sized by agenticcode.ingest.batch-size (default 200), and GraphRepository.mergeResults aggregates each chunk into batched UNWIND MERGEs (one per node class, one per edge type) in a single transaction — delivered earlier as P1-k below. Closed as done; it was item 24 above, not batching, that actually held persist back.
  • P1-m. Index the node MERGE key to kill O(N²) persist (2026-06-20) — on a whole-root ingest of upms (~6309 files), per-batch persist time grew with the graph (batch 1 ≈ 1 s, batch 8 ≈ 4.5 min) — quadratic. MERGE_NODES/ MERGE_POSITIONAL_NODES match on (type, name, sourceFile, project) but the only supporting index was (project, type, name); for low-selectivity names that recur project-wide (CONTROL_FLOW IF/FOR/REPEAT/DECIDE, VARIABLE/DATA_STRUCTURE FILLER/#CODE/SQLERR) that bucket grows with the graph and MERGE filtered sourceFile/startLine in memory over it. Added composite index ast_node_project_sourcefile_type_name on (project, sourceFile, type, name) so MERGE seeks per-file (a handful of nodes). Index-only; ensureSchema adds it IF NOT EXISTS.
  • P1-l. Split enrichment into per-statement transactions (2026-06-20) — finalizeProject ran all ~15 project-wide enrichment statements in one transaction; on the full upms graph that single transaction blew Neo4j's dbms.memory.transaction.total.max (5.4 GiB) and aborted, so placeholders never resolved. finalizeProject now runs each ordered, idempotent statement (enrichmentStatements()) in its own executeWriteWithoutResult, bounding peak per-transaction memory. Same graph result (ITs green); a lone over-large step would still need CALL { … } IN TRANSACTIONS (not yet required).
  • P1-k. Batch persistence in ingest all (2026-06-19) — ingestAll persisted one file per transaction (fresh session + ~8 statements + commit), sequentially — fine for a small slice but slow once the P1-j dedup fix made it ingest thousands of files (each round trip is fixed overhead × N files). New GraphRepository.persistBatch / AstIngestService.persistBatch persist a chunk of files in one transaction, aggregating all files' nodes/edges into batched UNWIND writes (one statement per node label / edge type per chunk instead of per file). ingestAll now chunks the non-conflicting results by agenticcode.ingest.batch-size (default 200) and runs one finalizeProject after. Correctness subtlety: different files emit the same placeholder (CALLNAT/USING target, sourceFile="") with distinct UUIDs but one MERGE key; persisting all nodes before all edges collapses them (last id wins) and would orphan earlier files' edges. mergeResults dedupes nodes by MERGE key (canonical = first id) and remaps every edge endpoint onto the canonical id — the per-file transactions previously masked this. Caught by AnalysisResourceIT (interface fan-out + shared-PDA field-flow) during implementation; full ingest/analysis ITs green.
  • P1-j. Fix over-aggressive duplicate detection in ingest all (2026-06-19) — ProjectIngestService.ingestAll flagged cross-file duplicates by scanning every real MODULE/DATA_STRUCTURE node, so two unrelated Natural programs that each contain an inline DEFINE DATA group with a common name (FILLER, SQLERR, #ERROR-GROUP, …) were treated as duplicates and both whole files skipped. On the upms codebase this silently dropped 4379 files (3894 of 3895 "duplicate" identities were inline DATA_STRUCTURE names; only 1 was a real MODULE), so e.g. W-MNT-N0/KDWWIFN0 never ingested and their callers/callees couldn't resolve. New fileIdentities(Parsed) keys duplicate detection on the file's own identity only: real MODULE nodes (Java FQN-aware) for program/class files, and a single (DATA_STRUCTURE, stem) for .lda/.pda data-area files — inline group nodes are excluded. Real MODULE and data-area-file duplicates are still reported. In-memory only (no schema change/re-ingest of node shape). New IT sharedInlineDataStructureNamesDoNotSkipModules; full AnalysisResourceIT + IngestModuleIT green. A persisted per-language source/module kind is tracked separately as future work.
  • P1-i. Defer enrichment in by-name ingest (2026-06-19) — POST /ingest/{name} (ProjectIngestService.ingestModule, the CLI ingest <module> path) called save() per file in its dependency BFS, which re-ran the full ~15-statement project-wide enrichment (placeholder resolution + dataflow + polymorphic fan-out) after every file — ≈O(N²) over a growing graph, the dominant cost for Natural modules that fan out through CALLNAT/PERFORM/ USING (observed on ingest KDWWIFN0). Now mirrors ingestAll: persist() per file (merge only), then one finalizeProject() after the BFS. Enrichment is idempotent/order-independent, so the end graph is identical (verified by IngestModuleIT, which exercises cross-module dataflow that depends on enrichment). Removed the now-unused AstIngestService.save() / GraphRepository.save().
  • P1-h. Run DB endpoints off the event loop (2026-06-19) — the Neo4j blocking driver is wrapped in Uni.createFrom().item(...) in every GraphRepository method, whose supplier runs on the subscribing thread. The Uni-returning JAX-RS query methods carried no @Blocking, so for those the driver ran on (and blocked) the Vert.x event loop — latent until a slow query tripped the 2 s BlockedThreadChecker (observed on DELETE /api/projects = CLEAR_ALL full-graph DETACH DELETE, blocked ~3.3 s). Fixed by moving @Blocking to class level on AnalysisResource and ProjectResource so all endpoints dispatch to a worker thread (the two now-redundant method-level @Blocking on the ingest endpoints were removed). Verified via access log (executor-thread-* instead of vert.x-eventloop-thread-*); ProjectResourceIT + AnalysisResourceIT green.
  • 23. ingest-all performance: enrich once, not per file (2026-06-19) — ingest-all ran the project-wide enrichment block (placeholder resolution for all RESOLVABLE_EDGE_TYPES, field-target resolution, bare-field redirection, placeholder deletes, LINK_ARGS_TO_PARAMS/_JAVA, LINK_CALLS_TO_IMPLEMENTATIONS) inside every GraphRepository.save, i.e. once per file — each pass scans the whole project graph, so cost grew ~O(N²) in the file count and dominated large Java scans. Split persistence from enrichment: GraphRepository refactored into mergeResult/runEnrichment helpers, exposing persist(project, result) (MERGE only) and finalizeProject(project) (enrichment once), with save retained as merge + enrich in one tx for single-file/by-name ingest. ProjectIngestService.ingestAll now persists each survivor then calls finalizeProject once (skipped when nothing persisted). The enrichment is idempotent and order-independent, so the final graph is identical; verified by the existing AnalysisResourceIT (ingest-all-based) suite. ingestModule unchanged (still save per file).

Call graph, edges & provenance

  • P1-g. Language-aware edgeKind (2026-06-19) — edgeKind in callers/callees was always derived from the target node type as Natural's CALLNAT/PERFORM, mislabelling Java calls (e.g. a new Foo() constructor showed as CALLNAT). Each parser now stamps a callKind property on the CALLS edge using the new CallKind enum (ac-parser-core): Natural → CALLNAT/PERFORM; Java → METHOD_CALL (intra- and cross-class invocations) / CONSTRUCTOR (new). Stamping at parse time lets Java distinguish constructors from method calls, which the previous target-type-only logic could not. CypherQueries.callers/callees now return coalesce(r.callKind, <legacy CASE>) so graphs ingested before the stamp fall back to the old labels (no forced re-ingest; re-ingest needed for Java precision). Synthetic inheritance edges (LINK_CALLS_TO_IMPLEMENTATIONS) copy s.callKind = r.callKind. New IT assertions for Java (METHOD_CALL/CONSTRUCTOR) and Natural (PERFORM/CALLNAT).
  • 6. Cross-file placeholder resolution (2026-06-15) — decided against a generic EnrichmentPipeline.run() (no second use case identified yet; enrichment/ stays empty). Instead, GraphRepository.save() runs CypherQueries.resolvePlaceholderTargets(EdgeType) for each of CALLS/INCLUDES/USES_TYPE/EXTENDS/IMPLEMENTS, then DELETE_RESOLVED_PLACEHOLDERS, in the same transaction as the node/edge merges. This redirects edges pointing at an unresolved placeholder (MODULE/DATA_STRUCTURE, sourceFile="") onto a real node sharing (type, name, project) once one exists, and deletes the now-orphaned placeholder — works regardless of ingest order.

Natural parser — statements, dataflow, dispatch

  • P1-a. SQL/ADABAS statement extraction for agentic Panache generation — new NodeType.DB_ACCESS node per SQL/ADABAS statement occurrence (table, access mode, raw statement text, line span), CONTAINS edge from the containing function/module, USES_TYPE edge to the accessed DB_TABLE and (for SELECT ... INTO VIEW) the associated DATA_STRUCTURE. New GET /api/projects/{project}/modules/{name}/sql-statements endpoint + CLI command, so an agent can read table/view/mode/statement text and the call-graph context and write the equivalent Panache query itself. Existing function -> DB_TABLE READS/WRITES edges and /db-accesses stay unchanged.
  • P1-b. MOVE/ASSIGN/COMPUTE variable read/write edges — Natural construct #7 (CLAUDE.md priority list). Declared VARIABLE/CONSTANT names from DEFINE DATA are tracked case-insensitively; for MOVE <src> TO <target>{...} and (COMPUTE|ASSIGN) <target> (=|:=) <expr>, READS/WRITES edges are added from the current function to known variables/constants referenced as source/target/expression operands. Exotic MOVE variants (BY NAME/BY POSITION, EDITED, SUBSTRING(...), ALL, ENCODED, NORMALIZED, *JUSTIFIED) are skipped. No new placeholder nodes for unknown identifiers/literals.
  • P1-c. Bare := assignment and qualified-target field lookup (2026-06-15) — NaturalParser previously required a COMPUTE/ASSIGN keyword to recognize an assignment; Natural's implicit-assignment form <target> := <expr> (no keyword) is now matched via a new BARE_ASSIGN pattern (fallback after ASSIGN_COMPUTE, restricted to := to avoid colliding with IF/comparison =). lookupVariable now also strips a <QUALIFIER>. prefix (e.g. CDBRPDA.SORT-KEY -> SORT-KEY) when the fully qualified name isn't a known local variable, so qualified field references resolve to the declared field name. New unit tests bareAssignWithoutKeywordIsRecognized / qualifiedAssignTargetMatchesUnqualifiedDeclaredField. New IngestModuleIT ingests the WGEAGB0S fixture tree (100 files).
  • P1-d. ADD/SUBTRACT/MULTIPLY/DIVIDE read/write edges (2026-06-15) — Natural's arithmetic statements read and write their operands but were not recognized at all by NaturalParser. New patterns ADD_STATEMENT, SUBTRACT_STATEMENT, MULTIPLY_STATEMENT, DIVIDE_STATEMENT add READS edges for every identifier in the operand-list/expression operands (via new addExpressionReads helper) and READS/WRITES edges for the accumulator/target (addOperandAccess helper): ADD ... TO x → reads+writes x; ADD ... GIVING x → writes only x; SUBTRACT ... FROM x [GIVING y]; MULTIPLY x BY y [GIVING z]; DIVIDE x INTO y [GIVING z] [REMAINDER r]. New unit tests covering all four statements and their GIVING/no-GIVING variants.
  • P1-e. lineNo on AstEdge (2026-06-15) — AstEdge gained an int lineNo field (the source line of the relationship/statement), threaded through both parsers' edge() helpers and all call sites. Persisted as a relationship property: buildMergeEdgeQueries() now does MERGE (a)-[:%s {lineNo: e.lineNo}]->(b) (so the same edge type between the same nodes at different lines becomes distinct relationships), and buildResolvePlaceholderTargetQueries() preserves r.lineNo when redirecting placeholder edges. Exposed via lineNo in VariableAccessLocation and CallReference (callers/callees now return one row per call site).
  • P1-f. DECIDE FOR/DECIDE ON as CONTROL_FLOW (2026-06-15) — per the construct-coverage review, DECIDE FOR/DECIDE ON (Natural's switch/case) appears in 8 of 26 WGEAGB0S fixtures but was not recognized. New DECIDE_STATEMENT/END_DECIDE patterns add a CONTROL_FLOW node (dataType="DECIDE", value = the full statement text) spanning to END-DECIDE, mirroring IF/FOR/REPEAT. Statements inside the block still get normal READS/WRITES edges. New unit test decideForIsRecognizedAsControlFlowBlock.
  • P1-g. Resolve qualified PARAMETER/LOCAL USING field references (2026-06-15) — MOVE/ASSIGN/etc. targets of the form STRUCT.FIELD (e.g. CDBRPDA.SORT-KEY) where STRUCT is a PARAMETER USING/LOCAL USING include now resolve to a placeholder VARIABLE node FIELD, CONTAINS-child of the placeholder STRUCT DATA_STRUCTURE. Previously such references either fell back to an unrelated same-named local variable or produced no edge at all. New Neo4j post-ingest step (resolvePlaceholderFieldTargets for READS/WRITES, DELETE_RESOLVED_PLACEHOLDER_FIELD_CONTAINS) redirects these placeholder fields onto the real field of the same name in the resolved structure. New unit test qualifiedFieldOfIncludedDataAreaResolvesToPlaceholderUnderThatStructure. Known limitation: two different included structures defining a same-named field converge (merge key has no parent reference); not an issue in current fixtures.
  • 18. Cross-file bare-field resolution (MOVE/ASSIGN/etc.) (2026-06-17) — closes the remaining half of the P1-c/P1-g gap: an unqualified reference (MOVE #X TO SORT-KEY where SORT-KEY is declared only in a LOCAL/PARAMETER USING area). lookupVariable gained an allowIncludePlaceholder flag (true only at explicit MOVE/ASSIGN/COMPUTE/ arithmetic operand sites). When set and the bare name is not a same-file variable, is identifier-shaped, and the module has USING includes, a module-level placeholder VARIABLE (sourceFile="") is created. New resolveBareIncludedFieldTargets redirects it onto a real included field only on a unique single match across m-[:INCLUDES]->(:DATA_STRUCTURE{sourceFile<>""})- [:CONTAINS*1..]->; orphan cleanup reaps unresolved placeholders so no spurious edge survives. New NaturalParserTest cases + IT bareReferenceToIncludedFieldResolvesToRealField.
  • 19. Field-level dataflow for shared PDAs (field-flow) (2026-06-17) — a shared PDA's fields are single nodes, so a field written in the caller and read in a transitively-called module is already connected through that one field node plus the CALLS edge — no new edge type needed. New GET /api/projects/{project}/variables/{field}/field-flow?module=X&depth=N correlates writers and downstream readers of a shared field across the call graph (wmod <> rmod, reachable via CALLS*1..depth, clamped 10), returning "field F is produced in MOD-A:120 and consumed in downstream MOD-B:45". New FieldFlow record + CLI field-flow. Caveat: reachability-based, not order-precise. Fixtures FF_SHARED.lda + FF_PROD.nat
    • FF_CONS.nat; IT fieldFlowTracesSharedPdaFieldAcrossCall.
  • 15. Dataflow analysis through the call hierarchy (2026-06-16, scoped) — captures which variables are passed at each call site. Parser (Phase 1): CALLNAT 'MOD' ARG1 ARG2 captures the positional argument list onto the CALLS edge args property; top-level PARAMETER fields tagged with paramPosition. Enricher (Phase 2): LINK_ARGS_TO_PARAMS runs after placeholder resolution, positionally joining split(r.args,',')[i] to the callee param with paramPosition = toString(i) and MERGE-ing an ARG_TO_PARAM {callSite, position} edge (new EdgeType.ARG_TO_PARAM). Query (Phase 3): GET /variables/{name}/flow-forward|flow-backward?module=&depth= follow ARG_TO_PARAM*1..depth, returning {variable, variableType, module, depth} (DataflowStep); CLI flow-forward/flow-backward. Placeholder-target resolution carries edge properties via SET r2 += properties(r) so args survive redirection. Deferred: PERFORM USING, whole-area PARAMETER USING position mapping, Java method-argument dataflow. Fixtures DF_CALLER.nat
    • DF_CALLEE.nat.

Java parser

  • J6. Ingest noise: JDK/framework filter + target/ exclusion (2026-07-07) — a by-name ingest (POST /ingest/{name}) follows every referenced type; JDK/stdlib/framework types never resolve to a project file, so they dominated the unresolved list and ballooned the dependency fan-out (one job pulled ~1158 files). New ExternalTypes denylist (uppercased simple names: java.lang/util/time/io/nio/math/concurrent/stream

    • CDI/JPA/Panache/Mutiny types) is consulted in the dependency BFS — matching MODULE refs are neither chased nor reported as unresolved, so only genuine gaps remain. The file walk now excludes target/ by default (merged with the project's excludeDirs) so build-output generated sources don't create duplicate module/entity definitions. Documented heuristic (a project class named like a JDK type would also be skipped — vanishingly rare). New JavaIngestNoiseIT (JDK types filtered while a genuine missing type is still reported; target/ copy not flagged as a duplicate); all 80 server ITs green.
  • J5. Cross-class Java dataflow (precise arg→param) (2026-07-07) — flow-forward/flow-backward across classes. Investigation found the generic LINK_ARGS_TO_PARAMS linker already matched Java cross-class (MODULE→MODULE) calls but imprecisely — (callee)-[:CONTAINS*1..]->(param at position i) linked an argument to the position-i parameter of every method in the callee class (no method disambiguation). J5 makes it precise: JavaParser stamps callerFn (calling method) and calleeMethod (invoked method) on cross-class METHOD_CALL edges; new dataflow-gated steps LINK_ARGS_TO_PARAMS_JAVA_CROSS (+ _SCOPED) build ARG_TO_PARAM from each caller-scope argument to the named callee method's parameter at that position; and the generic linker (+ its scoped variant) is guarded with r.calleeMethod IS NULL so it no longer cross-matches Java. flow-forward/flow-backward traverse ARG_TO_PARAM unchanged, so they now span classes correctly (deep-ingest only, like Natural). Overloads still match by name (over-approx); constructor-argument dataflow deferred. New fixtures fixtures/java/flow/* (CheckoutFlow.checkout(amount) → PriceCalculator.applyDiscount(basePrice)) and JavaCrossClassFlowIT (forward + backward); all 78 server ITs + 11 parser tests green.

  • J4. Virtual/override (template-method) dispatch (Java) (2026-07-07) — a template method in an abstract base (doProcessItem → clearTable/writeEntities) dead-ended at the base's abstract method. New EdgeType OVERRIDDEN_BY (base-class FUNCTION → same-named overriding FUNCTION in a subclass), materialized by the always-run, idempotent enrichment step BUILD_OVERRIDDEN_BY (name-based, transitive over EXTENDS; constructors excluded naturally since a super/sub pair never shares a name). Exposed via a new GET /modules/{name}/functions/{fn}/overrides endpoint (GraphRepository.functionOverrides → FUNCTION_OVERRIDES, returning FunctionOverride {module, name, sourceFile, startLine, endLine}) and the function_overrides MCP tool — kept function-level rather than threaded into the module-level call-tree. Name-based matching is heuristic (overloads over-match); documented. New fixtures fixtures/java/override/* (abstract template base + two concrete steps) and JavaOverrideIT (2 tests: all overrides for an overridden method, empty for a non-overridden one); all 76 server ITs + 11 parser tests green.

  • J3. Interface → implementation resolution (Java) (2026-07-07) — new EdgeType IMPLEMENTED_BY (interface → concrete impl), materialized by the always-run, idempotent enrichment step BUILD_IMPLEMENTED_BY (the derived inverse of a real IMPLEMENTS edge). Java MODULE nodes now carry an isInterface property (a small slice of the future module-kind item). New ?resolveInterfaces=true option on callees and call-tree (REST + the MCP tools): callees hops an interface callee to its IMPLEMENTED_BY implementation(s) via coalesce(impl, callee) — a single impl is a clean deterministic hop, multiple impls expand, the interface is dropped; call-tree drops interface nodes that have a known implementation (their impls are already reached via the CHA synthetic CALLS edges). Flag-gated so the default query is byte-unchanged. New fixtures fixtures/java/iface/* (single-impl Notifier/EmailNotifier) reusing the multi-impl gateway fixtures; JavaInterfaceResolutionIT (3 tests); all 74 server ITs + 11 parser tests green. Note: the interface-traversal dead-end J3 originally described was already fixed by the earlier polymorphic CHA step — this item adds the explicit edge and the caller-facing hop/suppress option.

  • J2. DI + class-literal wiring edges (Java) (2026-07-07) — call-tree/ callees on a job class returned empty because steps are wired by CDI injection and class literals, not method calls. Two new EdgeTypes: INJECTS (a class → an injected bean type: @Inject fields, and constructor params of an injection-point constructor — @Inject-annotated, or the sole constructor of a CDI-scoped class) and REFERENCES (a class → a type used as X.class in argument position, e.g. super(AccountKeyInitStep.class, …) / batchlet(refName(X.class))). Emitted by JavaParser.addWiringEdges as class-level (MODULE→MODULE) placeholder edges; added to RESOLVABLE_EDGE_TYPES so the existing resolve-placeholder finalize step resolves them cross-file. Surfaced in callers/callees (query match widened to CALLS|EXTENDS|IMPLEMENTS|INJECTS|REFERENCES) as edgeKind = INJECTS/REFERENCES, mirroring how EXTENDS/IMPLEMENTS already appear; deliberately excluded from call-tree by default (kept CALLS-only; see J9 below for the opt-in traversal). New fixtures fixtures/java/wiring/* (a job with a class-literal step + @Inject collaborator, a constructor-injected bean) and JavaWiringIT (3 tests); all 68 prior ITs green (shared callers/callees query change verified non-regressive), 11 parser tests green. Re-test 2026-07-07 confirms this works: all 3 probed PUR job digests now list their steps via callees.REFERENCES (RiskImportJob → RiskInitStep/RiskProcessingStep/ RiskEndStep; same for MultiTableImportJob, KeyTableExportJob), and the step's callers.REFERENCES names the job. The former job-root dead-end is fixed at digest level. Follow-up usability gap fixed by J9 (below).

  • J9. Opt-in traversal that follows REFERENCES/INJECTS (Java) (2026-07-07) — J2 wired the INJECTS/REFERENCES edges into callees/callers but kept them out of call-tree (CALLS-only), so there was no automatic transitive tree from a job to its steps to their repositories — an agent had to chain digest/callees calls by hand. New followWiring boolean, mirroring the J3 resolveInterfaces pattern end-to-end (REST call-tree query param, MCP call_tree tool arg, GraphRepository.callTree, CypherQueries.callTree): when true, the transitive-traversal relationship pattern becomes CALLS|INJECTS|REFERENCES instead of CALLS (both INJECTS/REFERENCES are MODULE→MODULE, same as the CALLS edges call-tree already traverses, so mixing them into one variable-length path is schema-compatible). Composes with resolveInterfaces. callees/callers unchanged (already surfaced these edges at one hop since J2) — only call-tree's transitive traversal needed the flag. Off by default. New test callTreeFollowsWiringOnlyWhenRequested in JavaWiringIT (default call-tree excludes INJECTS/REFERENCES targets; followWiring=true includes them), reusing the existing fixtures/java/wiring/* fixtures.

  • J1a. JPA/Panache repository calls as Java DB accesses (2026-07-07) — Java db-accesses/sql-statements were always empty; now repository/EntityManager/ Panache active-record calls resolve to READS/WRITES on the entity's DB_TABLE. Parser (JavaParser): tags a repository class with its managed entity (repositoryEntity, from a PanacheRepository<E>/JpaRepository<E,Id>/… generic supertype); treats a PanacheEntity[Base] subclass as an entity; and emits a DB_ACCESS candidate node (contained under the calling function, dataType = READ/WRITE/DELETE by method-name prefix, value = call text, plus javaReceiverType/javaMethod/javaArgType) for a persistence-shaped call — a write/delete verb on any non-JDK receiver, or a read verb on a repository-named / static-entity / EntityManager receiver. A small JDK denylist keeps candidate volume down (J6 will generalize it). Enrichment (RESOLVE_JAVA_DB_ACCESS, a new always-run, project-wide, idempotent finalize step): derives the entity — argument type for em.persist/merge/remove(x), the repository's repositoryEntity, else the receiver itself — follows the entity's MAPS_TO to the DB_TABLE, then MERGEs DB_ACCESS-[:USES_TYPE]->DB_TABLE (for sql-statements) and fn-[:READS|WRITES]->DB_TABLE (for db-accesses). No new EdgeType: DELETE shows as WRITES in db-accesses (edge mode) but DELETE in sql-statements (node mode), mirroring Natural. Unresolved candidates (entity not in graph) stay unlinked — never polluting db-accesses, incremental-ingest-safe (no deletion). New fixtures fixtures/java/jpa/* (Panache repo + entity, EntityManager service, active-record entity) and JavaDbAccessIT (3 tests); all 68 server ITs + 11 parser tests green. J1b deferred (roadmap): @Query/JPQL/native-SQL string parsing, derived-name filters, and no-generic custom repositories (need symbol resolution). Re-test 2026-07-07 fixed by J7 + J8 (below): JavaDbAccessInheritedRepoIT (repo → project base → Panache base, plus a constant-valued @Entity(name=…)) now passes end-to-end — db-accesses/sql-statements resolve to RISK. A regression check against the actual 9 PUR jobs is still worth doing as a follow-up acceptance pass, but the two traced root causes are fixed.

  • J7. Panache-ness inherited through a project base class (2026-07-07) — J1a only recovered repositoryEntity when a repository directly extended a Panache/JPA base type; a repository extending a project-specific abstract base (e.g. RiskRepository extends AbstractPurRepository<Risk, String>, where AbstractPurRepository<Entity, Id> implements PanacheRepositoryBase<Entity, Id>) got nothing, since the entity generic sits one inheritance hop away from the Panache marker. Parser (JavaParser): repositoryEntityType now returns null (not a bogus name) when the repository base type's first type argument is one of its own type parameters rather than a concrete class; a new panacheEntityTypeParam tags such a project base class with the name of that type parameter, alongside its own ordered typeParams; a new typeArgs edge property on EXTENDS records the concrete arguments a subclass supplies (e.g. Risk,String). Enrichment (RESOLVE_PANACHE_INHERITED_ENTITY, a new project-wide, idempotent finalize step run before resolve-java-db-access): finds the base class's panacheEntityTypeParam position in its typeParams, reads the subclass's EXTENDS typeArgs at that position, and SETs repositoryEntity on the subclass — after which RESOLVE_JAVA_DB_ACCESS (J1a, unchanged) resolves its repository calls exactly as for a directly-Panache repository. Resolves one level of indirection (the observed real-world shape). New fixtures fixtures/java/jpa/inherited/* and JavaDbAccessInheritedRepoIT (2 tests); full ac-code-server unit + integration suite green.

  • J8. @Entity(name=CONST) / @Table with a constant table name (verified 2026-07-07, no change needed) — the physical table name is not always a string literal on the annotation; RiskEntity-shaped entities use a same-class constant reference instead (@Entity(name = Risk.TABLE_NAME) with public static final String TABLE_NAME = "RISK";). Turns out resolveTableName already resolved this via existing same-class constant-value resolution (collectConstants/resolveAnnotationString, predating J1a) — confirmed by parsing the JavaDbAccessInheritedRepoIT fixture in isolation before any code change. No fix required; J7 (above) was the actual blocker for that fixture's db-accesses.

  • 12. Java parser — constructors, parameters, field READS/WRITES (2026-06-16) — constructors → FUNCTION nodes named after the class; method/constructor parameters → VARIABLE nodes (CONTAINS from the function); field READS/WRITES from this.field and unshadowed bare-name references (AssignExpr target → WRITES; compound-assign / ++/-- → READS+WRITES; else READS). Parameter/local names shadow field references. Intra-class CALLS now also scans constructor bodies. /variables/{name}/reads|writes gained FIELD to the matched node types. Known limitations (deferred to item 15): same-named params in one file and overloaded constructors converge on one node. New fixtures OrderService.java + BaseService.java.

  • 16. Java parser — constant value resolution (2026-06-16) — JavaParser.literalValue does type-aware extraction (StringLiteralExpr.asString(), Integer/LongLiteralExpr.asNumber() — strips L, Boolean/CharLiteralExpr); non-literal initializers fall through to a null value. collectConstants resolves in-class references (bare NAME and ThisClass.NAME) transitively with cycle guard. resolveAnnotationString resolves an annotation member to a string via the same-class constants (used by item 13 for @Entity(name = TABLE_NAME)). Deferred: cross-class constant resolution / ConstantRefEnricher.

  • 13. Java parser — Hibernate/JPA entity recognition (2026-06-16) — detects @Entity/@Table (table name resolved through item 16, fallback to class name), @MappedSuperclass (no MAPS_TO), and per-@Column field metadata (columnName, columnDefinition, nullable, converterType from @Convert, hibernateType from @Type, isId, declaredIn). AstNode gained an optional Map<String,String> properties (persisted via SET node += n.properties). New EdgeType.MAPS_TO. Inherited columns are collected at query time by walking (:MODULE)-[:EXTENDS*0..]->(:MODULE)-[:CONTAINS]->(:FIELD). New EntityColumn record + ENTITY_COLUMNS query + GET /api/projects/{project}/modules/{name}/columns + CLI entity-columns; DB_TABLE_COLUMNS gained a 3rd UNION branch. Fixtures SampleLegacyEntity + AbstractSampleHistorized + AbstractSampleBase.

  • 14. Java parser — general inheritance enrichment (all classes) (2026-06-16) — implemented query-time (no new edge types). New GET /api/projects/{project}/modules/{name}/functions?includeInherited= endpoint + CLI functions [--include-inherited]: walks (:MODULE)-[:EXTENDS|IMPLEMENTS*0..]->(:MODULE)-[:CONTAINS]->(:FUNCTION) (MODULE_FUNCTIONS_INHERITED) when inherited is requested, else own only (MODULE_FUNCTIONS_OWN); each row tagged with declaredIn. New InheritedFunction record. Deferred: explicit INHERITS/OVERRIDES edge types. Fixtures OrderService.java (overrides describe()) + BaseService.java.

  • 17. Java dataflow (extends item 15 to Java) (2026-06-16) — JavaParser tags each method/constructor parameter VARIABLE node with paramPosition and captures the positional argument list of each intra-class MethodCallExpr onto the CALLS edge args property. New LINK_ARGS_TO_PARAMS_JAVA enricher maps args[i] to the callee function's parameter with paramPosition = i, where the caller-side variable is the caller function's own parameter or a field of its enclosing class. The flow-forward/flow-backward endpoints are language-agnostic. Deferred: cross-class Java calls (resolved in item 20).

  • 20. Cross-class Java call graph (2026-06-17) — Java CALLS edges were intra-class only. Now resolved at the class (MODULE) level, like Natural CALLNAT: JavaParser resolves a method-call receiver to a target class — a typed field/parameter/local (svc.method()), a capitalized bare name (static Foo.bar()), or a new Foo(...) constructor — and emits a MODULE-[:CALLS]->MODULE edge to a placeholder for that class. The placeholder resolves via resolvePlaceholderTargets(CALLS) and surfaces as a DependencyRef. Scope: class-level resolution only. Fixture OrderController.java.

  • 22. Polymorphic Java call resolution + implementation ingest (2026-06-18) — (1) Fan-out — LINK_CALLS_TO_IMPLEMENTATIONS, run last in the save post-processing: for every (MODULE)-[:CALLS]->(base) where base has incoming IMPLEMENTS/EXTENDS, it MERGEs a CALLS edge from the caller to each subtype reachable via an IMPLEMENTS|EXTENDS*1.. chain (Class-Hierarchy- Analysis over-approximation). Synthetic edges carry resolvedVia:'INHERITANCE'; dataflow is intentionally not routed through them. (2) Implementation ingest — ingestModule builds a reverse index (buildImplementorIndex) so when the BFS ingests an interface/base, its implementations are enqueued too. Fixtures fixtures/java/gateway/*; IT javaInterfaceCallsFanOutToImplementations + IngestModuleJavaIT.

  • J1b. Parse @Query/JPQL/native-SQL strings + derived-name filters (Java) (found 2026-07-06, done 2026-07-07) — motivated by a batch-job analysis pass over the PUR pur-batch module (2026-07-06, 9 concrete AbstractPurBatchJob subclasses). J1a resolves calls whose entity is syntactically recoverable (repository generic param, em.persist/merge/remove(arg), Panache active-record receiver). Added: @Query JPQL/native-SQL string parsing (leading DML verb → mode, FROM/UPDATE/INTO target → entity/table, attached at the method declaration since abstract repository methods have no call site to scan); derived query-method name parsing (findAllByClientAndStatus → derivedFilter property, tagged on the existing call-site candidate); and a no-generic-entity fallback (repository interfaces with no resolvable generic type anywhere, e.g. a custom IRiskRepository, guess the entity from a declared method's return type). See x-docs/agent-api-usage-ac-implementation.md for semantics/limits.

Project model & ingest endpoints

  • 21. Per-project root folder + server-side scan ingest (2026-06-17) — every project now carries a required root (server-side path its sources live under) and optional excludeDirs, stored on the Project node. The server walks the root and parses files: new ProjectIngestService + SourceFiles helper, with AstIngestService split into parse/save. Two new endpoints (@Blocking): POST .../ingest-all (scan the whole root) and POST .../ingest/{name} (ingest one module + its transitive deps via a server-side BFS). Both return an IngestSummary ({ingested, unresolved, duplicates, failed}). Strict duplicate detection (Java FQN; Natural (name, kind) from extension). sourceFile stored relative to the root. Old content-upload path removed. New CLI: project create, project update, project clear, ingest <name>, ingest all.
  • 12. GET /api/projects/{project}/modules — list modules in a project (2026-06-15) — new CypherQueries.LIST_MODULES/GraphRepository.listModules/ AnalysisResource.modules, returns MODULE nodes (name, sourceFile), excluding unresolved placeholders. Optional ?sourceFile=... filters to the module(s) defined in that file. Also added type=<NodeType> filter to search/identifier (400 INVALID_TYPE for an unrecognized type). New IT tests.

Module analysis endpoints (P2 reengineering support)

  • 72. A dispatch row reports its whole guard chain, not just the innermost DECIDE (2026-07-17) — found by a manual VMULTMN4 audit against the Natural source, the second such audit to pay for itself. DispatchEntry's contract is "guardField = guardValue routes to assignedField := assignedValue" — a sufficient condition. NaturalParser.guardProps walked the control-flow stack and returned on the first (innermost) value-DECIDE with an active branch (its javadoc said so outright), so for a nested DECIDE the outer guard was discarded and a conjunction was served as one condition:

    128| DECIDE ON FIRST VALUE OF #SHORT-VIEW
    165|   VALUE 'TABL'
    170|     DECIDE ON FIRST VALUE OF #FIELD-NAME
    171|       VALUE 'TX-TABLA'
    174|         MOVE P-DESCRIPTION TO YTABLMA0.TX-TABLA
    reported: #FIELD-NAME = 'TX-TABLA'
    truth:    #SHORT-VIEW = 'TABL' AND #FIELD-NAME = 'TX-TABLA'
    

    12 of that module's 69 rows (17 %) carried an incomplete condition. It over-generalises, which is the dangerous direction: this table exists to port DECIDEs to Java, and a condition that is too weak ports into a branch firing where the Natural code never would. Corpus scale, measured on disk: 824 of 7,729 value-DECIDE blocks (10.7 %) are nested, across 178 modules (YFRAMN10.nat: 199). Additive, not breaking (correcting the roadmap's own note): whenField/whenValue/whenValues keep their exact meaning — the innermost guard — so existing consumers see what they always saw; the chain travels in new whenChainFields/whenChainValues and surfaces as DispatchEntry.guards ([{field, values}], outermost first, AND-joined; each link's values OR-joined). Encoding needs two levels, so RS (U+001E) separates links above item 64's US (U+001F) for alternatives — neither can occur in Natural source. Decoding is in Java, not Cypher: zipping two nested splits back together needs index arithmetic that Cypher makes unreadable. The ac-ui migration dossier rendered the worst possible form — guardField + the lossy guardValue (item 64 already said "prefer guardValues"), i.e. the innermost guard in its ambiguous encoding, in the one screen meant for porting. It now renders the full conjunction. Live proof after refresh: VMULTMN4:174 → #SHORT-VIEW = TABL AND #FIELD-NAME = TX-TABLA, legacy fields unchanged. Vacuity-checked: reverting guardProps fails 3 of the 4 tests (guards comes back []) while the legacy-field test stays green — which is the additive claim, proven rather than asserted. The chain's order is asserted explicitly: ArrayDeque iterates innermost-first, so a missing reversal would silently yield an inside-out chain that a "contains both" assertion would accept. Left behind: #73 (a NONE branch's condition is a negation no chain can express — guards is strictly better, not total) and #74 (a duplication in the graph that item 72 exposed, not caused).

  • 67. call-tree's depth counts module hops, and says when it truncates (2026-07-17) — the last of the raw-hop siblings. callTree bounded on raw CALLS edges and returned min(length(p)) as the depth column, so the number measured how deeply the calling statement happened to be nested rather than dependency distance. Measured live on upms: WGEAGB0S's seven direct module dependencies came back at depths 1..3 (raw path lengths 3,4,5,6,7,7,8) — BGEAGFN0 reported 3 — although all seven are one module hop away. All seven now report 1. Depth comes from the path, not from ownership: min(size([n IN nodes(p) WHERE n.type='MODULE']) - 1). The decided semantics were "a FUNCTION inherits the hop count of the module that owns it", but measurement killed the literal reading: a subroutine defined in a copycode is one shared node CONTAINSed by every including module — L4N-ENTER (L4NCOPY.cpy) has 137 owners, and 156 upms functions have more than one. "Its module" has no single answer. Counting module nodes on the traversed path yields one answer per route, sidesteps ownership entirely, and gives the intended result (a root's own subroutines are depth 0). The bound was the hard part, and the first design was wrong. A raw *1..N cannot express a module-hop bound. The initial plan — a raw budget of depth × (1 + budget) — was killed in validation: at depth=1 a budget of 7 means 8 raw hops, and 8 raw hops reach 8 module levels when internal chains are short, so a depth=1 request would have expanded ~8 module levels and discarded the rest (costing what depth=8 costs today). Worse, it could report a wrong depth: a target reachable at module-depth 1 via a long internal chain and at module-depth 5 via a short one would be found only on the short route. Fix: moduleTree (item 65) supplies the modules within maxDepth module hops, and a QPP with a per-hop predicate prunes the traversal to them while expanding, so it can never wander beyond the requested depth. Verified against live upms before writing any code (QPP had already refuted one of my designs in item 65). truncated (new response field). The raw quantifier survives as a pure safety cap (agenticcode.call-tree.internal-budget, default 20). Because pruning bounds the search space, the quantifier is free — measured on upms, quantifiers 8/20/40 gave identical results at 2.0/2.6/2.1 s — so it is sized far above the observed maximum. Sizing it from a 300-module sample would have been wrong: that sample said "internal chains ≤ 6", but WGEAGB0S alone reaches ISINDATE at 8 raw hops for one module hop, and the true longest path is 10. Since any cap can cut, exhausting it now sets truncated: true rather than returning a short list that reads as complete — the exact failure mode that made items 65 and 68 bugs instead of documented limits. Conservative: true is possible for a complete result. A "known imprecision" I filed here as #71 was retracted 2026-07-17 — it was my measurement error, not a defect. The claim was that Natural external subroutines (module A performs a subroutine owned by module B) do not increment depth, "4,808 such calls in upms". That count applied the single-owner filter to the target but not the source, so every copycode-shared source function (L4N-ENTER: 137 owning modules) was counted once per owner. Source and target single-owner returns 0 — and it must, because resolvePlaceholderTargets never resolves a FUNCTION placeholder across modules, so the cross-module FUNCTION edge the claim assumed cannot exist. Path-based depth has no gap here. Two regressions the existing suite caught, both mine. followWiring (item J9) returned an empty list: the bounding BFS followed only CALLS, so with CALLS|INJECTS|REFERENCES traversal every wiring-only neighbour fell outside the module set and the predicate pruned it away — a bound must follow the same edge types as the traversal it bounds (MODULE_HOP_OUT_WIRING). And AnalysisResourceIT encoded the old semantics; checking the fixture instead of just relaxing the assertion showed the old expectation was the bug: INITIALIZATIONS and R-ADDRESS_SP are both subroutines of YADDRBN0_SAMPLE.nat itself, yet depth=1 omitted R-ADDRESS_SP — a DEFINE SUBROUTINE of the very module being asked about, hidden because it sat two raw edges away. Consequence to expect: a given depth now returns more, since the internals of modules within the bound are inside it. Both halves vacuity-checked: reverting the formula made DEPTHLEAF vanish from a depth=1 answer entirely (expected: <1> but was: null), and the truncation test is paired with an assertion that truncated is false on a normal query, so an always-true flag could not fake it. Dead code removed on the way: four unused callTree overloads (CypherQueries ×2, GraphRepository ×2).

  • 68. field-flow bounds on module hops, via a derived CALLS_MODULE edge (2026-07-17) — the last of the three raw-hop siblings item 65 uncovered. fieldFlow gated producer→consumer pairs on EXISTS { (wmod)-[:CALLS*1..%d]->(rmod) }, so its depth counted a module's internal PERFORM jumps: a consumer called from two subroutines deep sat 3 raw edges away and fell outside depth=1. The endpoint then reported nothing downstream consumes this field — a confident false negative, the same failure mode as item 65's "no DB access". Unlike item 65 this is a pairwise reachability test between two arbitrary modules, and the producer is optional ($module IS NULL), so item 65's per-root moduleTree BFS does not apply — it would have to run once per candidate producer. Fix: an enricher materialises (MODULE)-[:CALLS_MODULE]->(MODULE) (the persisted module-level projection of the call graph, same shape as EGO_NEIGHBORS_OUT), and fieldFlow bounds on that, which makes a plain variable-length bound correct again. (Correction: this was expected to supply item 67's bound too. It did not — item 67 turned out to need a set of module names to prune a QPP with, which is what item 65's moduleTree BFS already returns, so CALLS_MODULE is used by field-flow alone. It earns its keep there: field-flow is a pairwise test between two arbitrary modules, where a per-root BFS does not apply.) DELETE before rebuild is the load-bearing part, not MERGE. The edge is derived, and MERGE is idempotent only for edges that still exist — it never removes ones that shouldn't. The item-58 sweep deletes stale nodes only, so a surviving module's relationships are never reaped. Proven by removing the DELETE step: after deleting the CALLNAT from the producer's source and refreshing, field-flow still answered [STALECONS] — a dependency that no longer exists in the code. Ordering matters too: the projection runs last, after every step that adds CALLS edges (placeholder + dynamic-CALLNAT resolution, link-calls-to-implementations) or removes them (the item-62 data-literal reaping), since a projection is only as correct as the graph at the moment it runs. Both fixes vacuity-checked: reverting the bound made the field-flow test fail with an empty producer list.

  • 69. The origin file is part of an edge's identity (2026-07-16) — data loss, found by item 66's own regression fixture rather than by review. An edge was identified as MERGE (a)-[r:%s {lineNo: e.lineNo}]->(b) — (source, target, type, lineNo), with no file. Copycode expansion (item 46a) made that key ambiguous: a host statement on line 10 and a copycode statement on line 10 share it entirely, so the second SET r += e.properties overwrote the first. The endpoint then returned one row where two statements exist, and a real write was gone from the graph. Proven by reverting the fix: with MOVE 'Y' TO #KEY on COPYHOST.nat:10 and MOVE 'Z' TO &1& on MYTABLECOPY.cpy:10, variables/#KEY/writes returned only [{sourceFile=MYTABLECOPY.cpy, lineNo=11, ...}, {sourceFile=MYTABLECOPY.cpy, lineNo=10, ...}] — the host's own write had vanished, and the surviving edge even claimed viaCopycode=MYTABLECOPY. Fix: MERGE (a)-[r:%s {lineNo: e.lineNo, originFile: coalesce(e.properties.originFile, a.sourceFile)}]->(b). The fallback is a.sourceFile, not '' (as first sketched in the roadmap): for a host statement the edge really does originate in its source node's file, so r.originFile holds the true file on every discriminated edge and item 66's coalesce(r.originFile, <src>.sourceFile) keeps returning exactly what it did before. An '' default would have broken all four of those queries, since coalesce replaces only null. Scope was 20 MERGE sites, not the 1 the roadmap named: the batch merge plus 7 placeholder-resolution re-merges (which re-key the edge) and 6 dynamic-CALLNAT merges. In the dynamic ones the origin had to be threaded explicitly through the WITH DISTINCT caller, dyn, lit, which drops the src call-site node. Deliberately excluded: CONTAINS/INCLUDES/USES_TYPE (unique per node pair by construction — a discriminator cannot prevent a collision there and would add a string property to the most numerous edge type in the graph); the Java {lineNo: a.startLine} merges (copycode is Natural-only); and ARG_TO_PARAM — its collisions are real but benign, since two colliding edges carry identical (callSite, position) and no per-site payload, so discriminating them would only add parallel edges for dataflow to walk. The type test is a switch, not a static Set field: the query maps are static finals that call it from their initialisers, and a field declared after them is still null at that point (this bit — compiled clean, failed at runtime). Corpus incidence remains unmeasurable after the fact: the collision destroys the evidence one would count. Existing graphs are not migrated — the stale sweep deletes only nodes, so old edges would linger beside the new key; delete + re-ingest the project (~179 s for upms).

  • 70. Placeholder field resolution kept the copycode provenance (2026-07-16) — found while fixing item 69, by reading the queries around it. Four of the six resolution queries copied only two properties when redirecting an edge onto its real target (SET r2.value = r.value; SET r2.lineNo = r.lineNo), while the other two used SET r2 += properties(r). The four discarded viaCopycode, includedAt and originFile — the very provenance item 66 had just added. This made item 66 only half-effective, which its upms verification could not reveal: a bare field reference (#W-OPTIONS) resolves through resolveBareIncludedFieldTargets, which copies everything — and that is the field item 66 was verified on. A qualified reference (MYLDA.Q-FIELD, a field of a LOCAL USING data area) takes resolvePlaceholderFieldTargets instead and silently came back with viaCopycode: null, i.e. indistinguishable from a host statement. The split is in NaturalParser.lookupVariable: bare names go to scope.bareFields(), qualified ones to scope.placeholderFields() under a placeholder DATA_STRUCTURE. Fix: SET r2 += properties(r) in all four. Proven by reverting it — the new fixture returned expected: <[QUALCOPY]> but was: <[null]>. Note the two fixes are complementary: item 69's key incidentally carries originFile through resolution, but viaCopycode/includedAt need the +=.

  • 66. Call sites and variable accesses name the file their line is in (2026-07-16) — found by the same manual WGEAGB0S audit as item 65. The graph was already right; the API threw the answer away. The parser stamps copycode-origin edges (item 46a) with viaCopycode + includedAt and gives copycode-origin nodes the real .cpy file — exactly what this doc promised — but viaCopycode appeared nowhere in ac-neo4j-store or ac-code-server, so a line belonging to a .cpy was served as if it were a host line:

    before: #W-OPTIONS writes → sourceFile WGEAGB0S.nat, lineNo 18/20/22
    truth:  WGEAGB0S.nat:18/20/22 = the generated comment banner ("* System : Versis", ...)
            ISICINDI.cpy:18/20/22 = &1& := '0' / '1' / ' '   (&1& = '#W-OPTIONS', INCLUDE at host line 676)
    

    Host and copycode lines sat indistinguishably in one list (lineNo: 670 in the same response was a real host line). Severity differed per endpoint: variables/{name}/reads|writes made an explicit false claim (sourceFile = host + lineNo = copycode line); callers/callees gave unattributed lineNos (their sourceFileIndex is the callee's file, so they never claimed a call-site file). Scale in upms: 11,008 CALLS, 1,325 WRITES, 996 READS edges carry viaCopycode. Root cause was one missing line: CopycodePreprocessor.withVia already held origin.sourceFile() where it stamped viaCopycode, but AstEdge has no sourceFile field, so the path was dropped — edges now carry it as originFile. On top of that: AggregatedCallRef.lineNos: [int] → sites: [CallSite] (lineNo, callSiteFileIndex, viaCopycode, includedAt) — per site, because one target can be called from both the host and a copycode, which entry-level provenance cannot express. Deliberately named callSiteFileIndex, not sourceFileIndex: at entry level that word means the callee's file, and reusing it would have baked in the next misreading. variables/{name}/reads|writes return coalesce(r.originFile, f.sourceFile) plus viaCopycode / includedAt. payload already did this correctly (per-entry sourceFile + lineNo) and was the model. Breaking change: REST/MCP/CLI stayed in sync for free (MCP returns the record, the CLI prints the raw JSON); ac-ui's IdentifierPopover was carrying the same confusion — keying PERFORM lines to the caller's definition file — and now reads each site's own file. Live: ADLML02 in WGEAGB0S's callees → lineNo 26, callSiteFile src/manual/copycode/ISIYESNO.cpy, viaCopycode ISIYESNO, includedAt 673, while BGEAGFN0 → lineNo 497, WGEAGB0S.nat, no copycode; the #W-OPTIONS writes above now name ISICINDI.cpy with includedAt 676, line 670 unchanged. Tests in CopycodeExpansionIT (call site + a variable written from both host and copycode), both vacuity-checked: without originFile they fail with expected: <MYTABLECOPY.cpy> but was: <COPYHOST.nat> — the defect verbatim. Building that fixture also turned up roadmap item 69 (two edges on the same line number from different files collapse in the MERGE).

  • 65. ?depth= counts module hops, not raw CALLS edges (2026-07-16) — found by a manual endpoint audit of WGEAGB0S (upms) against the Natural source. db-accesses?depth=1..4 reported zero DB accesses for a module that reaches YGEAGBNH's tables through exactly two module calls. An empty list reads as "this module touches no database" — a confident false negative, the worst failure shape for an agent. Cause: a CALLS edge starts at the statement that makes the call, not at the enclosing MODULE node, so (m)-[:CALLS*1..N]->(hop:MODULE) also steps through internal PERFORM jumps. The real path is MODULE:WGEAGB0S → FUNCTION:GET-DATA → FUNCTION:GET-MAIN-DATA → MODULE:BGEAGFN0 → FUNCTION:READ-FILE → MODULE:YGEAGBNH — 5 raw edges for 2 module hops, so the answer only appeared at depth=5, and the required value depends unpredictably on the callee's subroutine nesting. EGO_NEIGHBORS_OUT (item 49) already defined a module hop correctly as (a)-[:CONTAINS*0..]->()-[:CALLS]->(b:MODULE) and BFS'd one hop at a time — so ego_graph?depth=2 and db-accesses?depth=2 disagreed about the same graph. Fixed by giving the transitive endpoints the same definition: new MODULE_HOP_OUT (batched over a whole BFS frontier — one query per hop, not per module) + GraphRepository.moduleTree, consumed by DB_ACCESSES_FOR_MODULES, SQL_STATEMENTS_FOR_MODULES and the ?module= scope of variables/{name}/reads|writes (whose EXISTS { (root)-[:CALLS*1..N]->(owner) } had the same flaw). Done in Java because Cypher cannot express the hop: a quantified path pattern rejects both the variable-length CONTAINS*0.. inside it ("Variable length relationships cannot be part of a quantified path pattern") and nesting. Live: WGEAGB0S db-accesses now depth=1 → 0 (correct — none of its 7 direct callees touches a table), depth=2 → 7 incl. VDB2-VERSIS_GENAGREE via YGEAGBNH at line 2617 (matches the source's FIND (1) VDB2-VERSIS_GENAGREE), consistent with ego_graph. Regression test ModuleHopDepthIT (fixtures DEPTHROOT/DEPTHLEAF: one module hop, but the CALLNAT two PERFORM levels deep); vacuity-checked — dropping the CONTAINS*0.. step fails 2 of its 3 tests, while the depth=0 test correctly stays green.

  • P2-a. Control-flow parsing (IF/ELSE/END-IF, FOR/END-FOR, REPEAT/END-REPEAT) — Natural construct #6. New NodeType.CONTROL_FLOW node per block (dataType = IF/FOR/REPEAT, value = condition/loop-spec text, line span = full block incl. matching END-*), nested via CONTAINS edges.

  • P2-b. Module "context bundle" endpoint (GET /modules/{name}/context) — aggregates functions, callers/callees, db-accesses/sql-statements, and a variable read/write summary for a module in a single response. New GraphRepository.moduleContext() + CLI context command.

  • P2-c. DATA_STRUCTURE field schema — new GET /api/projects/{project}/data-structures/{name}/fields returns the flattened field list (name, type, dataType/length, const value, immediate parent) of a DEFINE DATA/DDM structure. New CLI data-structure-fields.

  • P2-d. DB_TABLE column schema — new GET /api/projects/{project}/db-tables/{name}/columns derives column info for a DB_TABLE from the DATA_STRUCTURE used as the INTO VIEW target of SELECTs against that table. New CLI db-table-columns.

  • P2-e. Transitive call graph — new GET /modules/{name}/call-tree?depth=N endpoint returns all modules/functions transitively reachable via CALLS edges from the module (each with its minimum hop-depth). depth defaults to 3, clamped to 10. New CLI call-tree --depth.

Correctness gaps (wrong/empty results)

  • 77. Bare-field resolution ignored which module it was resolving for (2026-07-17) — found while root-causing item 74, which it does not fix (see the note there). resolveBareIncludedFieldTargets documented itself as redirecting a bare reference "only when exactly one field of that name is reachable through the module's resolved INCLUDES", but resolved project-wide across all owners at once. An unresolved bare field is one node per (name, project) (item 76 — MERGE_NODES keys on sourceFile, "" here), so (m)-[:CONTAINS]->(ph) matched it once per owning module; WITH ph, collect(DISTINCT realv) then dropped m, and the redirect's MATCH (src)-[r]->(ph) never bound src to m. Two bugs in opposite directions, both measured in upms:

    • under-resolution — two owners disagreeing on the target made size(matches) = 1 fail for everyone, including owners for whom the name was unambiguous (38 placeholders);
    • misattribution — an owner with no matching include had its edges redirected onto another module's field anyway (28 placeholders, 199 module-field pairs). Live example: BMTABBP0 was recorded writing ##MSG-NR of CDPDA-M.pda, a data area it does not include, while the other 63 modules touching that node genuinely include it.

    Both surface as a fabricated dataflow: field-flow pairs producer and consumer only when both touch the same node, so two modules sharing nothing but a field name were reported as passing data between them — a dependency an agent would act on. variables/{name}/reads|writes cannot see any of it (it matches every node with the name and reports only the accessing side), which is why the defect survived this long.

    Fixed by carrying m into the aggregation and binding src with (m)-[:CONTAINS*0..1]->(src). The *0..1 bound is exact, not a guess: a placeholder edge's src is only ever the MODULE (35,572 edges) or a FUNCTION directly under it (157,616; all 16,231 such functions are direct children) — never a CONTROL_FLOW node, so the walk also stays clear of item 75's CONTAINS cycles. Also aligned the scoped and unscoped twins, which had silently disagreed (*1..10 vs unbounded → the same module could resolve differently depending on which pass ran); the shared bound is now INCLUDE_FIELD_DEPTH, and it is required, not an optimisation, because CONTAINS is not acyclic. DELETE_RESOLVED_BARE_PLACEHOLDER_CONTAINS became per-module in the same change — its global "no edges left" test would otherwise have kept a resolved module's CONTAINS alive on the strength of another module's unresolved edges.

    Cost: the resolver is ~4.4× slower on upms (34 s → 2 m 30 s for WRITES), paid on deep ingest only. Known residue (item 76): 132 of 18,539 source nodes are copycode FUNCTIONs shared by several modules; their edge to the placeholder is a single edge, so per-module binding cannot separate what is one node. Needs per-module placeholder identity, not a better query. Verified by BareFieldModuleScopeIT (2 of 3 tests fail without the fix; the third is a control proving resolution still works, since resolution failing everywhere would also yield no flow).

  • 75. Dynamic-CALLNAT resolvers wedged on the cyclic CONTAINS graph (code + ITs done 2026-07-17; corpus re-verify pending). A whole-root deep refresh of upms hung finalize step 17 (resolve-dynamic-callnat-intra-indirect) for ~2 h without completing, blocking every later step (including item 77's bare-field resolution — so item 77 could not be corpus-verified). Cause: that step joins three unbounded (caller:MODULE)-[:CONTAINS*0..]-> anchors, and per-caller path enumeration over item 75's cyclic copycode containment is cubic (a read-only probe of the exact query for the single caller DAGNTFN0 did not finish in 60 s). Fix: the six RESOLVE_DYNAMIC_CALLNAT_* queries (INTRA / INTRA_INDIRECT / CROSS + _SCOPED twins) find a module's own statements by sourceFile equality (a MODULE is 1:1 with its sourceFile; its CALLNAT sites and WRITES statements share it) instead of descending CONTAINS — an index hash-join that cannot cycle. The cross-module resolver's dispatch variable may live in an included PDA, so it is scoped to the caller's own file or a data structure the caller INCLUDES (name-only match would pull 20958 unrelated same-named vars → the scope filter keeps 307). Proven equivalent on upms: resolve-dynamic-callnat-intra returns the identical 31 (caller, target, lineNo) triples project-wide; the rewritten indirect step runs project-wide in ~6 s instead of wedging. Existing dynamic-dispatch ITs (intra / indirect / cross / scoped / unresolved- survival, 8 tests) stay green. Does not remove the underlying CONTAINS cycles (still open in roadmap.md #75); it removes this family's dependence on traversing them. Corpus-verified on upms v68 (2026-07-18): 251 resolved dynamic CALLS (107 indirect), cross-module 113 callers / 165 edges.

  • 75b. link-args-to-params (dataflow step 27) rewritten off CONTAINS* (2026-07-18) — the same item-75 blow-up: three unbounded CONTAINS* anchors (src, cv, pv) over the cyclic copycode graph made it the single slowest finalize step on upms (~25 min). All three now resolve through the (project, sourceFile, …) index instead of descent: the call site src by sourceFile equality (module ↔ file is 1:1), and the argument variable cv / callee parameter pv over the file-set {module's own file} ∪ {files it INCLUDES}, with cv resolved before pv is unwound so the file-lists are not cross-producted. Corpus-verified: ARG_TO_PARAM still built (436 edges on the v68 recreate) and the step no longer registers as a running query across 45 s poll windows (≈ seconds, not 25 min). After this, a full deep upms finalize (~28 min) is dominated entirely by step 19 (resolve-field- placeholder WRITES), which is an item-76 cardinality problem (shared placeholder nodes), not CONTAINS* — tracked separately.

  • 76. Per-module identity for field placeholders (2026-07-18) — an unresolved field (VARIABLE/CONSTANT, sourceFile="") used to be ONE shared node per (name, project), referenced by up to 916 modules. That sharing made resolve-field-placeholder WRITES (finalize step 19) join (ph)-[:CONTAINS]->(phv) × (src)-[:WRITES]->(phv) into a ~28.7M-row cartesian → ~55 min on upms. Fix: GraphRepository.placeholderOwner stamps an ownerModule (the referencing file) onto field placeholders only; it joins both the in-memory dedup key (mergeKey) and the DB merge key (MERGE_NODES/MERGE_POSITIONAL_NODES). MODULE/DB_TABLE/DATA_STRUCTURE placeholders keep ownerModule="" and stay shared (a CALLNAT/USING target still resolves once). A companion sweep DELETE_STALE_PLACEHOLDER_NODES reaps a re-ingested file's obsolete placeholders (the file sweep skips sourceFile=""). Corpus (v69): step 19 55 min → 55 s, whole deep finalize ~59 min → ~6 min, item-77 misattribution stays 0, graph grows only +2,207 nodes (resolved placeholders are deleted in step 26), ARG_TO_PARAM well-formed (30,583 edges, 0 position mismatch). Field/dynamic ITs green.

  • P1-h. Transitive DB access resolution (2026-06-16) — /db-accesses and /sql-statements returned [] for modules whose SQL lives behind one or more CALLNAT hops (e.g. WGEAGB0S → BGEAGFN0 → YGEAGBNH). Added ?depth=N param; when depth>0, new dbAccessesTransitive/sqlStatementsTransitive queries walk up to depth CALLS hops and annotate each result with via (the intermediate module name). Clamped to 10.

  • P1-i. /db-tables/{name}/columns is effectively broken (2026-06-16) — returned only one column for tables whose column schema is encoded in .pda files. parseDataArea now detects VIEW OF <table> and creates a USES_TYPE edge from the DATA_STRUCTURE to a DB_TABLE placeholder. DB_TABLE_COLUMNS upgraded to a UNION query covering both the PDA path and the SELECT INTO VIEW path. New unit test pdaViewOfCreatesUsesTypeEdgeToDbTable.

  • P1-j. SQL statement text is truncated (2026-06-16) — every sqlStatements entry had statement cut off at ...WHERE. The DB_READ handler now reads ahead through subsequent lines until an empty line or END-FIND/END-READ, accumulating all lines into DB_ACCESS.value. New unit test multiLineFindStatementTextIsCapturedFully.

  • P1-k. FIND (1) misparsed as a table named (1) (2026-06-16) — FIND (1) <view> produced a phantom DB-access entry. Fixed by adding (?:\\(\\d+\\)\\s+)? to the DB_READ pattern. New unit test findWithRecordLimitIsNotMisparsedAsTableName.

  • P1-m. Dynamic CALLNAT <var> calls are invisible in the call graph (2026-06-21) — the parser/enricher only recorded literal CALLNAT 'NAME' edges, so variable-target calls (CALLNAT #WIF ...) produced no CALLS edge. Real case: W-MNT-N0 dispatches ~90 Wxxxx*S XML-interface programs via a single CALLNAT #WIF; the target is resolved by KDWWIFN0. Done — and v1 went further, doing both intra- and cross-module. Parser detects CALLNAT <var> (new CallKind.CALLNAT_DYNAMIC), emits a variable-named placeholder + marker edge, captures multi-line argument lists; PARAMETER USING advances paramPosition. Enrichment adds RESOLVE_DYNAMIC_CALLNAT_INTRA and RESOLVE_DYNAMIC_CALLNAT_CROSS (follows ARG_TO_PARAM into the callee — recovers the W-MNT-N0 → KDWWIFN0 → WGEAGB0S chain), plus scoped variants and placeholder cleanup. Required fixing LINK_ARGS_TO_PARAMS to source calls via (:MODULE)-[:CONTAINS*0..]->(src)-[:CALLS] so subroutine-nested CALLNATs match. Live-validated on upms: deep-ingesting W-MNT-N0 (8.4 s) yields callers/WGEAGB0S = W-MNT-N0 (edgeKind=CALLNAT_DYNAMIC) and callees/W-MNT-N0 = 124 resolved dynamic targets.

  • P1-n. Data-structure fields lack scope and often have null dataType (2026-06-21) — the parser now stamps every DEFINE DATA field/variable with a scope property (PARAMETER|LOCAL|GLOBAL|INDEPENDENT), surfaced as a scope field on both DataStructureField and IdentifierMatch (/search/identifier) — the latter exposes the motivating inline param #P-CALLED-PROG. Covered by NaturalParserTest + AnalysisResourceIT.

  • P1-o. No search by literal value (2026-06-21) — a program name like WGEAGB0S can exist in the graph only as a string literal assigned to a field. Added GET /search/value?value={v} (SEARCH_BY_VALUE, ValueMatch): quote-insensitive, returning both nodes carrying the value as a constant (kind=NODE) and literal assignments to a variable (kind=ASSIGNMENT). Missing value → 400 MISSING_VALUE.

  • P1-p. Edge provenance on call results (2026-06-21) — done via the edgeKind=CALLNAT_DYNAMIC provenance carrier: stamped on every inferred edge and surfaced through the existing edgeKind = coalesce(r.callKind, …) projection, so callers/callees distinguish dynamic (inferred) from literal CALLNAT/PERFORM (static) with zero DTO changes.

  • P1-q. /data-structures/{name}/fields over-returns across modules (field pollution) (2026-06-21) — for a structure name shared by many modules (e.g. W-WIF-A1, used by ~80 Wxxxx0S subprograms), the endpoint returned a polluted list (269 rows spanning ~80 unrelated modules). DATA_STRUCTURE_FIELDS now selects the canonical definition(s) — those with a non-empty sourceFile — falling back to empty-sourceFile stubs only when no real def was ingested, and constrains each field's parent to lie within the chosen structure (p = s OR (s)-[:CONTAINS*1..]->(p)). Live: W-WIF-A1 266→115 rows with 0 module-parented leak; YGEAGROW 430→36. Covered by dataStructureFieldsAreScopedToCanonicalDefinition.

  • P1-r. Module→data-structures endpoint (2026-06-21) — added GET /modules/{name}/data-structures (MODULE_DATA_STRUCTURES, ModuleDataStructure) returning one row per referenced structure: name, relationship (USING copybook via INCLUDES / INLINE group via CONTAINS), area (PDA/LDA/GDA/INLINE/UNKNOWN), fieldCount, sourceFile. Live WGEAGB0S now exposes its whole interface (W-WIF-A1 PDA/118, W-WIF-A2 PDA/10, W-WIF-A4 PDA/5, the LDAs) and flags unresolved CDPDA-M/CDPDA-P as UNKNOWN/0. Note: precise PARAMETER/LOCAL/GLOBAL USING scope is not stored on the INCLUDES edge (area is the proxy). Covered by moduleDataStructuresListsReferencedAreas.

  • P1-s. Referenced PDAs/copybooks not fully ingested (empty field lists) (2026-06-21) — WGEAGB0S declares PARAMETER USING CDPDA-M, but GET /data-structures/CDPDA-M/fields returned []. Root cause: the file is ingested, but CDPDA-M.pda's sole top-level group is named MSG-INFO, while a module references it by file/DDM name. Fix: when a data area has exactly one top-level group whose name differs from the file name, parseDataArea wraps it in a root DATA_STRUCTURE named after the file. Multi-top-group areas and matching-name areas keep their existing shape. Covered by parsesPdaWithRedefineGroup + usingResolvesCopybookWhoseGroupNameDiffersFromFileName (fixtures MSGAREA.pda/MSGUSER.nat). Re-ingest required.

  • P1-u. Surface a module purpose/description (2026-06-21) — parsers stamp a description property on the MODULE node: Natural scans the leading comment banner (priority **SAG TITLE: → * Title : → **SAG DESCS(n): → * Function :); Java uses the class Javadoc's first line. Surfaced as ModuleContext.description (/context) via MODULE_SOURCE_FILE (fetched as a ModuleHeader), null when no banner. Covered by NaturalParserTest, JavaParserTest, AnalysisResourceIT. Re-ingest required.

  • P1-x. /callers silently omits incoming EXTENDS/IMPLEMENTS edges (2026-06-21) — "who extends/implements this class?" returned nothing (live on pur, GET /modules/AbstractUPMFESvc/callers was [] despite 22 controllers). Root cause: callers(scope) matched only [r:CALLS], whereas callees(scope) already traversed [r:CALLS|EXTENDS|IMPLEMENTS]. callers(scope) now mirrors callees: matches CALLS|EXTENDS|IMPLEMENTS and tags inheritance callers edgeKind=EXTENDS/IMPLEMENTS. scope=internal still excludes inheritance; scope=external includes it. Covered by callersSurfaceIncomingInheritanceEdges.

  • P1-y. Double-hash (##) field names truncated to # (2026-06-22) — the data-area field regex DATA_AREA_FIELD accepted only a single leading #, so a Natural variable named ##MSG was parsed with name = "#". Widened DATA_AREA_FIELD's name class to [#A-Za-z][#\w-]* so a leading run of # is captured whole. LEVEL_FIELD and IDENTIFIER_TOKEN already handled ##. Covered by doubleHashFieldNamesAreNotTruncated. Re-ingest required for ## names.

  • P1-z. Dispatch table queryable — DECIDE/IF value→assignment correlation (2026-06-22) — analyzing the router KDWWIFN0 could recover the set of 130 dispatched programs but not the mapping (#P-OBJECT-TYPE = 'genagree' → #P-CALLED-PROG := 'WGEAGB0S'). The NaturalParser now tracks the active DECIDE ON [FIRST] VALUE OF <subject> branch: it records the subject per DECIDE block and the current VALUE '<literal>' (handling NONE/ANY VALUE resets and comma-separated value lists), and stamps every literal assignment made inside that branch — including ones nested in an inner IF — with whenField/whenValue properties on the WRITES edge. New endpoint GET /modules/{name}/dispatch-table (DISPATCH_TABLE, DispatchEntry) returns rows {guardField, guardValue, assignedField, assignedValue, lineNo}. Covered by decideValueAssignmentsCarryDispatchGuard + dispatchTableRecoversObjectTypeToProgramMapping (fixture KDDISP.nat). Live-validated on upms (2026-06-22): GET /modules/KDWWIFN0/dispatch-table returned ~140 rows reproducing the hand-built kdwwifn0-prog-routing.csv mapping (genagree → WGEAGB0S/WGEAGX0S, etc.). Re-ingest required. Known limitations: the inner IF #L-LIST (list-vs-detail) sub-guard is not separately captured (both branch assignments share the DECIDE's whenValue); the field-placeholder resolution does not copy edge properties, so a guarded write to a copybook field would lose its guard on redirect.

Token efficiency (payload shape)

  • P1-l. Repeated absolute sourceFile paths inflate payloads (2026-06-16) — callers/callees and call-tree now return a wrapper with a deduplicated sourceFiles: [...] index and items that carry sourceFileIndex: int. ModuleContext gains a top-level sourceFile field. New DTO records: CallRefResponse, AggregatedCallRef, CallTreeResponse, CallTreeItem.
  • P1-m. Call-site aggregation (2026-06-16) — callers/callees queries use collect(r.lineNo) AS lineNos with a GROUP BY (name, type, sourceFile, edgeKind); db-accesses groups by (name, mode). lineNos sorted in Java.
  • P1-n. Separate PERFORM (intra) from CALLNAT (inter) (2026-06-16) — added edgeKind field (CALLNAT/PERFORM/EXTENDS/IMPLEMENTS) to AggregatedCallRef, and ?scope=external / ?scope=internal filter on /callers and /callees. CLI gains --scope.
  • P1-o. Field projection + sub-array pagination on /context (2026-06-16) — added ?include=functions,dbAccesses,... projection and ?limit=N&offset=N pagination applied per sub-array. CLI context gains --include, --limit, --offset.
  • 7. Pagination (2026-06-17) — added ?limit=N&offset=N to callers/callees/db-accesses/search-identifier (and the transitive db-accesses variant), default limit=50, offset=0. CLI gains --limit/--offset. New IT calleesPaginationLimitsAndOffsets.
  • P1-t. /context is heavy by default — make sub-arrays opt-in, return counts (2026-06-22) (found 2026-06-21) — the live WGEAGB0S /context was ~50 KB, ~70% of it the variableAccesses array (347 entries) which the analysis never used. ?include= projection existed but the default still returned every section fully expanded. Flipped the default to lean: heavy sub-arrays (variableAccesses, large sqlStatements) return a summary by default — e.g. variableAccesses: { count, byFunction: {...}, byMode: { READS, WRITES } } — and the full list is emitted only via ?include=. A summary line replaces 347 objects; the biggest single token lever found in the WGEAGB0S analysis.
  • P1-v. Tiny /modules/{name}/digest triage endpoint (2026-06-22) (found 2026-06-21) — /context is the "expand" call (tens of KB); there was no "should I dig deeper" call. Added GET /modules/{name}/digest with a deliberately small contract: description (P1-u), function count, callers/callees names only grouped by edgeKind, DB table names, referenced data-structure names + field counts (P1-r) — designed to a token budget, not a full dump.
  • P1-w. Names-only / field-projection mode on list endpoints (2026-06-22) (found 2026-06-21) — call-tree, callers, callees, search/identifier are often used only to enumerate names, yet each row still carried sourceFile(Index), line ranges, dataType, etc. Added a ?fields=name variant that drops per-row detail to just the name (+ type) when the agent is enumerating, on top of the existing sourceFiles-index dedup (P1-l). Complements P1-t/P1-v.
  • P1-x. Dynamic CALLNAT via lookup array + keep unresolved dynamic calls visible (2026-06-22) — a CALLNAT <var> whose dispatch variable is loaded from a lookup array (e.g. ASSIGN #TBL(1) = 'WPARTD2S', #W-ACT-PROG := #TBL(#I), CALLNAT #W-ACT-PROG — WPARTX2S L994) was dropped entirely: (1) the parser missed the subscripted literal write (ASSIGN #TBL (1) = …, space before (), and (2) the unresolved dynamic marker was deleted, erasing the call site. Fix: parser now captures subscripted assignment targets; a new intra-module indirect resolver follows the dispatch var's same-line READS to the source array and resolves its literals (tagged indirect); unresolved dynamic markers are now kept (only resolved-site markers are reaped) so an agent can still see/investigate the dynamic call in callees/context.

Agent API / MCP tooling gaps

  • 78. project recreate — re-initialise a project from its own stored config (2026-07-17) — POST /api/projects/{p}/recreate (+ ?deep=true) and ac project recreate <name> [--deep]. There was no way to say "wipe upms and set it up exactly as it was": recreating meant reading project list, writing down root/language/excludeDirs/generatedDir/userExitDir, deleting, and re-supplying them by hand. That is not hypothetical — on 2026-07-17 the ac project was deleted, the re-create failed with 400 LANGUAGE_REQUIRED (mandatory since item 47), and the project was gone for a while. A typo in generatedDir would not even error; the item-47 LoC split would just report nonsense. The (:Project) shell is not deleted, only its AstNodes. The shell holds nothing but config (verified: keys(p) = name, root, language, excludeDirs, generatedDir, userExitDir), so keeping it is observably identical to delete-then-create while removing the window in which the config can be lost. The root is resolved before anything is deleted (withResolvedRoot, which already existed for refresh): unlike create — where an unreachable root costs nothing — recreate would otherwise wipe a good graph and then find nothing to rebuild from. A moved root now fails 400 ROOT_NOT_FOUND with the graph untouched. Not atomic and not claimed to be: a failure after the delete leaves the project with an empty graph (re-run to finish), and the root could vanish between check and scan. It turns the common silent failure loud; it does not make the operation transactional. No MCP tool — MCP is deliberately read-only + ingest, and this is destructive. Verified by ProjectRecreateIT (3 tests); the missing-root guard was vacuity-checked by moving the delete ahead of the check, which fails exactly that test and no other.

  • 79. One delete verb, with the blast radius spelled out (2026-07-17) — clear used to be bound twice, at two levels, with drastically different reach: ac clear emptied the entire database while ac project clear <name> deleted one project — a forgotten word apart, and only the first prompted for confirmation. ProjectCommand.ClearCommand was also a byte-for-byte duplicate of DeleteCommand (same body, same endpoint), so it was dead weight on top of an overloaded name. Both are gone; deleting is now:

    ac project delete <name>   one project
    ac project delete --all    every project (prompts; -y skips)
    ac project delete          usage error — never "delete everything"
    

    --all and a name are mutually exclusive and one is required, enforced by picocli's ArgGroup(multiplicity = "1") rather than a hand-rolled check. The confirmation prompt is inherited verbatim from the old ac clear. Covered by ProjectDeleteArgsTest (6 tests, parse-level: apiClient() builds its client inside the call, so there is no seam to inject a fake and executing would fire real HTTP; what a valid parse does is covered by ProjectRecreateIT at the API level). Breaking change: ac clear and ac project clear no longer exist.

  • Batched deletes (2026-07-17, with item 78) — DELETE_PROJECT/CLEAR_ALL were a single DETACH DELETE in one transaction. At the corpus's real size — upms alone is 454,300 AstNodes joined by ~1.3M relationships, 536,210 nodes across all projects — that is one transaction the heap must hold at once, with no server.memory.heap.max_size configured (only a 512M pagecache). Deleting was rare enough to get away with; item 78 makes it routine. Both now use CALL { ... } IN TRANSACTIONS OF $batchSize ROWS (agenticcode.delete.batch-size, default 10000). This forced the delete path off session.executeWrite onto implicit transactions — Neo4j rejects CALL { } IN TRANSACTIONS inside an explicit one ("can only be executed in an implicit transaction") — which is why the existence check is now a separate statement. Trade accepted deliberately: an interrupted delete leaves a partially emptied project that re-running finishes, versus an unbatched delete that risks not completing at all.

  • deploy.sh → manage-ac.sh (2026-07-17) — renamed (via git mv, history preserved) and given a command list. manage-ac.sh deploy is the old up, unchanged. No argument now prints help instead of deploying — a full deploy bumps agenticcode.version and rebuilds everything, which a bare invocation should not trigger by accident. No up alias: keeping two names for one action is the exact mistake item 79 removed. Added restart (restart the server without a build — also the way to abort a running server-side deep refresh, whose HTTP client can be killed without stopping the job), status (containers, server version and projects; every probe guarded so it reports a down stack instead of dying on it under set -e), and down (whole stack incl. neo4j — deliberately not -v, which would delete the graph volume). All in-repo references updated, including the two user-facing CLI strings that told people to run ./deploy.sh cli; features.md's historical entries below are left as they were written.

Found while dogfooding AgenticCode on its own codebase for the J1b task (2026-07-07): friction points where the agent-facing tools didn't answer a question the graph already had the data for.

  • 27. Generic "inspect node" tool (done 2026-07-07) — added GET /nodes/{id} (+ MCP inspect_node): every property of a node, via Node.asMap() so new parser properties show up automatically. search/identifier and search/annotation now also return id so there's a way to get one — without that they'd have been unreachable in practice. Node ids are regenerated on every re-ingest (merge key is (type, name, sourceFile, project); id is unconditionally overwritten) — documented as a limit, not fixed.

  • 28. Source-snippet-by-node tool (done 2026-07-07) — added GET /nodes/{id}/source and GET /modules/{name}/source?startLine=&endLine= (+ MCP node_source/module_source), reusing the ingest-time root resolution and the same UTF-8/ISO-8859-1 fallback decoding the parsers use.

  • 29. Text/string-literal search (done 2026-07-07) — search_identifier only matches identifier names; there was no way to search annotation names or string-literal node values (e.g. "does anything reference @Query", "does any SQL string contain orders"). Added: search/value gained a contains parameter (case-insensitive substring, vs. the existing exact match) for literal/assigned string values; a new search/annotation endpoint (+ MCP tool search_annotation) finds Java classes/methods/constructors/fields carrying a matching annotation, backed by a new generic annotations property captured at parse time on every such node (independent of any annotation's own specific interpretation elsewhere, e.g. @Entity/@Query). See x-docs/agent-api-usage-ac-implementation.md section 7-8 for usage.

  • 31. Resolve inherited REFERENCES/wiring edges on concrete subclasses (found 2026-07-07/08, PUR pur-batch re-evaluation, done 2026-07-09) — call_tree/ module_digest only show a REFERENCES edge on the class that syntactically declares it. Where the wiring (e.g. JBeret step construction) lives in a shared abstract base class rather than the concrete subclass — AbstractKeyTableImportJob.jobSteps() wires KeyTableImport{Init,Processing, End,PassInit,Logging}Step — the concrete subclasses (KeyTableImportValidationJob/KeyTableImportCommitJob) show no REFERENCES at all and call_tree(..., followWiring=true) on them returns empty, even though call_tree(AbstractKeyTableImportJob, followWiring=true) reaches the full step tree. For 7 of 9 sibling job classes that declare their own wiring directly, followWiring=true already worked well — this item was specifically about the case where the wiring is declared on an ancestor. Fixed by also following REFERENCES edges declared on EXTENDS ancestors (transitively) when traversing followWiring from a concrete class, the same way method calls already resolve inherited members. Materialized as a synthetic edge (resolvedVia: 'INHERITANCE'), same class-level over-approximation as the rest of the CHA-style resolution (doesn't know whether the subclass overrides the specific method that declares the wiring). Covered by JavaInheritedWiringIT.subclassInheritingWiringAlsoResolvesIt (classDeclaringItsOwnWiringResolvesToday is the control test).

  • 32. Resolve DB-table access on the Repository/Entity itself, not only on the calling Logic class (found 2026-07-07/08, PUR pur-batch re-evaluation, done 2026-07-09) — db_accesses/sql_statements correctly resolve a table when a Logic class calls a repository method (via: "RiskLogic" etc., per item J1a/J1b), but db_accesses(RiskRepository)/db_accesses(RiskEntity) and dbTables in their own module_digest stayed empty — even though the repository interface statically carries its generic entity type (AbstractPurRepository<RiskEntity>/IRiskRepository) and the entity carries its @Entity(name = TABLE_NAME). Fixed by attaching the resolved table directly to the Repository/Entity module's own dbTables/ db_accesses, tagged mode: "DECLARES" (not a read/write, just "this module maps to this table") — besides the existing READS/WRITES from called functions. Lets a reference implementation's Repository/Entity pair reveal its table directly instead of requiring a call chain through whichever Logic class happens to use it first. Covered by JavaRepositoryOwnTableIT.repositoryOwnDigestListsItsTableWithoutAnyCaller and .entityOwnDigestListsItsTableWithoutAnyCaller.

  • 33. Hook-contract query for a base class (found 2026-07-07/08, PUR batch-job family comparison, done 2026-07-09) — when comparing/porting a template-method-style family (e.g. AbstractSteuertabellenProcessingStep with hooks like validate()/writeEntities()/clearTable()), there was no way to ask the graph which of a base class's methods are abstract (subclass must implement), final (fixed, subclass must not override), or a plain overridable default — this had to be read from the base class source by hand every time. Added GET /modules/{name}/functions?kind=abstract| final|overridable, sourced from the modifiers already visible to the parser (each function also carries a kind field in the unfiltered response, null for a Natural subroutine or a constructor). Turns "what must a new sibling implement" into one call instead of reading the whole abstract class. Covered by JavaFunctionKindIT's kindAbstractReturnsOnlyValidate/kindFinalReturnsOnlyCommit/ kindOverridableReturnsOnlyLog (unfilteredListsAllThreeMethodsToday is the control test).

  • 34. Bulk function_overrides for all abstract methods of a base class (found 2026-07-07/08, PUR batch-job family comparison, done 2026-07-09) — function_overrides(base, function) took one method name per call; profiling how an entire family (5+ hook methods × 5+ sibling classes) implements its contract needed one call per hook method. Added an optional function param — when omitted, GET /modules/{name}/functions/overrides returns overrides for every abstract method of base at once, grouped by method then by declaring subclass. Complements item 33 (which methods exist) with "how does each sibling implement them". Covered by JavaBulkFunctionOverridesIT.bulkEndpointGroupsOverridesByHookMethod (perMethodOverridesAlreadyWorkForEachHook is the control test).

  • 35. modules?extends={name} filter (found 2026-07-07/08, PUR batch-job family comparison, done 2026-07-09) — listing every concrete subclass of a base class previously required reading callers(base, ...) and filtering for the EXTENDS edge kind by hand. Added a direct filter on the existing GET /modules listing (?extends=AbstractSteuertabellenProcessingStep), making "show me every existing implementation of this pattern" a one-line query, consistent with the existing ?moduleKind=/?sourceFile= filters. Restricts to direct EXTENDS subclasses only (one hop, not transitive). Covered by JavaModulesExtendsFilterIT.extendsFilterListsOnlyDirectSubclasses (unfilteredListingContainsAllFourClassesToday is the control test).

Integration & ops

  • 9. docker-compose.yml committed + wired into the README (2026-07-12) — the root docker-compose.yml (Neo4j 5 + the ac-code-server container, built from src/main/docker/Dockerfile.jvm) is tracked in git and driven by the deploy.sh wrapper (up/stop/logs/cli). The README's "Getting Started" now leads with the Compose/deploy.sh full-stack path and documents a dev-mode variant that runs only the Neo4j service from Compose (docker compose up -d neo4j) alongside mvn quarkus:dev, replacing the previous manual docker run neo4j:5 command.

  • 5. MCP endpoint (HTTP/SSE) (2026-07-07) — implemented the previously empty ac-code-server/.../mcp package as a Quarkus MCP server (io.quarkiverse.mcp:quarkus-mcp-server-sse 1.9.1, augments cleanly against Quarkus 3.36.2), exposing the API to agents over HTTP/SSE at /mcp/sse (server name agenticcode). Two @ApplicationScoped tool beans mirror the REST surface by delegating to the same GraphRepository/ProjectIngestService — no query logic duplicated. McpQueryTools (22 read tools, all @Blocking, returning the same concise JSON as REST via McpSupport): list_projects, list_modules, module_digest, module_context, module_functions, module_data_structures, module_dispatch_table, module_columns, callers, callees, call_tree, db_accesses, sql_statements, data_structure_fields, db_table_columns, search_identifier, search_value, variable_reads, variable_writes, flow_forward, flow_backward, field_flow. McpIngestTools: ingest_all, ingest_module, ingest_call_graph. Deliberately read-only + ingest — destructive project create/update/delete/clearAll are not exposed. Errors reuse the REST {error,code,details} shape as MCP tool errors (PROJECT_NOT_FOUND, ROOT_NOT_SET, INVALID_TYPE, …); the deep-query tools return the same deep-ingest hint (pointing at ingest_module) when a module isn't deep-ingested. Extracted the project-root guard shared with REST into ProjectRootResolver so the validation can't drift between transports. New McpToolsIT drives the whole pipeline through tools only (connect → ingest_all → list_modules/list_projects

    • a PROJECT_NOT_FOUND error case); all 58 REST ITs still green after the resolver refactor. Docs: agent-api-usage-ac-implementation.md gained an "Access via MCP" section.
  • 30. Version number in startup log / API / MCP (done 2026-07-08) — agenticcode.version is a manually-bumped release counter in application.properties (deliberately independent of the Maven project version, which stays a build/packaging concern). Single source of truth (VersionInfo), exposed three ways: an explicit INFO startup log line (VersionLogger), GET /api/version, and the MCP version tool — plus the MCP protocol's own server-info.version handshake field, which was previously hardcoded to 1.0.0 (already drifted from the real build) and now references agenticcode.version instead of duplicating it.

  • 36. ac-cli parity with the REST/MCP API (done 2026-07-09) — the CLI was missing wrappers for 12 endpoints that already had MCP tools: modules, module-data-structures, dispatch-table, digest, function-overrides (bulk + single), search-value, search-annotation, inspect-node, node-source, module-source, and ingest call-graph. Added all as new ac-cli commands. Also added a hard rule (CLAUDE.md §0) requiring every new/changed REST endpoint to ship with both an MCP tool and an ac-cli command in the same change, so the three stay in sync going forward.

  • 37. ac version command (done 2026-07-09) — closes the last REST/MCP-vs-CLI gap: GET /api/version already had an MCP version tool but no CLI counterpart. deploy.sh now stamps ac-cli's bundled agenticcode.properties (version=) from agenticcode.version (the same release counter VersionInfo uses) at build time (stamp_cli_version, called from both build_all and build_cli_only), so a jar built outside deploy.sh reports dev instead of a stale number. ac version prints both the ac-cli and connected server's version and exits 1 with an error if they differ (skipped when the CLI is a dev build). The interactive shell also prints the ac-cli version on startup and the same mismatch error (VersionCheck, shared with the version command).

  • 38. Log every MCP tool call (name + arguments) (done 2026-07-09) — new @McpLogged/McpLoggingInterceptor CDI interceptor pair (com.agenticcode.codeserver.mcp), applied at class level to McpQueryTools and McpIngestTools. Logs INFO com.agenticcode.mcp: MCP tool call: {toolName}({arg=value, ...}) for every @Tool invocation, reading the tool name off @Tool.name() and real parameter names via reflection (relies on the existing -parameters javac flag). Verified live via McpToolsIT.

  • 54. Source-text regex search (2026-07-14) — GET /api/projects/{p}/search/source?regex=&limit=&ignoreCase= greps the source text of all modules from disk (deduped by file, bounded by a match limit → truncated), returning {module, sourceFile, lineNo, line}. Case-insensitive by default (legacy Natural). SourceSearchService reuses the parsers' UTF-8/ISO-8859-1 decoding; 400 INVALID_REGEX on a bad pattern. REST + MCP (search_source) + CLI (ac search-source <regex> [--limit --case-sensitive]) in sync; Web-UI explorer gains a names / source mode toggle with a results list (each hit opens its module at the line). Complements search_identifier (declared names). Covered by e2e/module-explorer.spec.ts.

  • 56. search_identifier sigil-insensitive name match (2026-07-14) — the name search now ignores a single leading Natural sigil (# user, & AIV, + GDA) on both sides, so name=K-OUT-MAX finds the declared #K-OUT-MAX (and #K-OUT-MAX still works). Implemented in the shared query path (CypherQueries.SEARCH_IDENTIFIER strips left(n.name,1); GraphRepository strips the search term) so REST + MCP (search_identifier) + CLI (ac search-identifier) inherit it by construction; exact matches for sigil-less names (Java, tables) are unchanged. Covered by AnalysisResourceIT.searchIdentifierIsLeadingSigilInsensitive. (Prompted by the Web-UI popover, where Ctrl-click always captures the # but the global search box did not.)

  • 52. Function-level callers ("who PERFORMs this subroutine") (2026-07-15) — new GET /modules/{name}/functions/{function}/callers returns the FUNCTION nodes that CALLS the named subroutine/method — the intra-module PERFORM sites (Natural) or cross-class method callers (Java) — with call-site lineNos, in the same CallRefResponse shape as module-level /callers. CypherQueries.FUNCTION_CALLERS (reuses toCallRefRow/buildCallRefResponse) + MCP function_callers + CLI ac function-callers <module> <function>. Covered by AnalysisResourceIT.functionCallersListsPerformSites (DYNAIDX: INIT-TBL ← DISPATCH; a subroutine called only from the module main body has 0 function callers). UI wiring: the identifier popover's definition-click picker now attributes each PERFORM site to the calling subroutine via function_callers (useFunctionCallers), falling back to "module body" for top-level PERFORMs — the complete site list still comes from callees?scope=internal so body-level PERFORMs (e.g. INITIALIZATION) are not lost. Covered by e2e/identifier-popover.spec.ts (ADD-XML-LINE ← GEN-XML-LINE attribution).

  • 53. search_identifier scoping (sourceFile/module filter) (2026-07-15) — added optional sourceFile=<relpath> and module=<name> filters to search_identifier so a caller can ask "this name, in this module" directly (previously the paginated, project-wide search could push the module-local declaration off the page). module= resolves to the module's source file via an EXISTS { MATCH (:MODULE {name}) WHERE .sourceFile = n.sourceFile } subquery; sourceFile= filters exactly. REST (?sourceFile=&module=) + MCP (search_identifier args) + CLI (ac search-identifier --module --source-file, plus a previously-missing --type) in sync. Covered by AnalysisResourceIT.searchIdentifierCanBeScopedToAModule.

Tests & tooling

  • 8. ac-cli test coverage (2026-06-17) — module had zero tests. Added VariableAccessQueryTest (query-string builder) and CommandResolutionTest (resolveProject() precedence, apiClient() URL normalization, picocli option-parsing). 12 tests.
  • P1-p. Fix pre-existing AnalysisResourceIT failures (2026-06-16) — six IT tests were red on main but unnoticed because *IT is excluded from the default surefire run. (1) Five tests double-encoded the # in the path; switched to RestAssured pathParam so # is encoded exactly once. (2) transitiveDbAccessesAndSqlStatementsIncludeCalleeResults passed inline Natural source as the classpathResource arg; added an ingestInline(...) helper.
  • 11. Natural/Java parser construct coverage review (2026-06-15) — checked NaturalParser against the CLAUDE.md priority construct list and the WGEAGB0S fixture set (26 files): all 7 priority constructs are covered. Found and fixed one gap: DECIDE FOR/DECIDE ON (see P1-f). Other constructs (ESCAPE, RESET, EXAMINE, COMPRESS, SEPARATE, WRITE) have no graph-relevant effects and are intentionally out of scope.

Web UI — code understanding & navigation

A React/TypeScript web UI (ac-ui/) that makes the AgenticCode graph navigable for humans — primary use case: understanding legacy Software AG Natural in order to migrate it to Java (who calls a module, what it calls incl. dynamic CALLNAT, which fields flow where, which DB tables it reads/writes, what breaks on change). Built on the existing REST API; query-driven (never load the whole graph — upms is ~1000–6000+ modules), WebGL graph rendering, OpenAPI-first contract. Full vision/architecture: x-docs/ui-proposal.md. Stack: React + TS + Vite, React Router (URL-driven, shareable deep-links), TanStack Query over the generated OpenAPI client, Sigma.js/graphology for the large call-graph, CodeMirror 6 with a custom Natural language mode + built-in Java. Out of scope: parser/ingest semantic changes (the UI is a consumer) and write access to code (read-only; notes are UI-side). Item 51; M6 (scale & polish) remains open — see roadmap.md.

Backend prerequisites (Milestone M0):

  • 48. OpenAPI spec + generated TS client (backend 2026-07-13) — added quarkus-smallrye-openapi and annotated all REST endpoints (@APIResponse/@Schema, since they return raw Response) so the spec is fully typed; served at /q/openapi (+ Swagger UI in dev). TS client generation (openapi-typescript + openapi-fetch) lands with the frontend unit.
  • 49. Ego-graph subgraph endpoint (2026-07-13) — GET /api/projects/{p}/modules/{name}/graph?depth=&direction=&limit= returns a bounded module-level neighbourhood (nodes + edges, truncated flag, unresolved styling) via a BFS in GraphRepository.egoGraph. REST + MCP (ego_graph) + ac-cli (ac ego-graph) in sync.
  • 50. SPA serving + CORS (2026-07-13) — CORS enabled (quarkus.http.cors.enabled=true), restricted to the Vite dev origins. ingestStatus/ingestDepth now joined into the GET /modules list rows (was only on inspect_node) for the UI's status badges.

Frontend milestones (item 51):

  • M0 Foundation + API contract (2026-07-13) — items 48–50 (backend) + React/Vite frontend (ac-ui/): project picker, virtualized module explorer (filter + search), ingest-status badges, project/module refresh, thin module-detail from context, generated OpenAPI TS client (openapi-typescript
    • openapi-fetch).
  • M1 Navigation (2026-07-13) — whole-file source endpoint (REST + MCP module_source + CLI ac module-source, omit range = whole file); CodeMirror 6 source viewer with a custom Natural StreamLanguage mode + built-in Java, Ctrl/⌘-click name-based identify (popover via search_identifier), caller/callee panels, lazy call-tree (expand-on-demand via callees, cycle-guarded), URL-driven tabs + line deep-links, STALE_SOURCE refresh prompt.
  • M2 Call-graph visualisation (2026-07-14) — interactive ego-graph (Sigma v3 / graphology, WebGL) as a lazy-loaded "Graph" tab: seeds on the current module via the item-49 /graph endpoint, ForceAtlas2 auto-layout, click-to-select, expand-on-demand (one hop, merged), direction/depth/limit controls + truncated badge, unresolved/dispatch (CALLNAT_DYNAMIC)/inheritance edge styling + legend, open-module (disabled for unresolved). No backend change (pure consumer).
  • M3 Migration dossier (2026-07-14) — a lazy per-module Dossier tab (MigrationDossier.tsx) bundling the "porting profile": payload (I/O contract), data structures (expand-on-demand → fields), DB-access matrix (READ/WRITE badges), SQL statements, and the dynamic-CALLNAT dispatch table. Line numbers deep-link into the Source tab (?tab=source&line=). Pure consumer of existing endpoints (payload, data-structures/data-structures/{name}/fields, db-accesses, sql-statements, dispatch-table). Follow-ups: (a) a program-view "ingest +callers/callees" button that deep-ingests the program's whole call-graph neighbourhood (the module + its transitive callers and callees, and via each module's USING/INCLUDE fan-out their data structures) — POST /refresh/{name}?scope=neighborhood (seeds the multi-module BFS ingest with the both-direction ego-graph closure; ProjectIngestService.ingestModules)
    • CLI ac refresh <name> --neighborhood (no MCP tool — refresh is deliberately REST+CLI only); nginx /api proxy timeout raised so the long request survives; (b) source-by-path so a USING data area's field line-links open its own file, not the current module — new GET /api/projects/{p}/source?file=&startLine=&endLine= (reuses SourceSnippetService, rejects root-escaping paths with 400 INVALID_SOURCE_FILE) + MCP file_source + CLI ac file-source; the Source tab gains a ?src=<file> mode with an "included file … back to module" banner.
  • M4 Data-flow & impact (2026-07-14, M4.1–M4.3) — flow-forward/backward / field-flow visualisation + "what breaks?" impact analysis.
    • M4.1a subroutine navigation — Ctrl/⌘-click a subroutine name resolves via the scoped module_functions (reliable despite dozens of same-named subroutines across modules — the global search is capped): clicking a reference (PERFORM) jumps to the definition; clicking the definition goes to its caller(s) — the PERFORM sites from callees?scope=internal lineNos (one → jump, many → a picker). Covered by e2e/identifier-popover.spec.ts (Playwright; 5 cases incl. the upms 60-match reproduction).
    • M4.1 actionable variable matches — identifier-popover non-MODULE matches now expand: show dataType/scope, a jump-to-definition (opens sourceFile:startLine, via the M3 source-by-path infra so a field defined in a USING PDA opens its own file), and lazy reads/writes lists (variable_reads/variable_writes, VariableAccessLocation) with each access site clickable to its sourceFile:lineNo. MODULE matches still navigate.
    • M4.2 flow visualisation — a dedicated Data-flow tab (DataFlowView.tsx) for a selected variable (URL ?tab=flow&flowvar=, seeded on the current module): backward (flow_backward) and forward (flow_forward) DataflowStep chains (indented by depth; variable re-targets the trace, module opens it) and field-flow producer→consumer pairs (field_flow, line numbers deep-link into Source). Reached via a "data-flow →" action in the identifier popover. The flow endpoints auto-deep-ingest the scoped module; a 409 (warm couldn't complete) shows a DEEP_INGEST_REQUIRED prompt wired to the neighbourhood-ingest button.
    • M4.3 impact ("what breaks?") — a per-module Impact tab (ImpactView.tsx): the transitive callers (blast radius if the module changes), via the ego-graph direction=in (useImpactCallers, depth 10 / limit 500), grouped by call distance (direct callers vs. N-hops-away), each dependent clickable to open; unresolved dynamic-dispatch callers shown but not navigable; truncated badge when the node cap is hit.
  • M5 Understanding boosters (2026-07-14, the consumer-feasible parts).
    • M5.1 source + structure side by side — a filterable, collapsible outline (SourceOutline.tsx, module_functions) beside the Source editor; clicking a function scrolls the editor to it (SourceView re-scrolls on line change without a remount — also fixes repeat dossier line-links). Hidden while viewing an included file.
    • M5.2 notes + saved views — per-module migration notes (ModuleNotes.tsx, in Overview) and header saved views (SavedViews.tsx), both persisted to localStorage via useLocalStore (the API is read-only, so annotations stay browser-side).
    • Deferred (need more than a consumer): diff after refresh (backend must expose before/after snapshots) and LLM summaries (no LLM in the stack; the existing module_context.description is shown in Overview as the current summary).
  • M6 — Explorer regex filter (2026-07-14) — the module-list search box accepts a case-insensitive regex over name/sourceFile (falls back to substring while the pattern is incomplete). Covered by e2e/module-explorer.spec.ts. (Rest of M6 — virtualisation/large-graph performance, multi-project, auth, theming, export — remains open in roadmap.md.)

UI-sweep bug fixes (2026-07-14/15)

Found via a UI test sweep + source cross-check.

  • 55. Dispatch table: multi-value DECIDE branch guardValue artifact (2026-07-14) — a DECIDE ON … VALUE 'X', ' ' branch (a real literal + a Natural "also match blank" catch) yielded a dispatch-table guardValue of "X, " (the blank ' ' trimmed to empty, leaving a trailing ", ") instead of X. Fix: NaturalParser.quotedLiterals now drops whitespace-only alternatives (keeping one blank only if a branch has nothing but blanks, so the guard isn't lost); genuine multi-literal branches ('A', 'B') stay comma-joined. Verified live on WGEAGB0S (3 rows @533/537/545 now clean, 0 trailing-comma artifacts across all 11 rows). Covered by NaturalParserTest.multiValueBranchDropsBlankAlternativeFromGuard.
  • Dossier data-structure fields shown out of order (2026-07-14) — the data-structures/{name}/fields API returns fields in graph order ([12,26,40,55,56,41,…]); the UI now sorts by startLine so a data area reads top-to-bottom (this out-of-order display was what made a correct line number look "wrong" earlier).
  • 57. Field format glued to the name (no space) parsed as one token (2026-07-15) — a DEFINE DATA field whose format is attached with no space — e.g. 01 KEY(A1/1:3,1:V), 01 #DELIMITER(A1) (common in NATURAL-CONSTRUCT generated code like CDRANGE) — was stored with the whole token as the identifier name (KEY(A1/1:3,1:V)) and dataType=null, and thus mis-typed as a DATA_STRUCTURE instead of a VARIABLE. 234 of 376 CDRANGE identifiers were affected, making them unsearchable (a search for KEY found nothing) and dataType-less. Root cause: the LEVEL_FIELD name group (\S+) greedily swallowed the attached (…). Fix: name group now stops at ( ([^\s(]+) — Natural identifiers never contain one — in both NaturalParser.LEVEL_FIELD (deep) and NaturalCoarseScanner.LEVEL_FIELD (Tier-1 identifier index), which carried the identical bug. Covered by NaturalParserTest.fieldFormatAttachedWithoutSpaceIsSplitFromName + NaturalCoarseScannerTest.fieldFormatAttachedWithoutSpaceIsIndexedByBareName. Verification note (superseded by #58): the old glued nodes once required a project re-create to clear; since #58 a plain refresh reconciles and purges them.
  • 58. Stale nodes survive refresh — diff-based reconciliation (2026-07-15) — nodes MERGE on (type, name, sourceFile, project), so a re-parse that renames/removes a field (e.g. after the #57 fix) produced a sibling node and the old one lingered — refresh MERGEd but never deleted. On the live ac/legacy graphs this left Tier-1 identifier-index nodes created at project-create time coexisting with the deep parse's nodes (CDRANGE: 460 = 234 stale-glued + 226 clean), inflating counts and returning ghost matches; the only remedy was a project re-create. Fix (roadmap option b): after a batch of fresh nodes is merged, GraphRepository.mergeResults collects each real source file's fresh (canonical) node ids and runs CypherQueries.DELETE_STALE_FILE_NODES — MATCH (n {project, sourceFile}) WHERE NOT n.id IN freshIds DETACH DELETE n. Because MERGE_NODES overwrites node.id to the fresh UUID, a surviving-key node keeps a fresh id (kept) while a renamed/removed field keeps its old id (deleted); moved positional nodes (DB_ACCESS/CONTROL_FLOW) are swept the same way. Gated by a reconcile flag threaded through persist/persistBatch: true for every full parse (whole-root refresh — both deep=true and the default call-graph pass — and the per-module deep refresh/{name}), false for the coarse Tier-1 scan (subset node set; runs only on an empty graph at create, so nothing to delete). sourceFile="" placeholders (shared cross-file targets) are never swept; reconciliation runs at persist-time, before enrichment, so enrichment-derived nodes/edges are unaffected and cross-module edges are rebuilt by finalizeProject. Does not remove nodes for deleted files (still roadmap #43). Covered by RefreshReconciliationIT (rename a declared field on disk → refresh → old identifier gone, new one present); verified live (ac whole-root refresh, 373 files, clean). No REST/MCP/CLI surface change — refresh already exists on all three; the behaviour change is internal.
  • 59. Shared Natural field-declaration tokenizer (2026-07-15) — the LEVEL_FIELD line pattern and its NN [REDEFINE] name (typeSpec) rest decode were duplicated verbatim in NaturalParser (deep) and NaturalCoarseScanner (Tier-1), so #57 had to be fixed twice and the two tiers could drift. Extracted NaturalFieldTokenizer (record NaturalField{level, redefine, name, typeSpec, rest} + static parse(line)), the single owner of the pattern; typeSpec is the raw parenthesised content (what the deep parser stores as dataType, byte-identical to before), with derived baseFormat()/ arrayDims() splitting it and constValue()/initValue() pulling the trailing clause. Both tiers now call the tokenizer instead of holding private copies; node-type/emission policy is unchanged per tier (a pure refactor — the #57 regression tests in both NaturalParserTest and NaturalCoarseScannerTest stay green). New NaturalFieldTokenizerTest (12 cases) covers the edge matrix: format with/without a leading space, A1/1:3,1:V dims, REDEFINE, CONST<>/INIT<>, #/##/&/+ sigils, dimension-only group arrays, and group (no-format) fields. Kills this class of tokenizing bug.
  • 61. Comments parsed as CALLNAT targets (2026-07-16) — found dogfooding WGEAGB0S (upms): the unanchored CALLNAT/CALLNAT_DYNAMIC patterns matched inside comments, so callees/unresolved were polluted by phantom modules. Evidence: tree-wide false unresolved WAS/RESULTED/DOES (from prose like * What the callnat does:, /* …last callnat resulted in end-of-data); WGEAGB0S's ISINDATE call-sites included commented lines 526 & 600 (* CALLNAT 'ISINDATE'); a phantom ADLML02 at copycode lines 18 & 22 (* CALLNAT 'ADLML02'). Root cause: the deep NaturalParser main loop matched on the raw line (no full-line * skip, no inline /* strip), and the coarse NaturalCoarseScanner stripped inline /* but not a leading *. (Anchored patterns — ^\s*PERFORM, ^\s*DECIDE, ^\s*INCLUDE — were already immune; only the unanchored CALLNAT ones leaked.) Fix: both parsers now skip a full-line comment (first non-blank char *, incl. **SAG) and strip inline /* before statement matching. Verified live: WAS/RESULTED/DOES gone (WGEAGB0S tree unresolved 21→17), ISINDATE now [597,626,1267], ADLML02 now [26] (real call only). Covered by NaturalParserTest.commentedAndInlineCommentCallnatsAreNotParsedAsCalls and NaturalCoarseScannerTest.commentedCallnatIsNotIndexedAsACall.
  • 62. Data literals recorded as false MODULE call targets (2026-07-16) — found in a corpus-wide sweep of upms (6311 files): 45 distinct data names (browse keys / codes such as CO-TABLA, COD-EMISOR, AGENT-SP, NAME-DESC-SP) sat in the graph as placeholder MODULE callees, carrying 282 false CALLS edges — polluting callees/call-tree/unresolved across the corpus. Root cause: a CALLNAT <bareword> (a browse key reaching the call site through a copycode/macro argument, e.g. INCLUDE YFRAMGC2 'C-MOD-GET' '"CO-TABLA"' expanding to CALLNAT &3& …) is parsed as a dynamic call and gets a placeholder target. It never resolves — no module of that name exists — and DELETE_DYNAMIC_CALLNAT_PLACEHOLDER_EDGES only drops markers that did resolve, so the false target was kept forever. Fix: a new project-wide enrichment step delete-data-literal-call-placeholders (CypherQueries.DELETE_DATA_LITERAL_CALL_PLACEHOLDERS), running after the dynamic-call resolution and marker cleanup and before stamp-unresolved-placeholders so the unresolved flags see the reaped graph. It reaps a placeholder only when all four hold: (1) the name carries no Natural sigil (#/&/+) — a genuine dispatch variable is always a sigil'd user variable, so CALLNAT #WIF-style markers stay; (2) the name is a real VARIABLE/CONSTANT of the project (DATA_STRUCTURE is deliberately excluded — a module and its interface PDA routinely share a stem name); (3) no real MODULE of that name exists (else it would simply resolve); (4) every incoming CALLS edge is inferred (CALLNAT_DYNAMIC/INCLUDE_MACRO) — a static CALLNAT 'X' is trustworthy, which protects the real external modules RPC-CNTX and USIX081X that collide with field names. Covered by DataLiteralCallCleanupIT (bareword CALLNAT CO-TABLA reaped; sigil'd CALLNAT #DISP marker kept; static CALLNAT 'REALMOD' resolved). Verified live on a clean upms re-create (2026-07-16): delete-data-literal-call-placeholders: 139 ms; rels +0/-274, nodes +0/-42, and the only survivors of the name-based suspect query are RPC-CNTX and USIX081X — both reached exclusively by static CALLNAT edges, i.e. gate 4 protecting them exactly as designed.
  • 63. CALLNAT matched inside a string literal (2026-07-16) — the last of the unanchored-CALLNAT family (#61 = comments, #62 = data literals). The CALLNAT/CALLNAT_DYNAMIC patterns match the keyword anywhere on a line, so prose inside a quoted string fabricated a call. Filed as "low, 1 occurrence"; that was wrong on both counts — a corpus sweep of upms found 3 distinct call sites (each doubled across src/ and generated_src/): PRINT '==> callnat before DREQUFN0' (DREQUDN0) → before; WRITE(#MSG) 'NACH CALLNAT ISINGEAG:' (JE0012N0) → ISINGEAG; #ERR-TYPE := 'Callnat USIA008N' (JA0016N0) → USIA008N. Two of the three are the dangerous class: ISINGEAG and USIA008N are real modules, so the phantom was not placeholder noise but a real → real CALLS edge in the call graph (PROCESS-AGENT-CHANGE → ISINGEAG, MAIN-PART → USIA008N) — and one invisible as a problem, since only before carried unresolved: TRUE while the other two were NULL precisely because a real module of that name exists. Item 62's reaper structurally cannot clean these (its gate 3 requires that no real MODULE share the name), so only a parse-time fix removes them. (The hundreds of 'Start of Callnat'-style lines are harmless: a quote directly follows the keyword, so no identifier matches.) Fix: new shared NaturalLines (the NaturalFieldTokenizer precedent from item 59) holding isInsideStringLiteral/findOutsideStringLiteral plus the formerly duplicated stripInlineComment; both tiers now gate the two CALLNAT matches on the keyword start lying outside a quoted literal — testing the start, not the whole match, is what keeps a real CALLNAT 'MOD' working (keyword outside, argument inside). Natural's '/" delimiters and doubled-delimiter escapes ('IT''S') are honoured. Considered and rejected: handling a /* inside a literal (which stripInlineComment would truncate, leaving an unbalanced quote) — measured 0 such lines that also contain CALLNAT, so it stays out rather than buying speculative complexity. Covered by NaturalParserTest.callnatInsideAStringLiteralIsNotParsedAsACall and NaturalCoarseScannerTest.callnatInsideAStringLiteralIsNotIndexedAsACall, both verified to fail against pre-fix behaviour. Verified live on a clean upms re-create (2026-07-16): the before MODULE node is gone entirely, and the fabricated CALLNAT_DYNAMIC edges into ISINGEAG/USIA008N (2 each) are gone while their genuine static CALLNAT edges (10 and 2) remain — the real → real corruption is removed without touching a single real call.
  • 64. Ingest summary contradicted the graph; dispatch guards lost their alternatives (2026-07-16) — found by dogfooding a deep ingest of WGEAGB0S (upms) and hand-checking every endpoint against the Natural source. Two independent defects: (a) unresolved reported data fields as missing modules. The deep ingest listed MODULE CO-TABLA, MODULE COD-ENTIDAD, MODULE NAME-DESC-SP, MODULE NAME-VALUE-SP — all (A8) fields, no such module anywhere — although the graph was already clean (the scoped enrichment reaped them: rels +0/-7, nodes +0/-5). Root cause: IngestSummary.unresolved is built during the BFS walk from a filename index (ProjectIngestService:581) and never consults the graph, so it applied none of item 62's gates. This made the item-62 fix incomplete: it cleaned the graph surface and missed the summary — an asymmetry item 63 doesn't share, since a parse-time fix stops the ref existing at all. Fix: filter the unresolved refs through item 62's gates — no Natural sigil, reached only by inferred (CALLNAT_DYNAMIC/INCLUDE_MACRO) edges (reconstructed from the parse result's CALLS edges, since DependencyRef carries no provenance), and the name is a real VARIABLE/CONSTANT. That last gate is asked of the graph, project-wide (CypherQueries.DATA_FIELD_NAMES), not of the ingested tree: a first, tree-local attempt fixed only 2 of 4 because NAME-DESC-SP is declared in BMTABBN2.nat, which is outside WGEAGB0S's 164-file tree. The static-CALLNAT gate keeps genuinely-missing modules whose names collide with field names reported (the RPC-CNTX/USIX081X class). Verified live: WGEAGB0S's unresolved list went 18 → 13, all four false positives gone, while the USR* modules, sigil'd #GETSHORT-MODUL, NDBERR and the VDB2-* views all stayed. Covered by DataLiteralUnresolvedSummaryIT (incl. the EXTMOD case: a declared field that is also a statically called missing module must stay reported), verified to fail pre-fix. (b) dispatch guards dropped their VALUE alternatives. guardValue is a single comma-joined string, so VALUE 'GENAGREE-WOUT-SP', ' ' (a Natural "also catch blank") silently lost the blank — quotedLiterals drops whitespace-only alternatives to avoid a trailing ", " — and a multi-alternative branch yields a synthetic "A1, A2" the guarded field never equals. Fix: the parser now also records whenValues, every alternative in source order including blanks, and DISPATCH_TABLE splits it into a real guardValues list. guardValue is untouched, so REST/MCP/CLI/UI keep working (all three surfaces return the record directly, so the field propagates automatically). Encoded with U+001F rather than modelled as a list because AstEdge.properties is Map<String, String>; US cannot occur in Natural source, unlike the comma that made guardValue ambiguous. Verified live: WGEAGB0S now reports 3 branches that also catch blank (0 before), and correctly distinguishes the DECIDE at 531 (blank alternatives) from the one at 602 (none). Covered by NaturalParserTest.multiValueDecideBranchKeepsEveryAlternativeIncludingBlank. Scale, measured honestly: 121 blank-bearing VALUE clauses corpus-wide, 3 real in WGEAGB0S; 0 joined multi-value guards observed in the deep-ingested subset (238 guarded WRITES), so the joined-string flaw is real in the code path but unobserved in practice — a full deep ingest would be needed to settle its true rate.

Qualified-field access & re-ingest reconcile (items 78, 79, 74) — 2026-07-18

Three defects found dogfooding a deep WGEAGB0S audit of upms, all in the Natural field READS/WRITES path, fixed together. All three verified with Testcontainers ITs (independent of the live server); full ac-code-server IT suite green (190/0/0) with no regressions.

  • 78 — group-name-qualified access dropped. A field written qualified by its data area's inner group name rather than the USING name (MSG-INFO.##MSG-PGM := *PROGRAM, where MSG-INFO is the top group of a USING CDPDA-M) produced no edge: NaturalParser.lookupVariable recognised only USING-name qualifiers and fell through to a local-only lookup. Fix: when the qualifier is not a known include and the field is not local, mint a bare included-field placeholder (gated allowIncludePlaceholder && hasIncludes && identifier-shaped) so resolveBareIncludedFieldTargets redirects it to the real PDA field. Test QualifiedFieldResolveIT.groupQualifiedFieldAccessIsCaptured.

  • 79 — qualified read inside an expression dropped. The read side of an assignment/expression/IF tokenised on IDENTIFIER_TOKEN (no dot), so #C := QCPDA.QC-DESC and IF MSG-INFO.##MSG-NR EQ … split into two non-resolving tokens, and IF conditions were not read-scanned at all. Fix: new OPERAND_TOKEN keeps a dotted operand whole; addExpressionReads resolves a qualified token with allowIncludePlaceholder (legitimate explicit field access) but a bare token without (no placeholder for the "expression soup", so item 74's by-name path is not fed); the assign-RHS loop and the IF handler both call it. Test QualifiedFieldResolveIT.qualifiedFieldReadInsideAnExpressionIsCaptured. (The originally-filed "qualifier does not disambiguate a shared name" was disproved by a controlled fixture — the qualifier resolves correctly; the live "unresolved" observation was item-74 stale state.)

  • 74 — USING field misattributed to a same-named local group in another subprogram. WGEAGB0S's BXFRABA4.#L-FORWARD-VIA-PRI was linked not only to the real BXFRABA4.pda but also into unrelated JX0124N6/JX0129N0. (First mis-diagnosed as a stale-edge coexistence surviving re-ingest — a clean recreate disproved that: the misattribution is created fresh in a single pass.) True cause: a Natural USING X references a data area (PDA/LDA/GDA) — always a top-level member of a data-area file — but the placeholder resolver (buildResolvePlaceholderTargetQueries + the two by-name field resolvers) matched USING X to any DATA_STRUCTURE named X, including a 1 BXFRABA4 group those subprograms declare inline, and linked every field under it. Fix: a DATA_STRUCTURE placeholder now resolves only to a real area not owned by a MODULE (NOT EXISTS { (:MODULE)-[:CONTAINS]->(real) }) — a file-level data area, never a program-internal group. Data-area files produce no MODULE node, so real PDAs are kept and inline .nat groups dropped. Test UsingResolvesToDataAreaNotLocalGroupIT.

    • Secondary safeguard (separate scenario, kept): mergeResults also runs DELETE_STALE_RESOLVED_FIELD_EDGES in the reconcile (deep-re-ingest-only) branch — deletes a re-parsed file's prior cross-file READS/WRITES to a VARIABLE/CONSTANT so finalize rebuilds them, so a module that changes which area it includes does not keep the old resolved edge. Scoped to VARIABLE/CONSTANT and cross-file targets; gated on reconcile. Two-phase test QualifiedWriteReconcileIT.reIngestDeletesTheStaleResolvedFieldEdge.

Parallel parse phase (item 24) — 2026-07-18

The ingest parse phase (read file + parse + shell/user-exit metric enrichment) ran as a sequential loop over all candidate files. It is per-file independent — the JavaParser/NaturalParser instances hold no mutable state, and copycodes/userExit are read-only — so ProjectIngestService.ingestRoot now submits one task per file to Executors.newVirtualThreadPerTaskExecutor() (extracted helper parseCandidate). Results are collected in candidate order (the futures list is parallel to candidates), so cross-file duplicate detection and persist order stay deterministic; a per-file parse/read failure is still recorded as an IngestSummary.Failure for that file only. Full IT suite green (192/0/0).

Qualified group-target resolution (item 80) — 2026-07-18

A Natural qualified reference can name a group (a DATA_STRUCTURE), not only a leaf field — e.g. WGEAGB0S writes BGEAGBA0.#P-DESC-NAME, a level-1 group of BGEAGBA0.pda. The qualifier-scoped field resolvers matched a real target of type VARIABLE/CONSTANT only, so such writes/reads stayed unresolved (sourceFile=""). Fix: the two qualifier-scoped (INCLUDES-gated) resolvers (buildResolvePlaceholderFieldTargetQueries + …ScopedQueries) now accept a DATA_STRUCTURE target as well (realv.type IN ['VARIABLE','CONSTANT','DATA_STRUCTURE']). Left the bare-field resolvers (which carry a size(matches)=1 guard) leaf-only, so adding groups cannot turn a previously-unique bare match ambiguous. Test QualifiedGroupTargetResolveIT.qualifiedWriteToAGroupResolvesIntoTheDataArea. (The #MAP-T.* reads that were unresolved in the same WGEAGB0S snapshot are leaves, a separate concern, not covered here.)

Group-qualifier field resolution (item 81) — 2026-07-18

A qualified reference can name a leaf via a group the parser cannot see as a USING member — e.g. WGEAGB0S reads #MAP-T.V-ID, where #MAP-T is a group inside the included BGEAGA01.pda and the leaf V-ID also occurs in dozens of other PDAs. After items 78/79 the read was captured but fell to the bare-field resolver, whose size(matches)=1 guard rightly refused the project-wide-ambiguous leaf, so it stayed sourceFile="". Fix: NaturalParser.lookupVariable now keeps the group qualifier as a qualifierGroup property on the bare placeholder (cached under the compound STRUCT.FIELD key so a group-qualified and a truly-bare reference to the same leaf are distinct); the two bare-included resolvers (buildResolveBareIncludedFieldQueries + …ScopedQueries) then keep only a candidate nested under a group of that name, so the qualifier pins the leaf to the one right PDA. The placeholder is a bare VARIABLE under the module (not a DATA_STRUCTURE area), so the by-name resolvers never see it — no #74-style misattribution risk. Test GroupQualifiedLeafResolveIT.groupQualifierDisambiguatesAnAmbiguousLeaf; full IT suite green (193/0/0).

Read-path CONTAINS bounding (item 75 read path) — 2026-07-19

The runtime read queries traversed a module's statements with an unbounded (m)-[:CONTAINS*0..]->(src). CONTAINS is not acyclic (shared copycode CONTROL_FLOW nodes, parser line-range artefacts), so on a heavy module (ACCNPE01) that expansion blew up and callees/digest/context hung (callees >2 min). Every edge source is a MODULE (depth 0) or a FUNCTION that is a direct CONTAINS child (depth 1) — verified corpus-wide (0 sources deeper; DB_ACCESS parents only FUNCTION/MODULE) — so the traversal is provably equivalent when bounded to *0..1, which cannot walk the cycles. 20 read-side traversals updated (callees, MODULE_HOP_OUT, DISPATCH_TABLE, EGO_NEIGHBORS_*, VARIABLE_ACCESSES, DB_ACCESSES, SQL_STATEMENTS, FUNCTION_CALLERS, SEARCH_BY_VALUE, fieldFlow, BUILD_CALLS_MODULE, and the *_FOR_MODULES batch variants). Query-only change (no recreate). Guard ReadPathBoundedTraversalIT (EXPLAIN asserts no unbounded CONTAINS expand); full IT suite 193/0/0. Live (v76): ACCNPE01 callees

2 min → 0.16 s, digest >10 s → 3.3 s, context >10 s → 3.1 s.

Version bump moved from deploy into the build — 2026-07-19

agenticcode.version (the integer build counter reported by /api/version and used by ac version for stale-CLI detection) used to be incremented by manage-ac.sh deploy (shell bump_version), so a plain mvn clean install never bumped and only a deploy did. It now lives in the Maven build: a new ac-mvn-plugins mojo bump-version, bound to ac-code-server's generate-resources phase, increments the counter and stamps the same number into ac-cli's agenticcode.properties. Binding to generate-resources (before process-resources) means the freshly built jar already reports the bumped number, keeping the jar's baked version and the CLI stamp in lock-step. The mojo only acts when a requested goal is package/install/deploy (read from MavenSession.getGoals()), so mvn test, mvn compile and quarkus:dev do not bump; deploy bumps because it runs a full install. Every such build bumps unconditionally (no source-change gating — high numbers are harmless). manage-ac.sh's bump_version was removed; stamp_cli_version is kept only for the CLI-only rebuild path (cli), which syncs the CLI to the current server version without bumping. Logic unit-tested (BumpVersionMojoTest, 7/0/0); live-verified: mvn generate-resources left the counter unchanged, mvn package bumped 77→78 and stamped ac-cli 78, and target/classes/application.properties carried 78 (timing correct).

Manual override for unresolvable dynamic CALLNAT targets (item 82) — 2026-07-19

Surfaced by the WGEAGB0S deep-API audit: the dynamic-CALLNAT resolvers cannot recover every target — e.g. YGEAGGNH's CALLNAT #GETSHORT-MODUL, whose name is assembled by MOVE 'YGEAGKEY' TO #GETSHORT-MODUL + MOVE 'GN0' TO SUBSTR(#GETSHORT-MODUL,6,3) → YGEAGGN0 (a real, ingested module). Such a site is left an unresolved CALLNAT_DYNAMIC placeholder. A human or agent can now pin it via a REST/MCP/CLI override, keyed by (originFile, lineNo) and fanning out to one or more target modules (deliberate branches). Stored as a :DynamicCallOverride node deliberately not labelled :AstNode, so DELETE_PROJECT_NODES (which matches :AstNode {project}) never wipes it on a refresh; an enrichment step apply-manual-dynamic-callnat (after the auto resolvers, before the placeholder cleanup) re-applies every override automatically — MERGEing a CALLS edge (callKind=CALLNAT_DYNAMIC, resolvedBy='manual') to each target and flagging the placeholder marker manualHidden (kept, not deleted) so the read queries suppress it. Precedence: an override applies only while the site is still unresolved; once an auto-resolver resolves it, the override is skipped and listed obsolete. Reset (DELETE, one site or all) clears the manual edges and un-hides the placeholder inline — no refresh needed, no fragile src-node tracking. Target validation rejects a non-module name (400 UNKNOWN_TARGET) so a typo can't reintroduce a phantom.

Bug B (same audit) fixed alongside: callees/digest now surface the unresolved flag the graph endpoint already exposed, so an unresolved dynamic target is machine-distinguishable without inspecting sourceFile. API surface (REST + MCP + CLI, in sync): GET /dynamic-calls/unresolved, GET/POST/DELETE /dynamic-calls/overrides; MCP list_unresolved_dynamic_calls, list_dynamic_call_overrides, set_dynamic_call_override, reset_dynamic_call_override; CLI ac dynamic-calls unresolved|overrides|set|reset. The agent system prompt (agent-api-system-prompt.md) now requires an agent to investigate and pin an unresolved dynamic call rather than report a dead end. Covered by DynamicCallOverrideIT (6/0/0): resolve, multi-target, unknown-target rejection, inline reset restore, obsolete precedence, and override survival across a refresh?deep=true.

MCP surface removed (item 26) — 2026-08-04

The MCP server was removed from the product. It existed as a second transport in front of the same services the REST API already exposes: McpQueryTools (770 LoC, 38 @Tool methods), McpIngestTools (2 tools), plus McpSupport/McpLogged/McpLoggingInterceptor — 973 LoC of main code and the 145-LoC McpToolsIT, all deleted, together with the quarkus-mcp-server-sse/-test dependencies (root POM dependencyManagement + mcp-server.version property + ac-code-server POM), the quarkus.mcp.server.server-info.* properties, and the repo-root .mcp.json client registration.

Rationale. Item 26 (MCP session reliability) never became fixable from this codebase: calls failed with "the first message from the client must be initialize: tools/call", at first intermittently and by 2026-08-02 from the first call of a session onwards, while the equivalent REST endpoints answered normally. Rather than carry a second, unreliable transport plus the CLAUDE.md rule that every REST change be mirrored into an MCP tool, the surface was dropped.

No capability lost. All 40 tools were verified to have a REST twin before deletion — including the non-obvious ones: version → GET /api/version, ego_graph → GET /modules/{name}/graph, file_source → GET /source?file=, project_loc → GET /loc, and the three dynamic-call-override tools → GET/POST/DELETE /api/projects/{p}/dynamic-calls/overrides. The MCP layer held no logic of its own; McpSupport only serialized service results into the same JSON the REST resources return.

Docs & rules updated. CLAUDE.md (sync rule now REST + ac-cli; tool priority now REST → CLI → grep/Explore; architecture diagram, project structure, ADR table), README.md (architecture diagram, tech stack, endpoint list, "Agent usage" section replacing "MCP"), x-docs/agent-api-system-prompt.md ("Access via MCP" section and its REST↔tool mapping table dropped), x-docs/agent-api-usage-ac-implementation.md (tool names replaced by their endpoint paths throughout; filename kept to avoid breaking ~6 inbound references), x-docs/agenticcode-ueberblick.md, x-docs/presentation.md, prompts/CLAUDE.md, both prompts/*-deep-api-audit.md, and .claude/settings.local.json. Historical MCP mentions in this file and in earlier roadmap.md entries are left intact as record.

Closed roadmap items (moved out of roadmap.md 2026-08-28)

roadmap.md states in its header that it tracks only open work, but 59 completed items had been left standing in it. They are moved here verbatim — no summarising, no rewriting — grouped by the section they came from and in their original order. Two sections whose content was entirely narrative about finished work moved with them.

One correction was made in transit: a second item had also been numbered 91 (search/identifier ?priorityModule=) and is now 151; the original 91 (copycode provenance on db-accesses/workfile-accesses/sql-statements) keeps its number.

High priority — Agent API gaps found in the UPMS→PUR reengineering (2026-08-18)

Six gaps reported after a week of daily agent use on upms / pur / app, ordered by the time they cost. All probes below were run against a freshly refreshed graph and are reproducible as written. Items 125 and 127 were the two said to make the API return a wrong answer rather than a missing one, which is why they led the list. 125 is fixed (2026-08-18); 127 was retracted the same day — its probe was not ambiguous, so the answer had been correct all along (see the retraction below). All six are closed: 125 and 127 (2026-08-18, 127 retracted), 126 and 129 (2026-08-18), 128 and 130 (2026-08-19).

  • 125. search/identifier did not match a type declaration's short name — and silently ignored contains ( fixed 2026-08-18)

    Symptom. The class is in the graph, but the search endpoint cannot find it.

    GET /pur/search/identifier?name=PartnerOtherUpdateLogic                 → []
    GET /pur/modules/com.uniqagroup.pur.logic.upms.partner.PartnerOtherUpdateLogic/context
                                                                            → the module, 8 functions
    GET /pur/search/identifier?name=PartnerOtherUpdate&contains=true        → []
    

    Only methods, fields and variables are indexed. This breaks the single most common question an agent asks before writing anything: does this name already exist? The only workaround is to guess the FQN and try /modules/{FQN}/context — and then a 404 means "guessed wrong" or "does not exist", which are not distinguishable.

    &contains=true works on /search/value but is accepted and ignored here — it returns [] rather than a match or an error. "Every identifier containing upd" was the central question of a rename pass and could not be answered at all.

    Diagnosis (2026-08-18). The reported cause was wrong in a way worth recording: type declarations are indexed — ?name=com.agenticcode.neo4jstore.graph.GraphRepository returns the MODULE node. SEARCH_IDENTIFIER matched n.name alone, and a Java module's name is its FQN (item 117) while the short form lives in n.simpleName. So the index was complete and only the short key was missing — which produced exactly the reported symptom. contains was not a parameter of the endpoint at all, so JAX-RS dropped it silently.

    Delivered. SEARCH_IDENTIFIER matches the sigil-stripped n.name or n.simpleName; ?contains=true (CLI ac search-identifier --contains) switches both to a case-insensitive substring match as on /search/value, and contains without a name is 400 MISSING_NAME rather than a full node dump. Each IdentifierMatch now carries simpleName and moduleKind (CLASS/INTERFACE/ENUM/RECORD, PROGRAM/SUBPROGRAM for Natural, null for non-MODULE hits) — the discriminator the item asked for. The unindexed substring scan is bounded to offset+limit rows via a $scanCap on the query's LIMIT, chosen so paginate()'s existing "limit <= 0 means unlimited" rule is untouched. Covered by IdentifierTypeDeclarationIT (8 tests); x-docs/agent-api-usage-ac-implementation.md updated.

    Measured after deploy (2026-08-18). pur?name=PartnerOtherUpdateLogic → the class (was []); pur?name=Builder&type=MODULE → 1 match, AuthorizationInterceptor → 2 (which is what item 127 could not check). The substring scan on upms (454,300 nodes) is 1.2–2.3 s warm — usable, but not free: it is a label scan, so keep a limit.

  • 126. /projects carried no ingest metadata, so "not found" was never evidence (fixed 2026-08-18)

    Symptom.

    GET /api/projects
    {"name":"upms","root":"…","excludeDirs":["target"],"language":"java", …}
    

    No ingestedAt, no file count, no failure list. The upms project is known to be incompletely ingested, but nothing in the API says so — so every negative answer has to be cross-checked against the file system, and the consuming project has had to write that rule into its own agent instructions. The same blind spot hides staleness: after a code change there is no way to tell whether an answer predates the edit short of running a refresh.

    Delivered. The (:Project) shell records what the last whole-root ingest did, returned as a nested ingest object on GET /projects and on the new GET /projects/{p} (CLI ac project show): ingestedAt, mode (tier1/call_graph/full), filesExamined, filesPersisted, filesFailed

    • failures (capped at 200 with an explicit failuresTruncated, so a short list is never read as the whole story), durationSeconds, serverVersion. Three deliberate semantics, each pinned by a test in ProjectIngestMetadataIT (6 tests):
    • ingest: null means never recorded, not "ingested nothing" — a zero-filled record would state a fact nobody measured, which is the same class of error as the one this item reports.
    • Only whole-root passes write it. A by-name refresh/{name}, a deep ingest or a fan-out warm ingests real files but walks a fraction of the tree; letting one move ingestedAt would advertise the whole project as freshly walked because one module was deepened.
    • ingestedAt is not a freshness guarantee — it dates the walk, not the match against disk. That is item 129's hash work; GET /{p}/source?file= (409 STALE_SOURCE) is today's real check.

    Kept apart from UPDATE_PROJECT (user config, COALESCEd) so neither can clobber the other, and ProjectInfo carries the state in a nested component rather than flat — the same record is the input to an ingest, and a run must not be handed the state it is about to replace.

    Not delivered (deferred, not forgotten). "Ideally every response carries the project's ingestedAt" is a cross-cutting envelope change — responses are bare arrays today. It belongs with item 130's scope-echo work, which rewrites the same envelopes.

  • 128. No endpoint for "every reference site of this symbol" (fixed 2026-08-19)

    Symptom. callers gives module-level call edges with line sites. There is no way to ask for all occurrences of a name — import, type position, field type, annotation argument, test reference. A rename of four classes had to be scoped by unioning the sourceFiles of four callers responses plus several search/identifier calls, and two affected files (PartnerControllerTest, AbstractUPMSServiceAcceptanceTest) were only found because the agent already knew they existed.

    Delivered. GET /search/references?name=&kind= (CLI ac references) returns {sourceFile, lineNo, kind, inModule, target} per mention, unioning CALL, IMPORT, TYPE (declared field/parameter/return), ANNOTATION, EXTENDS, IMPLEMENTS, INJECTS, CLASS_LITERAL and Natural's INCLUDE. The name may be the identity or the short form (item 125); an unknown kind is 400 INVALID_KIND, never an empty list. SearchReferencesIT, 9 tests.

    The Java parser now emits the mention edges it never had — imports, declared type positions and annotation usages — measured on pur as ~26.6k imports and ~13.4k declared-type positions against ~194k existing edges, so roughly +20%.

    Measured ingest cost (2026-08-19 deploy): deep refresh went ac 19s → 106s, app 37s → 106s, pur 29s → 125s, upms 1230s → 1496s. A 3-5x slowdown on the Java projects is the price of this index, and worth re-reading before assuming it is free.

    Verified live on pur: ?name=PartnerController returns 65 sites (41 CALL, 20 INJECTS, 2 IMPORT, 2 TYPE) — including AbstractUPMSServiceAcceptanceTest, one of the two files this item reported as "only found because the agent already knew they existed".

    Two decisions worth keeping:

    • A new MENTIONS edge type, not REFERENCES. Reusing REFERENCES regressed the call graph — callers/callees/call-tree/ego-graph follow it as wiring, so an import surfaced as a caller (javaCrossClassCallGraphSpansFiles caught it: edgeKind: REFERENCES where a method call belonged). A separate type keeps "mentions" out of "calls" by construction rather than by remembering to filter it in every existing query.
    • Imports are filtered to project-internal ones (sharing the importing file's first two package segments), and the JDK/framework name list moved to ac-parser-core (ExternalTypeNames) so the parser and the ingest apply the same exclusion. Otherwise every java.util/framework import mints a placeholder node that is created, persisted and swept again on every ingest.

    Known gaps, stated rather than hidden: local-variable types and generic type arguments are not indexed (List<Target> records List); same-package references have no import, so within one package the index rests on declared-type positions; Natural has no import or type-position concept and contributes call/include/inheritance kinds only. And the index is only as complete as the last re-parse — existing graphs need a refresh before it is populated.

  • 129. Refresh was all-or-nothing (fixed 2026-08-18)

    Symptom. Twelve changed files; POST /pur/refresh examined 2 981. The feedback loop after a code change is minutes, which discourages the verify-after-change step the workflow depends on. Worse, an aborted deep refresh leaves the graph half-updated with no marker, so every later query silently answers from that state — a hazard the consuming project documents in its own instructions because the API does not surface it.

    Delivered, in three independent pieces (CLI in step: ac refresh --paths / --changed-only):

    1. POST /refresh?paths=a/B.java,c/D.java — targeted deep re-ingest of the named files plus their dependencies, on the existing ingestFiles BFS. Paths matching no file come back in unresolved (a typo'd path must not read as a successful refresh), and paths escaping the root are refused by the same guard the source endpoints use. It deliberately skips the deleted-file sweep and does not move ingestedAt — both are only meaningful for a whole-root walk.
    2. ingest.incomplete on /projects — set before a whole-root pass, cleared on success, so a run that never finished stays flagged. It cannot self-heal (a killed process clears nothing), which is the correct direction to fail: a false "incomplete" costs one refresh, a false "clean" costs trust in every answer.
    3. ?changedOnly=true — hash-based skipping, opt-in rather than the default the item asked for.

    Why changedOnly is not the default. Implementing it that way would have silently corrupted the graph in three ways, all found while validating rather than after shipping:

    • Natural copycodes are inlined at parse time. A module whose .cpy changed parses differently while its own hash is unchanged — so it would be skipped and keep a stale expansion, with nothing in the graph to indicate it. A changed copycode now disables skipping for the whole run (tested). This alone rules out "incremental by default" for upms, where 137 modules share one copycode.
    • Duplicate detection groups the files it parsed — a subset can only confirm duplicates among changed files. Markers are never cleared (the query only MERGEs), so this loses discovery, not recorded facts.
    • User-exit LoC annotation is re-stamped only on re-parsed files.

    Two bugs found verifying against the live server (2026-08-19), both fixed with regression tests:

    • ?paths=pom.xml was accepted rather than reported — a path that exists but is not an ingestible source file passed the guard, was listed as examined, and left unresolved empty. That is exactly the "I ingested 2 of your 3 files" invisibility the list exists to prevent.
    • The copycode stand-down fired on a Java project: ac carries .cpy files as Natural test fixtures that its Java walk never ingests, so they had no stored hash, counted as changed, and disabled skipping entirely — 487 of 509 unchanged files were re-parsed, i.e. changedOnly did nothing at all. The guard is now Natural-only (an unknown language keeps the conservative path).

    Honest about the win: enrichment is project-wide and still runs in full, so this cuts parse+persist only. The premise in the symptom above has also moved — a full deep pur refresh measured 29 s on 2026-08-18 (upms: 20 m 30 s, the one project where this really pays). Covered by TargetedRefreshIT (8 tests).

    Not delivered: surfacing incomplete on every query response, as opposed to on /projects. That is the same envelope change item 126 deferred, and it belongs with item 130's scope echo.

  • 130. Two smaller ones: REST-path routing, and project scope invisible in the answer (fixed 2026-08-19)

    Delivered.

    • GET /rest-endpoints (CLI ac rest-endpoints) returns the composed path, verb, declaring class and handler, with ?module= to narrow. The parser now persists restPath (class and method) and httpMethod — annotations were stored by name only, so the path string was not in the graph at all. A @Path written as a constant reference resolves the same way a JPA @Column name does; a method with no verb annotation is not an endpoint. Narrow by choice: this is two properties, not a general annotation-argument store, which would be a JSON blob per node against item 111d's grain.
    • Scope and freshness are echoed as response headers — X-AC-Exclude-Dirs, X-AC-Ingested-At, X-AC-Ingest-Incomplete — on every project-scoped response. This also closes the "every response carries ingestedAt" half that items 126 and 129 deferred.

    Why headers rather than body fields. Most endpoints answer with a bare JSON array (db-accesses, functions, search/identifier, …). Adding a scope field there means restructuring array → object, which breaks the web UI's generated client, the CLI printers and any agent that indexes [0] — too high a price for a hint. The cost of the choice is stated in the docs rather than hidden: an agent reading only the body will not see them. The project shell is cached ~10 s so the headers add no query per request, and the ingest path invalidates that cache explicitly — a stale incomplete=false during a running refresh would point exactly the wrong way.

    RestEndpointsIT, 11 tests, including the headers riding on a bare-array response.

    Three bugs that only real data exposed (found verifying against pur after the 2026-08-19 deploy, fixed the same day):

    • //file — the path halves were joined with a single replace('//', '/'), and replacement is non-overlapping, so '///upload' collapsed to '//upload'. Each half is now stripped of its own leading/trailing slash before joining.
    • 183 duplicate rows of 436 — the graph holds more than one CONTAINS edge between the same module and function (see item 75), so every such endpoint was emitted twice. RETURN DISTINCT.
    • POST / — JAX-RS inherits @Path from a base class or interface, which this codebase uses heavily (AbstractFileTransferUiSvc). Those endpoints reported the bare /: a wrong answer, not a missing one. The class path now falls back to the nearest ancestor's @Path.

    Worth recording that none of the three was visible in the fixture-based tests that passed first time — they were found by looking at the real output on pur and disbelieving it. A fourth followed from the same habit: three @RegisterRestClient interfaces were listed as served endpoints, which states the traffic's direction backwards. They now carry outbound: true rather than being dropped, since "what does this application call out to" is a real question.

    Not changed: the projects still have different excludeDirs (app: ["test","target"], pur/ac: ["target"], upms: []). The header makes the asymmetry visible; silently normalising someone's ingest scope to make answers look consistent would be the wrong fix.

High priority — the silent-truncation defect of item 103 is still open on the three

search/* endpoints (2026-08-19)

  • 131. search/identifier, search/value and search/annotation silently truncated at the default limit=50 — a bare array with no total and no truncated flag (fixed 2026-08-19)

    This is item 103's own follow-up, in 103's words: "If a truncated flag is wanted anyway, it should be a separate item covering all paginated endpoints, not just these two." The fix for 103 was deliberately narrowed to db-accesses/workfile-accesses; the three search endpoints match 103's criterion exactly and were left as they were.

    Symptom. Measured on pur, freshly refreshed:

    GET /pur/search/annotation?name=Immutable              → 50 rows, ends at FolderEntity
    GET /pur/search/annotation?name=Immutable&limit=500    → 95 rows
    

    The response is a bare JSON array, and there is no header either — checked for X-Total-Count, Content-Range and Link: none is sent. The caller cannot tell the two answers apart from the outside.

    This produced a wrong finding in a real audit, not a hypothetical one. TASKS.md 58 of the UPMS→PUR project recorded that 17 of 35 read-only tables had lost their @Immutable annotation, and concluded that a generator run would silently break optimistic locking on them. The entry was derived from the truncated page. Re-checked on 2026-08-19 against all 114 rows of Tables_meta.csv in both directions: 94 with Writable=false carry @Immutable, 20 with Writable=true do not, zero deviations. 16 of the 17 named entities are precisely the rows the cut dropped; the 17th, FolderEntity, is the hit on the boundary. The finding cost a day and was pure artefact.

    Cause. AnalysisResource.searchIdentifier (:872), searchByValue (:789) and searchAnnotation (:812) all pass effectiveLimit(limit), i.e. DEFAULT_PAGE_LIMIT = 50 (:176-180). uncappedLimit (:198) — added for item 103 with a Javadoc that states the criterion as "endpoints where a silently-capped default is actively harmful… a bare JSON array with no total and no truncated flag" — is not applied to any of them.

    Why these three and not every paginated endpoint. The cap only bites where the query means enumerate, and only search/annotation does at scale. Measured spans:

    | Endpoint | Sample | Range | | :--- | :--- | ---: | | search/identifier | 10 queries | 1 – 41 | | search/value | 5 queries | 0 – 4 | | search/annotation | 12 queries | 14 – 3 630 |

    identifier and value never approach 50 in practice — uncapping them is free. annotation is where the damage is, and also where an uncapped default is expensive: a row costs ~272 bytes (~68 tokens), so @Column (3 630 rows) would be ~1 MB / ~250 k tokens in one response. Removing the cap outright, as 103 did, is therefore not obviously right here.

    Proposed fix, in order of preference.

    1. Keep the default, make the cut detectable. Send X-Total-Count and X-Truncated headers on the paginated bare-array endpoints. Zero bytes in the body, no contract change for any existing REST/MCP/CLI/UI consumer — exactly the objection that stopped 103 from adding an envelope. This also makes paging usable for the first time: the caller learns after call one whether paging is worth it.
    2. If only the number may change: uncappedLimit for identifier and value (their result sets are small), and raise annotation to ~300 — above the sets anyone asks a completeness question about (@Path 276, @RunInTransaction 233, @Entity 160, @Immutable 95) and far below @Column.
    3. ?countOnly=true. A completeness question is a counting question. It would answer "how many classes carry @Immutable" for a few bytes instead of 20 k tokens.

    Paging already works and is not the fix. Verified on the 95-row set: offset is honoured, the order is stable across repeated calls (five identical checksums), five pages of 20 reassemble the single fetch exactly — 95 rows, 95 distinct, no gap, no duplicate, same order — and reading past the end returns 200 with [] rather than an error. So "short page = done" is a sound stop signal. But paging is opt-in, and the caller who does not know the answer was cut is exactly the caller who will not page. Two caveats worth recording: the order is by internal node id, not alphabetical (KeyTableEntryEntity sorts last), so it is stable only within one graph state — a refresh between pages can shift the set; and a total that divides evenly by the limit costs one extra empty call to detect.

    Documentation gap, independent of the code. limit/offset are documented in agent-api-system-prompt.md only for context (§2). For the three search endpoints (§8) the parameters are not mentioned at all, so an agent reading the guide has no reason to suspect a cap. Fixed in the same pass — see §8 and the pitfalls list.

    Delivered (2026-08-19), option 1 + option 3 of the three proposed above.

    • X-AC-Total-Count and X-AC-Truncated on all three endpoints. Bodies stay bare arrays — no contract change for UI/CLI/MCP/agents, which is the objection that stopped item 103 from adding an envelope, and consistent with item 130's X-AC-* scope headers.
    • ?countOnly=true (CLI --count-only) returning {"count": n} — the completeness question answered for a few bytes rather than ~250 k tokens on @Column.
    • Default limits unchanged (option 2 rejected): raising them moves the cliff and makes the expensive case worse. With the cut visible and countOnly available, 50 is the right default.
    • CLI parity, which headers alone would have broken. ApiResponse was (statusCode, body) and dropped headers, so a truncated CLI answer would have looked complete — the same silent cut, one layer out. It now carries headers and prints a note to stderr (stdout stays pipeable into jq).

    Two implementation decisions worth keeping:

    • Each query is split into a shared predicate core plus row/count projections, composed from one constant. Two hand-maintained copies would drift, and a total that disagrees with its rows is worse than no total — it turns a visible truncation into a confident wrong number.
    • The count query runs only when the page comes back full. A short page is provably the end of the set, so its total is arithmetic. That matters most for search/identifier?contains=, an unindexed scan measured at 1.2-2.3 s on upms; counting unconditionally would have doubled it. The critical trap, pinned by a test: the count must not inherit item 125's $scanCap, or the total equals the row count every time and the whole item is silently undone.

    Covered by SearchTruncationIT (8 tests).

Follow-ups found while closing 125-130 (2026-08-19)

  • 134. A post-deploy API smoke test (2026-08-20)

    Why. Items 129 and 130 both shipped with bugs that no test caught and that only manual live probing exposed: //file in a REST path, 183 duplicate rows of 436, an inherited JAX-RS @Path collapsing to POST /, ?paths=pom.xml being accepted as a source file. The integration tests pin semantics against fixtures; nothing checked the deployed server against the real graph.

    Delivered. x-scripts/verify-api.sh — curl + python3 only, no Maven, no Docker, ~10s against ac/app/pur, ~25s against upms. ./x-scripts/verify-api.sh [-p <project>], exit 0/1, one PASS/FAIL/SKIP line per check. Five groups: reachability and version; the item-130 scope and freshness headers; the item-131 paging contract (X-AC-Total-Count numeric, countOnly agrees with the full total, limit=1 sets X-AC-Truncated, consecutive offsets are disjoint); data plausibility per endpoint family; and the negative cases (structured PROJECT_NOT_FOUND / MODULE_NOT_FOUND, and the item-129 ?paths=pom.xml regression).

    Deliberate limits. Every assertion is an invariant (> 0, no duplicates, header present, required field set) — never a fixed row count, because counts move with each refresh and differ per project. Endpoint families that a project legitimately lacks (rest-endpoints in a pure Natural project, an annotation search with no hits, a leaf module with no callees) report SKIP, not FAIL. The run is read-only apart from refresh?paths=pom.xml, which resolves no file and starts no ingest. The duplicate-row check on rest-endpoints cannot surface item 75 — that query already applies DISTINCT; it only guards against the DISTINCT being dropped again. This is not a test replacement and must not be treated as a quality gate.

    It paid for itself on the first run — three real defects, now items 135 and 136 below.

  • 135. search/references and rest-endpoints were left out of item 131 — they still truncate silently ( 2026-08-20)

    Symptom. verify-api.sh reports on all four projects: search/references: X-AC-Total-Count is numeric — got '' (status=200), same for rest-endpoints. Neither response carries X-AC-Total-Count or X-AC-Truncated.

    Cause. Both handlers in AnalysisResource return ok(...) on a plain List instead of paged(...) on a Page<>. searchReferences uses effectiveLimit(limit), so it caps at the default 50 with no signal at all — exactly the failure item 131 set out to remove. restEndpoints uses uncappedLimit(limit), so it does not lose rows today, but it is equally silent about how many there are.

    Delivered. Both queries split into *_CORE + a shared row projection + a *_COUNT that is WITH DISTINCT <the same columns> RETURN count(*) — the same column set as the row projection, on purpose: a narrower distinct key would count rows the page never delivers. Neither count carries $scanCap, for the reason already pinned on SEARCH_IDENTIFIER_COUNT. restEndpointsPage / searchReferencesPage go through the existing withTotal(...), so the count query runs only when the page comes back full — and with rest-endpoints' default uncapped limit it never runs at all. Both endpoints gained ?countOnly=true, both CLI commands --count-only; warnIfTruncated already sat in printResponse, so the stderr note started working the moment the headers appeared. SearchTruncationIT grew from 8 to 15 tests.

    One thing the fixture taught us: a reference to a type in the same package produces no edge at all — there is no import, and TypeResolver cannot turn the bare simple name into an identity. The test fixture had to put the referenced class in its own package to have anything to page over. That limit is documented on JavaParser#addReferenceEdges and is worth knowing before scoping a rename inside a single package.

  • 136. search/identifier faults with an unstructured 500 on a deep page (2026-08-20)

    Symptom. GET /api/projects/ac/search/identifier?name=e&contains=true&limit=500 answers 500 with a plain-text Quarkus error page. limit=300 is fine, so it is not the limit itself but which rows land in the page. Server log: org.neo4j.driver.exceptions.value.Uncoercible: Cannot coerce NULL to Java int.

    Cause. The identifier row mapper coerces startLine/endLine unconditionally, and at least one node in ac carries NULL there. Only reachable past ~300 rows, which is why no test and no manual probe hit it.

    Two defects, not one. Beyond the mapping bug, the response violated the API principle that errors are structured JSON — an agent got an HTML-ish body with no code to branch on.

    Delivered. Three changes, because one alone would have been a patch over a symptom:

    1. Cause. MARK_DUPLICATE_IDENTITIES (item 114) was the only node-creating query that set neither startLine nor endLine. It now sets both via coalesce(n.startLine, 0) — coalesce because its MERGE key is deliberately MERGE_NODES', so marker and reference placeholder can be one node whose real lines must survive. 0 is a convention, not a truth: a marker has no line. The alternative — letting two properties be null on a handful of nodes — pushes null handling into every row mapper in the project, and that is exactly how this bug happened.
    2. Defence. The identifier row mapper uses the existing intOrZero(record, key) helper. A row mapper must never be the thing that faults an endpoint. The other asInt() call sites were reviewed and left alone: they read FUNCTION/FIELD/SQL rows that always carry lines, and churning 50 call sites would have been a different, larger change.
    3. Contract. New ApiExceptionMapper (@Provider, ExceptionMapper<Throwable>) returns 500 INTERNAL_ERROR with an errorId in details matching the logged stack trace, which is not echoed to the client. The WebApplicationException pass-through is load-bearing — Throwable is the least specific mapper there is, and without it every deliberate 404/400/409 would have become a 500. ApiExceptionMapperTest pins both directions.

    Not retroactive. The write-side fix takes effect on the next ingest; the 11 broken nodes in ac stay until then. The mapper hardening makes the endpoint correct immediately, refresh or not.

Lazy / deferred ingest (three-tier model)

Reworks ingest from eager whole-project parsing into a lazy, on-demand model. Three tiers: Tier 1 = cheap eager reference index (per file: nodes, identifiers, coarse call/DB references — no deep bodies); Tier 2 = lazy deep ingest (control flow, statement-level dataflow, precise reads/writes) triggered on demand; Tier 3 = source served from the filesystem, no longer stored on nodes. Reverse queries (callers, search_identifier, flow_backward) stay answerable because Tier 1 pre-indexes coarse references globally.

Items 36–43 are done (Tier-1 reference index + tri-state status, Tier-2 lazy deep-ingest, depth/node caps, unresolved-reference nodes, Tier-3 source-from-disk with stale check, the refresh surface, and hash-based auto-invalidation) — see x-docs/features.md. No open items remain in this track.

Ingest performance

(Items 24 — the stale-file sweep's missing (project, sourceFile) index, the actual persist bottleneck — and 25 — batch persist, found already implemented — completed 2026-07-16 and moved to x-docs/features.md. A full upms call-graph refresh (6311 files) now takes 179 s end-to-end (~63 s parse, 104.5 s persist across 32 batches, ~12 s finalize), against ~2.5–2.8 min per batch before. The parked "parallel parse phase" idea was implemented 2026-07-18 (item 24) — see x-docs/features.md. No open items remain in this track.)

  • 111a. NodeType as a Neo4j label, stage 1: written on every persist (done 2026-08-05)

    Measured problem. The node type lived only in the type property, so every expand-and-filter read it out of the property store. Profiled on upms (570,739 nodes, 2,002,935 relationships):

    | query | dbHits | rows | |---|---|---| | (:AstNode{type:'MODULE',project})-[:CALLS]->(:AstNode{type:'MODULE'}) | 1,598,264 | 5,384 | | same, without the target's type filter | 663,156 | 25,843 |

    ~935k dbHits — 58% of the query — spent reading one string property, ~36 hits per candidate node because type sits in an 11-property chain. A label lives in the node record and costs no property-store access. Of the 200 type: '…' occurrences in CypherQueries, 106 filter on type without binding name, and none bind both — so the composite index (project, type, name) is never used with all three columns, and the win is in traversal filtering, not in index seeks.

    Delivered (stage 1, deliberately additive). SET node:$(n.type) in both persist paths (MERGE_NODES, MERGE_POSITIONAL_NODES — dynamic labels verified available on Neo4j 5.26.27), plus BACKFILL_TYPE_LABELS run from ensureSchema(), batched IN TRANSACTIONS so 570k nodes do not build one transaction state larger than the 2 GB Neo4j heap. Labels take no part in the MERGE identity, so merge semantics are unchanged. The type property stays, so none of the ~200 queries change behaviour and nothing has to migrate at once. Covered by TypeLabelIT.

    Not yet done — stage 1 alone buys nothing. It only creates the precondition:

    • 111c: drop the type property and the four AstNode collection indexes once no query reads it (−9 chars/node).
  • 111b. Hot-path queries anchored on the label instead of the type property (done 2026-08-05, builds on 111a)

    Migrated in CypherQueries: callers (both scopes), callees (both branches and its scope filter, which runs on every expanded candidate), MODULE_HOP_OUT / MODULE_HOP_OUT_WIRING (the call-tree traversal), FUNCTION_CALLERS, and the two module anchors in SEARCH_IDENTIFIER. Its $type filter stays a property comparison on purpose — it is a bound parameter, not a literal, so it cannot become a static label.

    Two new indexes were required, not optional. A Neo4j index serves exactly one label, so (m:MODULE {project, name}) cannot use the :AstNode indexes. Without module_project_name / function_project_name the anchor lookup would degrade from an index seek to a label scan — turning the migration into a regression at the very point it is meant to help.

    ~158 type: '…' literals remain in the enrichment queries; they migrate incrementally, each with its own measurement.

    Baseline captured before deploying (upms, best of 3 per endpoint, version 160):

    | module | endpoint | before | |---|---|---| | WGEAGB0S | callers / callees | 24 ms / 14 ms | | WGEAGB0S | call-tree?depth=3 / depth=5 | 266 ms / 312 ms | | ZINERR01 (1689 callers) | callers | 329 ms | | YFRAMN04 (916 callers) | callers / call-tree?depth=5 | 335 ms / 262 ms |

    Measured after deploying (version 162) — the performance case did not hold up.

    | module | endpoint | before | after | |---|---|---|---| | WGEAGB0S | callers / callees | 24 / 14 ms | 21 / 16 ms | | WGEAGB0S | call-tree?depth=3 / depth=5 | 266 / 312 ms | 255 / 292 ms | | ZINERR01 | callers | 329 ms | 323 ms | | YFRAMN04 | callers / call-tree?depth=5 | 335 / 262 ms | 277 / 141 ms |

    At the Cypher level, where dbHits are deterministic and HTTP noise is absent:

    | query | before | after | change | |---|---|---|---| | callers YFRAMN04 | 16,202 dbHits | 14,369 | −11% | | module hop (call-tree core), 4 seeds | 1,864 dbHits | 1,850 | −0.8% |

    Why the 15.8× microbenchmark did not transfer. It scanned every module in the project and filtered the expansion; the real queries seek one module by name and expand over hundreds of edges. There are simply almost no property reads left to eliminate. The YFRAMN04 call-tree 1.9× is not attributable to the migration — the module-hop core it is built from improved by 0.8%. A 12-run repeat of the one apparently slower endpoint (WGEAGB0S/callers) gave min 21 ms / median 43 / max 66: no regression, just a measurement dominated by noise.

    Kept despite this, on a different justification than the one it was proposed under: it is a small consistent improvement and no regression, and — the actual reason — 111c cannot happen without it. Dropping the type property requires that no query reads it. The remaining value of 111a/b is the storage reduction 111c unlocks, not query speed.

    Lesson recorded deliberately: the pre-implementation benchmark was chosen for convenience, not for resemblance to the production query shape, and overstated the benefit by more than two orders of magnitude. Benchmark the query the code actually runs. 295/295 ITs pass.

    Write cost — measured 2026-08-05, no regression. A project recreate upms --deep with the label writes in place took 541 s for the parse/persist phase (6311 files), against 853 s for the same phase without them. Not a like-for-like comparison — a recreate creates nodes rather than merging onto existing ones and skips per-file reconciliation, so it is expected to be faster — but there is no sign of the slowdown that would have forced a revert.

    Payoff confirmed on the real graph — bigger than estimated. Same graph, same query, same result (5381 rows), only the filter style differs:

    | | dbHits | time | |---|---|---| | (:AstNode{type:'MODULE',project})-[:CALLS]->(:AstNode{type:'MODULE'}) | 1,601,071 | 512 ms | | (:MODULE{project})-[:CALLS]->(:MODULE) | 101,120 | 25 ms |

    15.8× fewer dbHits, 20× faster — but do not read this as the payoff of the migration. It is not. This query starts from every module in the project and expands, so the property filter is applied to hundreds of thousands of candidates. The real API queries anchor on a single module by name via an index seek and expand from there, over hundreds of edges rather than hundreds of thousands. See 111b for what the migration actually delivered on those (11% and 0.8%, not 15×). The number above is a property of this microbenchmark, not of the codebase.

  • 111d-1. The 36-char node UUID is gone — replaced by two inline longs (done 2026-08-05)

    Measured share. Of 100,000 sampled nodes, 100% carried an id over Neo4j's ~15-char inline threshold (vs sourceFile 96%, name 33%, value 13%, type 0%). At ~2.45 long strings per node against a 385 MB string store, the UUID is roughly 41% of it (~157 MB).

    What it actually did — and the mistake that nearly shipped. The first proposal asserted the id had "exactly one load-bearing use" (the item-58 reconciliation sweep). That was wrong: it is also the join key wiring a batch's edges to the nodes merged in the same transaction (mergeEdgesBatch). The assertion came from grepping n.id patterns and classifying the hits, without following the dataflow — the edge query was in the same file and was missed. Removing the property on that basis would have produced a graph with nodes and no edges, failing silently. The change was reverted and re-proposed. Enumerate by dataflow, not by grep sample.

    Design. Per persist transaction: one ingestGen stamp (wall-clock seeded, monotonic, so a fresh JVM can never reuse an earlier run's generation) plus a per-transaction nid counter. Both inline longs. Verified precondition: persist and persistBatch both wrap mergeResults in a single executeWriteWithoutResult, so nodes and their edges are always in one transaction — which is why a counter suffices where a UUID was needed.

    • edges: MATCH (a:AstNode {nid: e.sourceNid, ingestGen: $ingestGen}), (b …)
    • reconciliation (items 58/76): WHERE n.ingestGen IS NULL OR n.ingestGen <> $ingestGen. The IS NULL half is required — Cypher's three-valued logic makes NULL <> x yield NULL, so legacy nodes would otherwise be undeletable.
    • new index (ingestGen, nid) — required, not an optimisation: without it every one of upms's 2M edge endpoints degrades from a point seek to a label scan.
    • ast_node_id_unique dropped explicitly (deleting the CREATE alone would leave it in place on every existing database), and DROP_LEGACY_NODE_IDS strips the retired property from pre-existing nodes — SET node += n.properties does not remove properties, so without it every surviving node would keep its old id forever.

    The API contract got stronger, not weaker. /nodes/{id} now returns Neo4j's elementId (4:<db-uuid>:<internal-id>), which survives a re-ingest — where the UUID was overwritten on every single merge, which is what forced the old "same ingest generation only" caveat. VALIDATE had called this a deliberate weakening; that was backwards. Affects ac inspect-node / ac node-source; the web UI never used node ids.

    Also removes the ~1.1M UUID strings shipped as query parameters per full refresh (two id lists). Covered by NodeIdentityIT; 300/300 ITs pass.

    Not yet verified: the net store saving after subtracting the new (ingestGen, nid) index. Only measurable on the store files after a full re-ingest.

  • 112. A whole-project refresh logs no completion line (done 2026-08-06)

    A finished refresh was only visible as POST /api/projects/upms/refresh?deep=true -> 200, an access-log line naming neither the project's outcome nor how long it took. Anything watching the log for completion — a monitoring script, an agent, rebuild-and-refresh.sh — had to grep HTTP status lines, which is why a watcher written during the 111d-1 verification missed the end of a 1191 s run outright.

    ProjectIngestService.refreshProject now emits one INFO line covering both modes:

    Refresh finished: project='upms', mode=deep, files=6311, modules=6311, failed=0, 1191 s
    

    Placed after sweepDeletedFileOrphans so the timing covers the deleted-file sweep too — the line means "the refresh is done", not "the ingest is done". Per-module refreshModule is deliberately left alone: it is short, called often, and would only add noise.

  • 113. Every server operation logs when it finishes (done 2026-08-06, extends 112)

    Rule applied: an operation that logs a start must log an end. Four violated it.

    The access log was already there — its duration was just never recorded. Every request has been logged as POST /api/projects/upms/refresh?deep=true -> 200 (-ms); the (-ms) is not a missing feature but quarkus.http.record-request-start-time defaulting to false, so %{RESPONSE_TIME} has nothing to render. Setting it gives every endpoint a completion line with a real duration for one property and one System.nanoTime() per request. It adds no new log lines — the access log already covered every call. The property lives in VertxHttpConfig, not VertxHttpBuildTimeConfig, i.e. it is a runtime property, so it does not have to be repeated in the test-side application.properties.

    Two domain completion lines, both at the shared worker rather than the entry points:

    • ProjectIngestService.ingestRoot → Project ingest finished: project=…, mode=tier1|call_graph|full, files=…, persisted=…, failed=…, duplicates=…, N s — covers project create, the call-graph pass and the deep whole-root pass.
    • ProjectIngestService.bfsIngest → Deep ingest finished: project=…, seed=…, files=…, failed=…, unresolved=…, truncated=…, N s.

    bfsIngest is the single worker behind both ingestModules (by name) and ingestFiles (the fan-out warm). Lines first written into DeepIngestCoordinator at the two call sites were removed again once that was traced — they would have double-logged every deep ingest. The files count is per operation, deliberately: the walk emits one Ingesting <module> … start line per file, so the single end line is what those N start lines add up to.

    No try/finally with an outcome flag, though the proposal called for one. All three sites already log their failure path (Fan-out warm … failed, Auto deep-ingest … failed, and an ingestRoot exception surfacing as a 500 with duration once the access log works). Buying symmetry would have meant restructuring the definite-assignment flow of a 90-line method on the hot ingest path for a log line — a bad trade.

    A refresh now logs two lines by design: Project ingest finished (parse + persist) and Refresh finished (item 112, additionally covering the deleted-file sweep). The difference between them is the sweep's cost, which was not visible anywhere before. refreshModule stays silent — short, frequent, pure noise.

Known bugs

  • 119. Enums, records and annotation types are modules (found 2026-08-06 in the pur source-vs-API cross-check; done 2026-08-06)

    The Java parser iterated ClassOrInterfaceDeclaration. Everything else was invisible — in pur 76 annotation types, 67 enums and 25 records, 168 types with no MODULE node at all (plus 384 package-info.java, correctly ignored). GET /modules/BatchParam/digest answered 404; a record referenced from another file stayed an unresolved placeholder and answered 409 NOT_INGESTED. The API did not lie — item 107 saw to that — but a whole category of type was unanalysable.

    The loop now iterates TypeDeclaration. Exactly three things differ between the four kinds, and they are resolved once in a TypeFacts adapter instead of scattering instanceof through a 250-line body: only a class/interface can extends, only an annotation type can do neither, and isInterface is a class/interface notion. moduleKind gains ENUM | RECORD | ANNOTATION.

    What the probe corrected in the plan. Rather than reason about the JavaParser API I ran a parser over a fixture with all four kinds, and two assumptions were wrong:

    • A record's compact canonical constructor (Point { … }) is not returned by getConstructors() — CompactConstructorDeclaration does not extend CallableDeclaration and so does not fit the callable machinery. Left uncovered deliberately (1 occurrence in pur), and named here because the calls in its body stay invisible.
    • An annotation type's String value(); is not a method — getMethods() returns none. Building it as planned would have given 76 annotation types a module node with an empty body: "analysed, nothing found" again. Members are modelled as FIELDs with defaultValue, because the question asked of an annotation is which attributes it carries.

    Record components and enum constants are captured for the same reason — a record DTO reporting zero fields is worse than no answer. The TypeResolver and the enclosing-type walk had to follow, or a reference to an enum declared in the same file would have been unqualifiable — item 117 running backwards for precisely the types it just gained.

    Kept class/interface-only on purpose: the JPA/Panache repository heuristics. A record is not a Panache repository, and generalizing guesswork to types it was never written for produces confident wrong answers rather than silence.

    Watched, not asserted: the CHA fan-out grows, since an enum implementing a project interface is now a real implementation target. And 168 additional types can make a previously unique simpleName ambiguous, so a short name that used to work may now answer 409 — the honest answer under 115, but a visible change.

  • 118b. A control character in one constant disabled two guards, and the tests that should have caught it asserted the same wrong literal (found 2026-08-06 in the pur cross-check; done 2026-08-06; both regressions from 116b)

    A — the marker leak. UNRESOLVED_FIELD_RECEIVER_PREFIX contained a stray U+0001:

    JavaParser.java:157   = ".field:";
                      hex:  3d 20 22 01 66 69 65 6c 64 3a 22 3b
    

    So the parser wrote field:x while both Cypher guards matched STARTS WITH 'field:' — the cleanup in DELETE_UNRESOLVED_FIELD_RECEIVERS and the exclusion in RESOLVE_SIMPLE_NAME_REFERENCES. Neither ever matched. In pur, 202 marker nodes survived, 180 of them still wired with 850 edges, and the scaffolding was served from the public API:

    GET /modules/…KundeService/callees → { "name": "field:e", … }
    

    Why it survived review and tests. The two ITs written to guard exactly this checked m.name STARTS WITH 'field:' and not hasItem("field:repo") — the same wrong literal. They were true because they matched nothing. Code and test were wrong in the same way, so the test could not see the bug it existed for.

    Fixed at the source, and both predicates now match CONTAINS 'field:': a ':' cannot occur in a Java or Natural module name (verified across all four projects), so it is equally sharp — and it also clears markers left by an earlier build, which a refresh would never reach otherwise (placeholders have no sourceFile, so the per-file sweeps do not touch them).

    B — lambda parameters taken for inherited fields. isProbableFieldReceiver asked "lower-case and not in declaredTypes?". Lambda and catch parameters are in neither the callable's parameter list nor its VariableDeclarationExprs, so .map(e -> e.getX()) looked like an inherited field named e — 24 markers from one class alone, none of them ever resolvable. The bound names now go into the existing shadowed set. Deliberately not into declaredTypes: an implicit lambda parameter has type UnknownType, and feeding that to the receiver resolver would turn a silent omission into a confident edge to a module named after a non-type.

    Both fixes proven by reverting them. With A restored to its broken form, three assertions fail — including field:somethingUndeclared reaching /callees. The first B test I wrote was itself vacuous: with A fixed the cleanup deletes every marker, so nothing is observable end-to-end. It moved to JavaParserTest, where reverting B turns it red with [field:entry, field:ex, field:inherited]. The new assertions match the marker anywhere in the name and, separately, assert the invariant that was actually violated — no module name contains a control character.

  • 118. The Java DB resolvers scanned the project once per candidate row (found 2026-08-06 while watching a pur refresh; done 2026-08-06; regression from [117])

    Symptom. A deep refresh of pur (2988 files) spent 391 s of 447 s in two of 48 finalize steps — and both produced zero edges:

    [10/48] resolve-java-db-access  : 305 382 ms   rels +0/-0, props 0
    [11/48] resolve-java-query-jpql :  85 537 ms   rels +0/-0, props 0
    the other 46 steps              : ~10 s total
    

    Cause, and it was mine. 117 added OR m.simpleName = … so endpoints accept the short form, and an index to keep that a seek — but declared it as FOR (n:MODULE), while these enrichment queries anchor on (:AstNode {type: 'MODULE'}). The label indexes are therefore invisible to them, and the plan confirms what that costs:

    +Union
    | +Filter        repo.simpleName = a.javaReceiverType …
    | +NodeIndexSeek RANGE INDEX repo:AstNode(project, sourceFile)     | 216674 rows
    | +NodeIndexSeek RANGE INDEX repo:AstNode(project, type, name)     |    198 rows
    

    The simpleName branch scans the whole project — under an Apply, i.e. once per DB_ACCESS row. 3563 candidates × 55 030 nodes ≈ 196M property reads. The same 117 note that (correctly) refused to put OR simpleName inside the 18 API queries for exactly this reason missed that four enrichment queries already had it.

    Fix. Anchor on the dynamic labels: (a:DB_ACCESS …)<-[:CONTAINS]-(fn) instead of a label scan over all 573k AstNode + 639k CONTAINS, and (repo:MODULE …) / (entity:MODULE …) so all four OR branches become seeks (216 674 estimated rows → 1 and 14). Applied to RESOLVE_JAVA_DB_ACCESS, RESOLVE_JAVA_QUERY_JPQL, RESOLVE_JAVA_QUERY_NATIVE_SQL and the two db-accesses API queries that share the pattern.

    This is not a pure plan change, and it was verified as such. It depends on every node carrying its type label; if one lacked it, the module would silently drop out of resolution and report "no DB access" — the [114] failure class. Checked both directions on live data (all four projects, including the freshly refreshed pur with its placeholders): AstNode{type:'MODULE'} without :MODULE = 0, :MODULE without the property = 0, same for DB_ACCESS. Labels come from SET node:$(n.type) in both node-merge paths, i.e. at ingest time, and no other query creates AstNodes.

    Equivalence, measured rather than argued. Same 200 candidate rows through both formulations, result triples (function, access, table) sorted and hashed: 6c829e01… for old and new alike — 12.3 s vs 1.7 s. Full query over all 3563 rows: 1.7 s (was 305 s).

    Not claimed: the JPQL step cannot be verified on pur — it has zero JPQL rows there, so its 86 s were pure anchor cost and only the anchor fix applies. And every timing above is warm-cache; the 305 s arose cold, right after 2988 file writes. What the finalize actually costs after this gets measured on the next refresh, not predicted here.

  • 117. A Java module's identity is its fully-qualified name (done 2026-08-06; the root fix behind [115])

    115 stopped the wrong answers; this removes the cause. A Java module node's name is now the FQN (com.example.OrderService, and a.b.Outer.Inner for a nested class — verified against JavaParser, which also returns empty for local/anonymous classes, so those keep the simple name). simpleName carries the short form for display. Endpoints accept either: the FQN resolves exactly, the short form is the convenience and falls back to 115's 409 AMBIGUOUS_NAME.

    The rename is the easy half; references are the hard one. All eight reference sites in the parser created their placeholders from simpleType(...). Renaming only the definitions would have left every reference unable to find its definition — the whole Java call graph, silently. A TypeResolver per compilation unit now qualifies from the file's own declarations and its explicit imports, and refuses to guess: "not imported, therefore same package" would have produced com.example.String. What stays unqualified is picked up by the new resolve-simple-name-references enrichment step, which matches on simpleName and only when exactly one module matches. Measured justification: 12 of 2988 files in one real codebase use a wildcard import (0.4%).

    Resolution happens once, in the API guard. MODULE_INGEST_STATE matches name OR simpleName and returns the resolved identity, which the guard hands to the endpoint; every downstream query then runs on the identity alone. The alternative — OR simpleName inside all 18 module queries — would have dropped them from an index seek to a label scan on 500k nodes, undoing items 111a/b. A (project, simpleName) index was added to keep the short form a seek as well.

    What the tests caught that review did not. Two regressions, both silent:

    • Cross-class dataflow went empty. The ?module= scope filters are not behind the guard, so they got a short name and matched nothing. Fixed by resolving the scope centrally (resolveModuleName), also used by the module-hop BFS seed and the ?extends= filter.
    • The em.persist(...) write path disappeared. RESOLVE_JAVA_DB_ACCESS tests javaReceiverType IN ['EntityManager', 'Session', 'StatelessSession'] — hardcoded simple names. Qualifying EntityManager (it is imported, so the resolver could prove it) made that branch dead. The fix was to stop qualifying the DB type properties rather than qualify three hardcoded lists: the module lookups accept the simple form via simpleName, so nothing is lost where it matters.

    Isolating the first one is worth recording: disabling the new enrichment step and seeing the test still fail is what proved the cause lay in the query, not in the resolution.

    Two of my own errors, for the record. The reference-resolution query was first written with APOC (not available here) and then registered as enrichment step 24 although its own comment said it must run first — leaving every later join reading a placeholder about to vanish.

    Also unified on the way: declaredIn is a display label in every producer (column metadata and both function queries), i.e. the short name. Two of the three had drifted to the identity.

    UI (done 2026-08-06, after the deploy that unblocked the codegen). The client is generated from the running server's OpenAPI, so this had to wait for a build that knows simpleName. Regenerating pulled in exactly the 115/117 contract and nothing else: 18 sourceFile query params, 11 reworded 409 descriptions, 1 simpleName.

    The find was that this is not a cosmetic task. Explorer.tsx filters the module list by regex over name, and name is now the FQN — so an anchored pattern on the class itself returned nothing:

    ^AbstractLogic$   →  0 Treffer   (name = com.uniqagroup.common.base.AbstractLogic)
    

    The existing explorer spec runs against upms (Natural, no dots in any name) and stayed green throughout. The filter now also matches simpleName; the package path and the FQN remain searchable.

    Display: eight sites render the short name with the identity in the title — module list, module header, graph node labels and the selected-node chip, callers/callees, call tree (including its cycle markers), impact list, dataflow steps. Identity is untouched everywhere it matters: routing, query keys, key=, requests, and the graphology node key. Where the server sends simpleName it wins over the client-side rule, because it knows the cases a string rule cannot — a class in the default package, or a local class for which JavaParser reports no qualified name.

    The derivation ("text after the last dot") was checked against the server rather than assumed: for all 4734 qualified modules of pur, it equals simpleName — 0 divergences; and 0 of 3587 upms module names contain a dot, so Natural is provably unaffected. Worth noting the graph is currently mixed-generation (only pur is re-ingested post-117), which is what makes the prefer-then-derive fallback necessary rather than nice.

    Accepted loss: a nested class shows as Inner, so two Outer.Inner in different outers look alike in a list. The FQN is one hover away and the list also shows the source file.

    New e2e spec java-fqn.spec.ts (the only one that runs against a Java project) — and it was verified to actually catch the regression by reverting the filter fix and watching it go red. tsc --noEmit + vite build clean; Playwright 21 passed, 1 pre-existing failure in identifier-popover.spec.ts (expects line=323, gets 324 — reproduced with all UI changes stashed, so it predates this work and belongs to the upms data, not the UI).

  • 115. Module endpoints address Java classes by simple name and silently merge distinct classes that share one (found 2026-08-06, pur source-vs-API cross-check; done 2026-08-06)

    Symptom. GET /modules/BrokerHistoryTests/functions returns 627 functions. The class it names has 107. The other 520 belong to four unrelated classes:

    MATCH (m:MODULE {project:'pur', name:'BrokerHistoryTests'}) RETURN m.sourceFile
    → HistoryCalculationServiceTest.java   84
      HistoryComplexServiceTest.java      203
      HistoryRecordPageServiceTest.java   128
      HistoryRecordServiceTest.java       107
      HistoryVariantServiceTest.java      105   = 627
    

    Five distinct @Nested classes, one per test file, each legitimately its own module — correctly ingested, correctly distinguished in the graph by sourceFile. The API collapses them, because /modules/{name}/… matches on name alone and the path offers no way to disambiguate (verified against the OpenAPI spec: the only parameters are project and name).

    Not a test-only problem. WorkingStorage is six nested classes in six production files (LastAgentLoopLogic, AuthorizationCheckStandardProcessingLogic, InitializationService, …). Its digest reports all six enclosing classes as CONSTRUCTOR callers, reading as one shared class constructed in six places when in truth each file has its own.

    Scale in pur: 163 ambiguous names covering 385 modules — ~8% of all 4738 modules. Worst cases collide 7-fold (ProcessTests, VermittlerstammTests). Java makes this normal: nested test classes, Builder, Config, Handler, WorkingStorage.

    Why it matters. This is worse than a wrong count. Every derived answer — callers, callees, db-accesses, call-tree — is a union across unrelated classes, presented with no indication that a merge happened. An agent asking "who uses WorkingStorage?" gets six answers for six different classes as though they were one, and cannot tell.

    Solution sketch.

    1. Make ambiguity visible before making it addressable. When a name resolves to more than one module, answer 409 AMBIGUOUS_NAME with the candidate sourceFiles in details rather than silently unioning. This is the same shape as the 107 fix and stops wrong answers immediately, at the cost of breaking queries that currently "work".
    2. Then make it addressable. Accept an optional sourceFile (or fully-qualified name) query parameter on the module endpoints to pick one candidate; search/identifier already returns the sourceFile per hit, so an agent has what it needs to disambiguate in one extra call.
    3. Natural is unaffected in practice — its module names are file-stem-unique by language convention, and genuine collisions are already handled (skipped) by the duplicate detection; see [114], which is the other half of this problem.

    Not yet decided: whether Java modules should carry the FQN as name instead. That would fix addressing at the root but changes every response body and the whole UI, so it needs its own proposal rather than being folded in here.

    Delivered. Both halves together — the guard alone would have made 385 modules in pur and 326 in app unreachable by name, trading one wrong answer for another.

    • ModuleIngestState carries candidates (the real, non-placeholder sourceFiles for the name) and ambiguous(). MODULE_INGEST_STATE returns them and honours a $sourceFile filter.
    • withModule/withIngestedModule (the item-107 guards) answer 409 AMBIGUOUS_NAME with the candidates in details. Guard order matters: absent → ambiguous → placeholder. A placeholder is never a candidate, so one real module beside a placeholder is not ambiguous and #SUBPROGRAM-style references keep answering NOT_INGESTED.
    • ?sourceFile= on the module endpoints, threaded through 18 Cypher queries and the repository.
    • ac --source-file on all 17 module commands, via a picocli @Mixin whose append() picks ? or & — several commands add their own parameters after the base path and a fixed ? emitted two.
    • UI: 409 no longer maps blindly to NOT_INGESTED; the body's code decides, and the banner lists the candidates. Listing them is the point — a bare "ambiguous" is barely better than the merge.

    Two corrections worth keeping. First, a line-based inventory found 27 affected Cypher sites; a block-based one found 33 — the same grep-instead-of-dataflow error as item 111d-1, 20% low. Second, EGO_NEIGHBORS_IN/OUT were patched and then unpatched: they are keyed on the BFS frontier's name, not the root's, so filtering them by the root's sourceFile would have broken traversal. Ambiguous neighbour names still merge inside a call tree — the same scope boundary item 107 has, and it is not closed here.

    The runtime hazard predicted in review actually fired. Cypher rejects an unbound parameter at runtime, not compile time; moduleFunctions built its parameters as a hand-rolled HashMap (because kind is nullable and Map.of rejects nulls), so it was the one site the central moduleParams helper did not cover, and it returned 500. Caught by AmbiguousModuleIT, which asserts a real selection (1 method vs 2) rather than just a 200.

  • 116. A method call on a field the calling class does not declare produces no edge (found 2026-08-06, pur source-vs-API cross-check; done 2026-08-06)

    Symptom. HistoryRecordService (53 methods, exercised by 185 @Test methods) reports callers: {} — nobody calls it. The tests call it constantly:

    class HistoryRecordServiceTest {
        private HistoryRecordService service;     // line 33, outer class
        @Nested class BrokerHistoryTests {
            … service.applyControlCardFilters(…)  // line 50, inner class → no edge
    

    Neither the class nor the invoked method names appear among BrokerHistoryTests' callees; only receivers that are type names (static calls, constructors) resolve.

    Cause, narrowed by counter-example. Instance-field calls resolve correctly within one class: AbstractPartnerLogic.partnerRepository.findFirstByHistorySpOptional(…) yields IPartnerRepository, confirmed in both directions (33 callers). The failure is specific to the field being declared in the lexically enclosing class — field-type resolution does not cross that boundary.

    Why it matters. @Nested is the standard JUnit 5 layout, so an entire codebase's test → production call graph can be missing while the API reports it as empty rather than unknown. "Which tests cover this service?" and "is this method still used?" both answer wrongly, in the direction that looks like a clean result.

    The original write-up was too narrow, and the counter-example that produced it did not hold. It read as an enclosing-class problem because AbstractPartnerLogic.partnerRepository.findFirst(…) resolved correctly — but that class declares the field. PartnerCommonLogic, which inherits it, produces no edge either: the IPartnerRepository edge there is INJECTS (propagated by LINK_INJECTS_TO_SUBCLASSES), not METHOD_CALL. The real rule is any field not declared in this very type, and it splits by what the parser can see:

    116a — lexically enclosing types (done). JavaParser builds fieldTypes from the findAncestor(ClassOrInterfaceDeclaration) chain, applied outermost-first so an inner class's own field shadows an enclosing one. Purely in-file, so no enrichment involved. This is the @Nested case — the standard JUnit 5 layout, which had been costing the entire test→production call graph.

    116b — inherited fields (done). The parser reads one file and cannot know a supertype's fields, so it no longer drops the call: it records the receiver's identifier on a field:<name> placeholder edge. The new resolve-inherited-field-receivers enrichment step walks the EXTENDS/ IMPLEMENTS chain the graph does know, finds the declaring field, normalises its declared type (generics and package stripped in Cypher, since modules are keyed on the simple name) and re-points the edge. delete-unresolved-field-receivers then removes the scaffolding, so field:* never reaches the module namespace. Ordered before link-calls-to-implementations, so a recovered edge is still eligible for the polymorphic fan-out.

    A bug the test caught that review did not. The resolver first ended with a global ... ORDER BY target.sourceFile DESC LIMIT 1, which collapses the whole query to one row — exactly one marker edge resolved per project. On a one-level hierarchy that looks like success. The two-hop assertion in InheritedFieldCallIT failed and exposed it; the fix collapses per marker edge via collect().

    Still open: receivers that are neither locals, own fields, enclosing fields nor inherited fields — static imports and chained calls. Those markers are cleaned up rather than resolved.

  • 114. A module skipped as a duplicate identity is indistinguishable from one that does not exist, and caller/callee lists drop it without a word (found 2026-08-06, upms source-vs-API cross-check; same failure class as 107, one level deeper; step 1 done 2026-08-06)

    Done: the state is now visible. Each skipped identity gets a marker carrying the conflicting files, so module endpoints answer 409 DUPLICATE_IDENTITY with details.paths, and GET /projects/{p}/duplicates (+ ac duplicates) makes the set queryable long after the ingest response that used to hold it. The marker shares the MERGE key with the reference placeholder, so a referenced duplicate is one node carrying both facts and the more specific one wins — which also fixes the instability described above, where one cause produced 404 and 409 NOT_INGESTED depending only on whether anyone happened to call it.

    What the tests caught that the plan did not. A third sweep: delete-resolved-placeholders erases every edgeless placeholder, which is exactly what an unreferenced marker is — so DUPE kept answering 404 while the referenced CALLED worked. VALIDATE had checked the two file-level sweeps and missed this one. Markers are now excluded from it explicitly.

    Also fixed on the way: the duplicate paths were absolute server paths (/tmp/junit-…/DUPE.nat), unlike every other path the API returns. They are relative to the project root now — in the ingest response too, which had the same flaw since the duplicate detection was written.

    Deliberately not done — step 2. The missing edges stay missing: JE999 → USIX045N exists only if JE999's body is parsed, and it was not. Caller lists can still be short, they just no longer pretend otherwise. Choosing which of the conflicting files wins is a design decision (a configurable precedence such as src/manual over generated_src), not a bug fix, and this entry must not be read as closing that.

    Symptom. GET /modules/USIX045N/callers returns two callers. The source has three: the call in generated_src/subprogram/JE999.nat:321 is uncommented and real. JE999 is not in the graph at all — digest → 404 MODULE_NOT_FOUND, search/identifier?name=JE999 → []. The callers response carries no hint that anything was left out.

    Cause — a deliberate decision with an undocumented consequence. JE999 has two colliding identities, and ingestRoot filters every conflicting path out of toPersist rather than silently picking one file:

    MODULE JE999          generated_src/subprogram/JE999.nat   vs  src/manual/program/JE999.nat
    DATA_STRUCTURE JE999  …/local_data_area/new/JE999.lda      vs  …/parameter_data_area/new/JE999.pda
    

    Skipping is right — picking arbitrarily would be worse. Verified as the only colliding stem in the whole walked tree, in both categories, which matches the item-112/113 completion line exactly: files=6315, persisted=6311, duplicates=2 — two identities, four skipped files.

    Why it matters. It creates a third module state that 107 does not model, and reports it as the first:

    | graph state | answer today | truth | |---|---|---| | no node at all | 404 MODULE_NOT_FOUND | correct | | placeholder (sourceFile = "") | 409 NOT_INGESTED | correct | | skipped as duplicate | 404 MODULE_NOT_FOUND | exists twice, deliberately not ingested |

    Worse, the answer is not even stable: a duplicate-skipped module that something references gets a placeholder and answers 409, while an unreferenced one answers 404. Nothing calls JE999, which is why it vanished entirely. The same cause produces two different HTTP answers.

    The duplicate list is computed — it rides in the IngestSummary of the ingest/refresh response — but it is unreachable afterwards: Duplicate appears in the OpenAPI schema only inside that response, and no endpoint exposes it. After a 17-minute deep refresh nobody has that body.

    Solution sketch — two genuinely separate problems; the second is the hard one.

    1. Make the state visible. Persist a marker node for each skipped identity (a placeholder carrying the conflicting paths), so the module endpoints answer 409 DUPLICATE_IDENTITY with the paths in details instead of 404. Add GET /api/projects/{p}/duplicates so the set is queryable after the fact rather than only in an ingest response. This alone fixes the misleading status.
    2. The missing edges stay missing. A marker node does not restore the JE999 → USIX045N CALLNAT edge: that edge only exists if JE999's body is parsed, and it was not. So USIX045N/callers would still be short one entry, just no longer silently. Closing that needs a decision the graph cannot make alone — e.g. a configurable precedence (src/manual over generated_src) that ingests one side and flags the module ambiguous. That is a design question, not a bug fix, and must not be smuggled in with step 1.

    Scope note: fixing only step 1 leaves caller lists incomplete. That is still a strict improvement — an agent can see that a duplicate exists and go read the sources — but the roadmap entry must not imply the class of bug is closed, exactly as 107's scope note does.

  • 107. Every module endpoint answers 200 with an empty shell for a module that does not exist — indistinguishable from a real but empty module (found 2026-08-02, upms webservice-layer audit; contradicts the documented 404 MODULE_NOT_FOUND) — done 2026-08-05

    Symptom. Measured against the live server, upms:

    GET /modules/WXSPOD0S/digest      → 200 {"name":"WXSPOD0S","description":null,"functionCount":0,
                                              "callers":{},"callees":{},"dbTables":[],"dataStructures":[]}
    GET /modules/NOSUCHMOD123/digest  → 200 {"name":"NOSUCHMOD123", … identical shell … }
    

    Byte-identical answers apart from the echoed name — for a module that exists nowhere in the graph and for a name typed at random. callees, call-tree and context behave the same (all 200). search/identifier?name=WXSPOD0S&type=MODULE correctly returns [], so the graph knows; only the module endpoints invent the row. agent-api-system-prompt.md promises 404 NODE_NOT_FOUND / MODULE_NOT_FOUND — unknown id / module name.

    Why it matters — this produced a wrong analytical result, not just an ugly response. The question was "can any W* webservice module reach the commission calculation?". Five dispatchers (WPOLIX0S, WGARCX0S, WACOMX0S, WCLAIX0S, WOBJPX0S) dispatch dynamically to 13 targets (WXSPOD0S, WXSDAD0S, WXSCMD0S, WCARLD0S, WXSCLD0S, WXSCLD2S, WXSFUD0S, WXSGAD0S, WACOMD0S, WCLAID0S, WOBJPD0S, WCATEX0S, WCATED0R). call-tree?depth=4 on each returned 200 with 0 modules and 0 provenance hits, which reads as "analysed, nothing found". The truth is "not analysable" — all 13 source files are absent from the checkout (find finds none). An agent that trusts the 200 concludes "these paths trigger no commission processing"; the honest answer is "unknown". Same failure mode as item 103: an incomplete answer that looks complete.

    Fix (as implemented 2026-08-05). ModuleIngestState now carries sourceFile, so the three graph states are separable: present() (any node), placeholder() (node exists, sourceFile = ""), ingested() (real parsed node; the old exists(), renamed). Two guards in AnalysisResource replace the project-only withProject on every /modules/{name}/… endpoint:

    • withModule → 404 MODULE_NOT_FOUND when no node exists at all.
    • withIngestedModule → the same 404, plus 409 {status:"NOT_INGESTED", module, detail, nextAction} for a placeholder, reusing the existing DeepIngestRequired record.

    Applied to all 15 endpoints whose answer comes from the module's own source (digest, context, call-tree, callees, db-accesses, workfile-accesses, sql-statements, functions, functions/overrides, functions/{fn}/overrides, functions/{fn}/callers, data-structures, dispatch-table, payload, columns). callers and graph get the 404 but stay 200 for a placeholder — their data comes from the calling modules and is genuine, so the 409 copy points callers there. /modules/{name}/source was already correct. Covered by ModuleNotFoundIT (54 cases, including an explicit assertion that the fixture really produces all three states). ac-ui maps both statuses to one banner instead of ~8 per-panel failures; the CLI needed no change (printResponse already prints any body and exits 1).

    Deviation from the fix sketched above. The placeholder case returns 409 rather than a placeholder: true / sourceFile: null payload marker. A marker cannot be attached to the array-shaped responses (db-accesses, functions, dispatch-table, …), so the marker approach would have fixed the object-shaped endpoints only and left the rest indistinguishable — exactly the gap this item is about. The 409 is uniform, needs no DTO changes, and carries an actionable nextAction.

    Scope limit — the class of bug is NOT closed. This guards only the root module of a request. A call-tree that traverses into placeholder targets still reports that subtree as empty without flagging it, so the original WPOLIX0S-style dispatcher whose 13 targets are all placeholders still returns 200 with a silently truncated tree. Callers must still cross-check individual targets (each now answers 409). Flagging unanalysable nodes inside a traversal result is item 103's territory.

  • 75. CONTAINS is not acyclic — 22 self-loops and 162 two-cycles in upms (fixed 2026-08-23) (found 2026-07-17 while root-causing item 74; cause NOT established — do not treat the notes below as settled). The containment hierarchy that dozens of queries traverse with CONTAINS* contains cycles:

    (IF  @ JX0031N0.nat:781-966) -[:CONTAINS]-> (FOR @ YFRAMBC0.cpy:65-15)
    (FOR @ YFRAMBC0.cpy:65-15)   -[:CONTAINS]-> (IF  @ JX0031N0.nat:781-966)
    

    Neo4j's variable-length patterns use trail semantics (no relationship repeats in a path), so queries terminate rather than hang — but the blow-up is real: three probes using an unbounded CONTAINS* over this region were killed at a 2-minute timeout during the item-74 investigation. The shared-node identity below has a second consequence — see item 124. Because a copycode node is MERGEd per (type, name, sourceFile) and thus shared by every includer, item 124's call-edge reap cannot touch the 605 call edges whose source subroutine lives in a .cpy (14 of them stale): reaping them during one module's refresh would delete edges other modules contributed. Items 86 and 106 have the same hole. Whoever gives copycode-resident nodes a per-module identity closes all three at once — worth knowing before designing a fix for either symptom alone.

    Candidate, unconfirmed: copycode CONTROL_FLOW nodes are shared by every including module (one node per .cpy line), and a .cpy that opens a block it does not close (YFRAMBC0.cpy opens FOR at line 65; the END-FOR lives in the includer) gets containment edges from every includer's nesting context accumulated onto that one shared node. Related: 625 nodes have endLine < startLine (531 DB_ACCESS, 94 CONTROL_FLOW) — the FOR above is 65 -> 15. But this explains only 26 of 162 cycles and 4 of 22 self-loops, so it is not the main cause. Left open on purpose rather than guessed at. The cycles are not confined to the statement tree: the DATA_STRUCTURE node named "," in BSUPLFN0.nat (itself a parser artifact worth its own look) carries self-loops, so the INCLUDES -> field walk the bare-field resolvers do runs through cyclic ground too. That is why item 77's INCLUDE_FIELD_DEPTH bound is a correctness requirement, not a tuning knob — an unbounded CONTAINS* there is what killed the probes above. (Item 77's redirect does not traverse this region: src for a placeholder edge is only ever FUNCTION (157,616 edges) or MODULE (35,572) — never CONTROL_FLOW — and all 16,231 such functions are direct CONTAINS children of a module, so its *0..1 bound avoids the cycles entirely.)

    2026-07-17 — concrete impact established and fixed for the dynamic-CALLNAT family (pending corpus re-verify). The blow-up is not merely theoretical: a whole-root deep refresh of upms wedged finalize step 17 (resolve-dynamic-callnat-intra-indirect) for ~2 h without completing, which blocks every later step — including item 77's bare-field resolution, so item 77 could not be corpus-verified. Measured cause: that step joins three unbounded (caller:MODULE)-[:CONTAINS*0..]-> anchors, and the resulting per-caller path enumeration over the cyclic copycode region is cubic — a read-only probe of the exact query for a single caller (DAGNTFN0) did not finish in 60 s. Fix: the six dynamic-CALLNAT resolvers (RESOLVE_DYNAMIC_CALLNAT_INTRA / …_INDIRECT / …_CROSS

    • their …_SCOPED variants in CypherQueries) no longer descend CONTAINS to find a module's own statements. A MODULE is 1:1 with its sourceFile (verified: 3587 files, max one module each), and a module's CALLNAT sites and WRITES statements all carry that same sourceFile, so the anchors become sourceFile-equality hash-joins that cannot cycle. The dispatch variable in the cross-module resolver can live in an included PDA, so it is scoped to the caller's own file or a data structure the caller INCLUDES (matching by name alone would pull in 20958 unrelated same-named vars; the scope filter keeps the 307 in-scope ones). Proven equivalent on upms: resolve-dynamic-callnat-intra yields the identical 31 resolved (caller, target, lineNo) triples project-wide, and the rewritten indirect step completes project-wide in ~6 s. Existing dynamic-dispatch ITs (intra / indirect / cross / scoped / unresolved- survival) stay green. Still to do: deep-recreate upms and confirm finalize reaches 36/36; then finish item-77 corpus verification. This does not remove the underlying CONTAINS cycles — other CONTAINS* traversals remain exposed if a future step joins several of them; the cycles themselves (parser line-range/shared-copycode artifacts) are still open above.

    2026-07-18 — same blow-up confirmed on the READ path (frontend-facing, NOT yet fixed). Only the finalize write-queries were rewritten above; the runtime read-queries the UI (ac-ui) renders still use unbounded (m:MODULE)-[:CONTAINS*0..]->(src). Measured live against server v71 on a heavy module (ACCNPE01): callees >2 min / hangs, digest >10 s timeout, context >10 s timeout, callers ~3.6 s; light modules (BMTABBP0) and db-accesses/call-tree/graph/functions stay <1.2 s. So the UI's Callees / module-overview / context panels spin on large modules. Same cure as the dynamic- CALLNAT fix (src.sourceFile = m.sourceFile hash-join, INCLUDES-scoped for variables). Affected read constants: CALLEES, CALLERS, DIGEST/CONTEXT query, FUNCTION_CALLERS, DB_ACCESSES, variable READS/WRITES, *_FOR_MODULES. 2026-07-19 — fixed. Rather than the sourceFile hash-join, the read queries use a tighter, provably equivalent bound: every {@code CALLS}/{@code READS}/{@code WRITES}/{@code DB_ACCESS}-parent edge source is a {@code MODULE} (depth 0) or a {@code FUNCTION} that is a direct {@code CONTAINS} child of the module (depth 1) — verified corpus-wide (0 sources deeper, 0 non-direct-child edge-source functions, and {@code DB_ACCESS} parents are only {@code FUNCTION}/{@code MODULE}, never {@code CONTROL_FLOW}). So (m)-[:CONTAINS*0..]->(src) becomes (m)-[:CONTAINS*0..1]->(src), which returns the identical set but cannot walk the cyclic copycode region. 20 read-side traversals updated (callees, MODULE_HOP_OUT(+ wiring), DISPATCH_TABLE, EGO_NEIGHBORS_*, VARIABLE_ACCESSES, DB_ACCESSES(+FOR_MODULES), SQL_STATEMENTS(+FOR_MODULES), FUNCTION_CALLERS, SEARCH_BY_VALUE(+_CONTAINS), fieldFlow, BUILD_CALLS_MODULE, FLOW_FRONTIER_SOURCE_FILES). Finalize/resolve queries left as-is (they completed). Guarded by ReadPathBoundedTraversalIT (EXPLAIN plan asserts no unbounded CONTAINS expand in callees); full IT suite green (193/0/0). No recreate needed — a query-only change against the existing graph. Verified live (v76): ACCNPE01 callees >2 min → 0.16 s, digest >10 s → 3.3 s, context **>10 s → 3.1 s; WGEAGB0S calleesunchanged (7). The underlyingCONTAINS` cycles (parser artefacts) still exist, but both the finalize and the read consumers are now bounded — item 75 no longer has a practical impact.

    2026-08-21 — cause established and half of it fixed (parser artefacts, scope A). The entry above said "cause NOT established"; it is now. There are exactly two families, on unrelated code paths:

    • A — parser artefacts (fixed). A Natural filler is declared <level><byteCount>X with no name (489X = level 4, 89 bytes). DATA_AREA_FIELD_EXPORT's occurrence/length column swallows part of the digits, so each filler line produced a DATA_STRUCTURE literally named X at a different, bogus level; the (type, name, sourceFile) merge key collapsed all of a file's fillers onto one node, which then contained itself. 27 such nodes in upms — 20 of the 24 self-loops. The "," node (BSUPLFN0.nat 11x, JB0067N0.nat 5x) is the same bug on the .nat source path: the continuation lines of a multi-line INIT<...> (166 , /*CREDIT NOTE) tokenize as level 166, name ",". Not cosmetic — a swallowed level pops the whole group stack, so the fields after a filler were silently re-parented under it: in SNA27R01.pda, C63P0027-GESCHL (line 28) hung under X instead of its real group. Fix: NaturalParser.DATA_AREA_FILLER skips filler lines on both data-area call sites (it requires at least one count digit, so a field genuinely named X still parses), and NaturalFieldTokenizer.opensMultiLineValue()/closesMultiLineValue() let both ingest tiers skip INIT</CONST< continuation lines — the grammar decision lives in the shared tokenizer so the two tiers cannot drift (item 59). Guarded by three tests that fail on the pre-fix parser — NaturalParserTest's fillerDeclarationProducesNoNodeAndDoesNotReparentFollowingFields and multiLineInitValuesAreNotParsedAsFields, plus NaturalCoarseScannerTest.multiLineInitValuesDoNotEnterTheIdentifierIndex — and NaturalParserTest.singleLineInitIsUnaffected, which pins the common single-line INIT<...> form against the new skip. The stale nodes leave the graph on the next deep refresh of upms.
    • B — shared copycode nodes (open, own round). The candidate above is confirmed: YFRAMBC0.cpy opens a FOR at line 65 and ends at 68 (the END-FOR is in the includer), so the one shared node accumulates every includer's nesting context — hence the 2-cycles with JX0031N0.nat:781/828/874/… and the 65 -> 15 range. Same in YFRAMBC4/CH/CI/CM.cpy and JX9901C6.cpy; JX0030C2.cpy:37 <-> 41 is the intra-file variant. Sized for the first time: 1185 shared .cpy nodes, avg 13.7 includers, max 773 -> a per-includer identity (ownerModule, as item 76 does for field placeholders) yields ~16,199 nodes, +3 % on upms's 459,093 — affordable. This also closes items 124/86/106. Drift measured on the current graph (vs. the 2026-07-17 numbers in the heading): self-loops 22 -> 24, two-cycles 162 -> 21, inverted ranges 625 -> 383. Scope A removes 20 self-loops and 2 two-cycles; the item stays open for B, so a corpus-wide CONTAINS-acyclicity assertion is not yet possible — the guards above are deliberately fixture-level.

    2026-08-22 — scope B shipped (per-module identity), and it does NOT fix the cycles. Copycode-resident nodes now carry the including module's file as ownerModule (GraphRepository.nodeOwner, generalising item 76's placeholderOwner): a node whose sourceFile differs from the parse's own module file came from a copycode — no extension test needed, since CopycodePreprocessor is the only thing that can produce one. MODULE/DB_TABLE are excluded: a module can be declared inside a copycode (ZDTSTBP6 in ZDTSTBC6.cpy) and every module lookup binds (project, name, sourceFile) but never ownerModule, so an owned MODULE node would be invisible to them. All five sweeps/reaps (DELETE_STALE_FILE_NODES, ..._RESOLVED_FIELD_EDGES, ..._NATURAL_TABLE_ACCESS_EDGES, ..._NATURAL_USING_EDGES, ..._NATURAL_CALL_EDGES) now key on the (sourceFile, ownerModule) pair instead of the file alone — mandatory, not cosmetic: the copycode file is in every includer's fresh-file set, so a file-only sweep would delete the other includers' nodes (one transaction older) together with their edges. That pair key also closes the documented scope limit of items 124/86/106: the 605 call edges whose source subroutine lives in a .cpy are now reapable, scoped to the re-parsed owner. Guarded by CopycodeNodeOwnershipIT (4 tests, 3 of which fail on the pre-fix store).

    Migration caveat, learned the hard way: a deep refresh does not clean up after an identity change. Legacy shared nodes key as (cpy, "") and no fresh parse produces that pair any more, so nothing sweeps them — after the refresh upms held 2217 legacy nodes beside 27,551 new per-owner ones (486,543 nodes). A recreate (item 78) is required whenever the node identity changes. After recreate?deep=true of upms/pur/app: 484,327 nodes, legacy shared down to 1 — exactly the excluded MODULE ZDTSTBP6.

    The cycles survive the recreate, so the current parser produces them — they are not stale-edge residue (a hypothesis that looked plausible because nothing ever reaps CONTAINS edges, and was wrong). Post-recreate: self-loops 15, two-cycles 16, inverted ranges 1007 (up from 383 — the 65 -> 15 range defect is no longer folded onto one shared node but exists once per includer: not worse, just no longer hidden). Every remaining cycle is intra-module — the copycode node already belongs to one host and its cycle partners are IF nodes of that same host. Cause:

    JX0031N0.nat  INCLUDE YFRAMBC0  16x      JE0018N0.nat  INCLUDE YFRAMBCH  4x
    JB0025N0.nat  INCLUDE YFRAMBCI   5x
    

    A module that includes the same copycode at many differently-nested sites gets one node for all of them (same file, same line), so the copycode's unterminated FOR accumulates the nesting context of every site. Per-module identity cannot separate those; per-include-site identity can.

    Scope C (implemented 2026-08-22, awaiting a corpus recreate) — identity per expansion site. ownerModule for a copycode-resident node is now <hostFile>#<includePath>. Not #includedAt, as first planned: item 104 makes includedAt the host's INCLUDE line at every nesting level, so all expansions of a member that a copycode includes repeatedly share it — VPARTC02.cpy includes L4NLOGIC 136 times, JB0028C7.cpy includes ISICDAYS 4 times. includePath (the full file:line>file:line chain) is the only unique site identity; verified present and non-empty on all 27,551 copycode nodes of upms (max 111 chars, avg 39.7). The key is a parser-set property, so the keys live in ac-parser-core's CopycodeProperties rather than as magic strings on both sides, and a missing/empty chain falls back to per-module identity (the item 75-B behaviour — a re-collapse, never an invented identity).

    Proven on a fixture, not yet on the corpus. CopycodeNodeOwnershipIT's HOSTB includes the same copycode at two differently-nested sites — the shape JX0031N0.nat has 16 times over. On the 75-B store that fixture produces 2 two-cycles; with per-site identity it produces 0, and 3 of the IT's 5 tests fail without the change. The corpus numbers (15 self-loops, 16 two-cycles) can only be confirmed by a recreate?deep=true.

    Projected cost (recursive include-chain expansion, cycle-guarded, depth-capped at CopycodePreprocessor.MAX_DEPTH): copycode nodes 27,551 -> ~77,349, i.e. upms 484,327 -> ~534,125 (+10.3%). Extremely skewed: JX0030N0.nat includes JX0030C2 91 times and accounts for +16,650 nodes on its own; XUPD009P.nat reaches YFRAMEC2 742 times transitively (but that member has 2 nodes). Two things will look worse afterwards and are expected: inverted ranges (endLine < startLine) multiply again from 1007, because each site's node carries the same 65 -> 15 defect, and JX0030N0 will own ~16,800 copycode nodes — its digest/context/graph endpoints need a latency check after the recreate. ownerModule also grows to ~80 chars on ~77k nodes, which is item 111d-2's territory (a hash would kill graph-side diagnosability, so plain text was kept deliberately).

    2026-08-23 — closed. Corpus-verified after recreate?deep=true of upms (server v252, 6311 files, full, incomplete=false): self-loops 15 -> 0, two-cycles 16 -> 0, and a bounded search for longer cycles (CONTAINS*3..6 from every multi-parent DATA_STRUCTURE) finds 0 as well. CONTAINS is acyclic on the corpus. Copycode nodes 27,551 -> 51,895 across 19,565 distinct owners, all but one site-keyed (hostFile#includePath); the one exception is the excluded MODULE ZDTSTBP6, by design. Project total 484,327 -> 508,670 (+5.0%), i.e. half the projected +10.3% — the projection expanded every include chain independently, but sites that resolve to the same includePath legitimately share a node. No legacy residue: the only 2 nodes without an ingestGen are the unresolved-target placeholders MODULE/DATA_STRUCTURE JE999 (empty sourceFile), which never carry one. Latency on the worst-case module JX0030N0.nat (91 INCLUDEs of JX0030C2) is unaffected: digest 0.83 s, context 0.25 s, graph 0.19 s. The corpus-wide acyclicity assertion mentioned above was not added as a test — the IT suite runs against Testcontainers fixtures, not the corpus, so it would have nowhere to live; CopycodeNodeOwnershipIT (5 tests, 3 failing pre-fix) stays the guard, and the corpus number is a measurement recorded here. Inverted ranges rose 1007 -> 1437 exactly as predicted (each site now carries the 65 -> 15 defect once) — that defect is tracked separately, see item 137. pur and app still hold 75-B-keyed nodes and need the same recreate?deep=true.

  • 140. In a project without JPA entities every Java DB_ACCESS is a false positive (found 2026-08-27 while investigating why app has no USES_TYPE edges; fixed 2026-08-27). A DB_TABLE node is created only from a JPA @Entity or a Panache active-record class (JavaParser.java:1297). With none in the project, not one DB_ACCESS candidate can resolve — while addDbAccessCandidate (JavaParser.java:1063) over-approximates on purpose: its read gate is mode != READ || isRepositoryReceiverName(...) || staticReceiver || ENTITY_MANAGER_TYPES..., and || staticReceiver admits every static call whose method starts with get/find/read/ list/count/... In a legacy codebase built on static utility classes that gate filters nothing.

    How much of the corpus this was:

    | project | DB_TABLE | DB_ACCESS | resolved | unresolved | |---|---:|---:|---:|---:| | upms (natural) | yes | 14,303 | 14,302 | 1 | | pur (java) | 157 | 3,951 | 1,937 | 2,014 (51%) | | ac (java) | 71 | 335 | 114 | 221 (66%) | | app (java) | 0 | 2,219 | 0 | 2,219 (100%) |

    app contains no @Entity, @Table, @Repository, JpaRepository, PanacheEntity, @Query or EntityManager anywhere, and its four java.sql imports are all Timestamp — it has no database access at all. Its 2219 candidates were led by UserContext.getCurrent() (561x), UpmsSessionUtils.getSupportDaten (140x), Config.getInstance() (84x). Natural, by contrast, is exact, because a READ/FIND/STORE is an access regardless of view resolution.

    The damage was agent-visible, contrary to the first version of this entry. That version said both endpoints join through the table so unresolved candidates never surface. True for db-accesses (plain MATCH); false for sql-statements, which uses OPTIONAL MATCH and returned all 76 candidates of ...vermittler.VermittlerGeschaeftsregeln as "table": null, "mode": "READ", "statement": "UserContext.getCurrent()".

    Fixed by a new enrichment step reap-java-db-access-without-tables (CypherQueries.REAP_JAVA_DB_ACCESS_WITHOUT_TABLES), after the three Java resolvers: if the project holds no DB_TABLE, delete its Java DB_ACCESS nodes. Java only — the same reaping in Natural would destroy real accesses.

    Deliberate limits of the chosen rule, all three pinned by JavaDbAccessNoEntityIT:

    • Project-level, not per node. It fixes app and nothing else: pur's 2014 and ac's 221 unresolved candidates keep leaking through sql-statements, because those projects have tables. 2219 of 4454 corpus-wide false positives are removed, roughly half. Whether the rest are false positives or real accesses whose entity lies outside the ingested root is unmeasured — that question, and any tightening of the heuristic itself, is untouched here.
    • A project whose DB access is exclusively native SQL has no entity, hence no table, so its accesses are reaped too. They were already unresolvable (RESOLVE_JAVA_QUERY_NATIVE_SQL matches an existing DB_TABLE, which only an entity creates), so this loses no working behaviour — but it turns a silent false positive into a silent false negative. Not present in any of the four projects; constructible.
    • Recovery needs a full refresh. Once such a project gains its first entity the reaper stops firing, but the deleted nodes only return for files that are actually re-parsed; changedOnly skips unchanged ones.

    Rejected alternative (the first proposal in this entry): marking the nodes unresolved: true instead of deleting them. Deleting loses the only available measure of how well the Java heuristic aims — for app that number is now recorded above, but future projects of this shape will not report it.

  • 99. DATA_AREA_FIELD mis-split data-area lines that carry a marker column — the field name was lost and the level invented (found and fixed 2026-07-27 while implementing item 98)

    Symptom. Natural data-area exports (.lda/.pda/.gda) contain lines with a marker character between the type/length columns and the level. For those lines the parser emits an invented level and uses the marker (or type) letter as the field name; the real field name never reaches the graph at all.

    Cause. DATA_AREA_FIELD is ^\s*(.*?)(\d)([#A-Za-z][#\w-]*)\s*(.*)$ — the prefix is non-greedy, so the regex takes the first digit followed by an identifier as the level. Normally that is right, because the prefix ends in <TYPE><spaces><LENGTH> and <LEVEL><NAME> follows directly:

    A        60  2##COMMAND      ->  prefix 'A        60', level 2, name '##COMMAND'   (correct)
    

    But when a marker is glued to the length, the regex stops too early — inside the length:

    A         4C 2#C-PADRE_START-PATTERN
        parsed  : level 4, name 'C'          <- '4' is the length, 'C' the constant marker
        correct : level 2, name '#C-PADRE_START-PATTERN'
    
    S0001A        50M 2CRITERIA
        parsed  : level 1, name 'A'          <- 'S0001' is the marker column, 'A' the type
        correct : level 2, name 'CRITERIA'
    

    Measured impact (whole upms corpus, 2726 data-area files):

    • 60 files affected, 931 lines mis-split.
    • Of those, 333 lines in 31 files produce a phantom group at level 1.
    • The invented names are nearly always marker/type letters: C (739×, the constant marker), A (117×), I (44×), N (20×), M (9×), B, P.
    • The invented levels come from the length digits and range from 0 to 6.

    Consequences.

    1. The real field name does not exist. search_identifier cannot find #C-PADRE_START-PATTERN; data_structure_fields shows a field called C instead.
    2. Group nesting collapses. The level is arbitrary, and at level 0 the groupStack is emptied completely (while peek().level() >= level) including the root — every following field in the file loses its parent.
    3. The wrapper root disappears. parseDataArea only creates the file-named root when topLevelNames reports exactly one top-level group. A phantom level-1 group makes it two, and a module's USING <area> then references a node that does not exist — exactly what happens at YLORDVL1.lda:30.
    4. Identifier-index pollution. 739 nodes named C across the corpus, which can collapse together at the sourceFile="" placeholder level.

    Relation to item 98. Item 98 was not blocked by this: its enricher deliberately joins on area.sourceFile instead of walking CONTAINS from the root, precisely because that root was not guaranteed to exist. That join stays — it is the more robust one regardless.

    Done. The export is column-oriented — [<occ>] <TYPE> <LENGTH>[<marker>] <LEVEL><NAME> — where <occ> is an occurrence/superdescriptor column (S0001, 0013) glued to the type and <marker> (C constant, M multiple-value, *) is glued to the length. New anchored DATA_AREA_FIELD_EXPORT is tried first, with the permissive DATA_AREA_FIELD kept as a fallback, so anything the anchored form does not recognise keeps its previous behaviour exactly — regression is impossible by construction. topLevelNames uses the same split, otherwise a phantom level-1 group would still suppress the wrapper root. Deliberately a targeted grammar, not a complete one: the zero-padded two-digit level in A 8*03COD-GENAGREE is left to the fallback, which already resolves it correctly.

    Measured over all 2726 data areas: 59 282 lines parse identically, 931 are corrected (in 60 files), 1 338 fall back to the previous behaviour verbatim. Tests: NaturalParserTest#dataAreaMarkerColumnsDoNotStealTheLevelAndName (one case per marker shape) and #dataAreaLinesWithoutAMarkerColumnAreUnchanged (regression guard — the source form, the export view marker and the plain field all depend on regex backtracking past the occurrence column, so they are pinned explicitly rather than assumed).

    Knock-on: YLORDVL1.lda regains a single top-level group, so its wrapper root reappears and its USING reference resolves — item 98's enricher now reaches VDB2-VERSIS_LISTORDER too.

    Open. What the * marker means is still unknown (the adjacent comment on A 1002* 2V25C6961-RECORD reads /* #01 - alte Länge, hinting at a superseded field). Both the old and the new grammar treat it as a live field, so including it changes no outcome — but if it marks a removed field, those declarations are wrong in the graph either way, which would be its own item.

  • 98. View aliases declared in a USING data area were reported as tables (2026-07-27, follow-up to item 95). Item 95's alias pre-scan is per-module over the copycode-expanded lines, but a LOCAL USING data area is a separate module, so a view declared there stayed unresolved: YGEAGBNH.nat:2617 does FIND (1) VDB2-VERSIS_GENAGREE with the view declared in YGEAGVL1.lda, so db-accesses reported the alias instead of VERSVW_GENAGREE. Root cause was deeper than cross-file scoping: parseDataArea did not recognise the data-area export view marker at all. In an export a view is V 1VDB2-VERSIS_GENAGREE VERSVW_GENAGREE DA:00,00… — there is no VIEW OF text, and the V prefix was read as a data type, so the view became a VARIABLE with no USES_TYPE to its DDM and no group push, letting its columns escape to the file root. 45 data areas use this form; 32 modules do DML on an alias only declared there. Done: (a) parseDataArea treats prefix V as a view — DATA_STRUCTURE + USES_TYPE to the DDM (first token of the rest), fields now nesting under it; (b) new enrichment steps resolve-view-alias-tables READS|WRITES + resolve-view-alias-access-nodes redirect the module's READS/WRITES and its DB_ACCESS node onto the real table, scoped by the module's own USING set — a name-based redirect would be arbitrary, since NEXT-VIEW alone is declared over 100 different tables corpus-wide. size(reals) = 1 leaves a contradictory USING set unresolved rather than guessed; the join is on area.sourceFile, not CONTAINS from the area root, because that root only exists when the file has one top-level group (see item 99). Tests: NaturalParserTest#dataAreaExportViewMarkerLinksToTheTableAndNestsItsFields, IT NaturalCrossFileViewAliasIT (incl. two modules resolving the same alias name to different tables); the enricher was verified load-bearing by disabling it and watching the IT go red.

  • 97. call-tree leaked the dynamic-call placeholder a manual override only hides (2026-07-27, third WGEAGB0S deep API audit). call-tree for WGEAGB0S listed #GETSHORT-MODUL — a variable (YGEAGGNH.nat:443, CALLNAT #GETSHORT-MODUL) — as a MODULE in the closure, while callees for the same module correctly reported only the resolved target YGEAGGN0. Cause: a manual override does not delete the marker edge to the variable-named placeholder, it sets manualHidden = true and relies on the read queries to suppress it (DELETE_DYNAMIC_CALLNAT_PLACEHOLDER_EDGES). callees/callers filter it; the BFS behind call-tree did not — MODULE_HOP_OUT/MODULE_HOP_OUT_WIRING did not even bind the relationship. Everything driven by that BFS inherited the pollution (graph, db-accesses?depth=N, sql-statements?depth=N). Done: both hop queries bind r and apply coalesce(r.manualHidden, false) = false, matching callees/callers. Characterization IT DynamicCallOverrideIT#callTreeHonoursTheOverrideLikeCallees (placeholder present → override → absent → reset → present again); verified red against the pre-fix query.

  • 96. Natural UPDATE(ref.) / DELETE(ref.) were dropped, hiding every access layer's write path (2026-07-27, third WGEAGB0S deep API audit). 34 statement sites across 13 of the 65 modules in the WGEAGB0S closure — every Y****MN0 CRUD module — produced no WRITES edge, so db-accesses showed them as read-only plus a single STORE. Cause: DB_WRITE's (?!\() guard (added by item 90 to stop a phantom (OLD.) table) suppressed the phantom but never recovered the real table, and DELETE was only handled in its SQL DELETE FROM form. Done: new DB_WRITE_BY_REF plus a pre-scan mapping each FIND/READ statement label to its (alias-resolved) table; an unresolvable reference still records nothing, so item 90's no-phantom guarantee holds — its two tests stay green unchanged and now serve as the negative cases. Shared with NaturalCoarseScanner so tier-1 and deep agree. Tests: NaturalParserTest#updateAndDeleteByReferenceResolveToTheEnclosingLoopTable, #byReferenceWriteWithoutAResolvableLoopNamesNoTable, NaturalCoarseScannerTest, IT NaturalViewAliasDbAccessIT. A label may also introduce a SQL SELECT loop rather than a FIND (YELEMMN0, YMULTMN0 hold their record that way) — those resolve through the FROM clause; #byReferenceWriteResolvesThroughALabelledSelectLoop. Verified on live upms: all 34 by-reference sites in the WGEAGB0S closure now recorded, 0 missing.

  • 95. Natural view aliases were reported as DB tables (2026-07-27, third WGEAGB0S deep API audit). db-accesses named the Natural view variable of a DML statement, not the DDM it is declared over: 58 rows across 13 of the 65 modules in the WGEAGB0S closure, 32 alias names standing in for 20 real tables. Worst effects — the generator's boilerplate alias NEXT-VIEW became one DB_TABLE node shared by 11 modules meaning 11 different tables (and reporting no columns), and 11 VDB2-*-VLOG aliases hid every write to VERSVW_LOGFILE, so "who writes the audit log?" answered nothing. Cause: VIEW OF was only recognised in parseDataArea (.pda files); parseModule — which parses every .nat — never built an alias map, and the DML branches passed the operand verbatim to dbTable(...). Done: VIEW_DECL pre-scan over the copycode-expanded lines feeds resolveViewAlias into the DB_WRITE/DB_READ branches; dbTable() now upper-cases (Natural is case-insensitive and DB_TABLE merges on the name). Shared with NaturalCoarseScanner so a shallow and a FULL module cannot report different names for the same statement. Tests: NaturalParserTest#viewAliasResolvesToTheUnderlyingTable, NaturalCoarseScannerTest, IT NaturalViewAliasDbAccessIT (incl. the same alias in two modules resolving to two tables). Verified on live upms: alias rows in the WGEAGB0S closure 58 → 1, the phantom NEXT-VIEW node gone, VERSVW_LOGFILE reachable for the first time. The remaining row is the cross-file case, item 98.

  • 94. call-tree no longer enumerates paths; followWiring usable again (2026-07-20, JX0034N0 ↔ MultiTableImportJob functional comparison). call-tree?followWiring=true timed out on pur at depth ≥ 2 (>120s; depth 1 already took 5.3s), which made the Java wiring closure unobtainable. Measured cause — a single quantified path pattern over CALLS|INJECTS|REFERENCES, bounded by maxDepth × (1 + internalBudget) (= 42 at depth 2), recovering each target's depth as min(#MODULE nodes on path) - 1, i.e. by enumerating every path. With the CHA-materialized wiring edges (items 31/92) that is combinatorial:

    | rawBound | Java, followWiring, depth 2 | Natural JX0034N0, depth 5 | |---|---|---| | 3 | 1.6s, 71 targets | — | | 4 | 1.7s, 71 targets | 2.6s, truncated (42) | | 6 | 18.9s, 71 targets | 3.2s, truncated (79) | | 8 / 12 | >120s | 2.2s / 2.4s, truncated (120/172) | | 21 | >120s | 3.9s, converged (182) | | 42 (production) | >120s | 3.2s, 182 |

    So the budget is necessary for Natural (whose result converges only near 21) and useless for Java (converged at 3) — lowering it globally would silently truncate Natural, the exact failure its javadoc warns about. The blow-up comes from the wiring edges, which JavaParser.addWiringEdges and the CHA steps only ever emit class-to-class, so they can never reach a FUNCTION. Fix: split the query. MODULE rows now come straight from the moduleDepths BFS (its hop index is the module-hop depth — verified equal to the old query's module set: 71/71 for Java, 46/46 for Natural), and only FUNCTION rows still traverse, CALLS-only and bounded within one module. Behaviour change: the budget can no longer hide a module whose call site sits behind a long internal PERFORM chain — DEPTHLEAF is now reported at depth 1, which also removes a standing contradiction with db-accesses/sql-statements, whose module set always came from the same BFS. truncated accordingly now means "some module's internal subroutine chain may be cut off". Covered by CallTreeTruncationIT (both tests).

    Measured live after deploy — call-tree?followWiring=true on MultiTableImportJob: depth 2 1.9s (was >120s) with the same 71 modules, depth 6 3.0s, depth 10 2.3s converging at 852 modules. On upms/JX0034N0 at depth 5 the module set grew 46 → 53 with nothing lost; the seven that had been hidden are NDBERR, NDBNOERR, USIX009N, USIX052N, USIX053N, YELEMGN0, YLITEMN0. They are real: USIX052N is CALLNATed by ISI173N0 (line 474), itself a direct callee of JX0034N0, so it sits at module depth 2; and YLITEMN0 was already reported by db-accesses?depth=5 as a via, which is the contradiction this item removes.

  • 93. Transitive db-accesses lost the DECLARES rows (2026-07-20, JX0034N0 ↔ MultiTableImportJob functional comparison). DB_ACCESSES resolves a table from three sources — READS/WRITES, an entity's own MAPS_TO, and a repository's repositoryEntity (item 32) — but its transitive counterpart DB_ACCESSES_FOR_MODULES (item 65) only ever had the first. The transitive view was therefore not a superset of the direct one: asking the same module with depth silently dropped its table. Minimal repro: db-accesses on MultiTableEntryEntity returns multi_table_entry/DECLARES, db-accesses?depth=1 on that same module returns []. Consequence: a Java caller's transitive db-accesses came back empty even though the entity it persists through maps to a real table, which made the Java side of a Natural↔Java DB comparison impossible to obtain from the API. Fix: DB_ACCESSES_FOR_MODULES now carries the same three UNION branches, with via naming the module that declares the table. SQL_STATEMENTS has no such branches, so SQL_STATEMENTS_FOR_MODULES needed no change (verified). Covered by JavaRepositoryOwnTableIT.entityTableAlsoResolvesInTheTransitiveView / repositoryTableAlsoResolvesInTheTransitiveView (both red before the fix).

  • 92. Java inheritance/CHA wiring: three defect classes fixed (2026-07-19, MultiTableImportJob deep API audit — manual Java source pass, project pur). Three systemic errors in the callees/wiring materialization, all found by comparing callees against source:

    • A — inherited INJECTS/REFERENCES lost their origin file (635× in the MTIJ closure, 255 with a lineNo past the caller file's end). LINK_REFERENCES_TO_SUBCLASSES/LINK_INJECTS_TO_SUBCLASSES (item 31) copied the base class's lineNo onto the subclass edge but never set originFile, so the callees sites (coalesce(r.originFile, source.sourceFile)) fell back to the subclass file. Worked example: MultiTableImportJob reports PurBatchJobListener REFERENCES lineNo=133, but that file has 93 lines — line 133 is in AbstractPurBatchJob.java. Same as the Natural copycode bug (items 66/91), for Java inheritance. Fix: materialized edges now carry originFile = coalesce(r.originFile, base.sourceFile) + inheritedFrom = base.name.
    • B — CHA fanned constructor calls out to subtypes (48× phantom CONSTRUCTOR callees). LINK_CALLS_TO_IMPLEMENTATIONS applied class-hierarchy analysis to callKind='CONSTRUCTOR' CALLS, so new ArrayList<>() produced a phantom → InputConstraintHolder [CONSTRUCTOR] (it extends ArrayList), and new BaseException() fanned out to every exception subtype. A constructor is statically bound. Fix: AND coalesce(r.callKind,'') <> 'CONSTRUCTOR'.
    • C — a qualified same-name supertype resolved to self (3× self-EXTENDS). DateUtils extends org.apache.commons.lang3.time.DateUtils (and NumberUtils/StringUtils) were resolved by simple name to the project's own same-named class → a DateUtils EXTENDS DateUtils self-loop that also poisoned the inheritance materialization. Fix: JavaParser.supertypeName keeps the FQN when a qualified supertype's simple name equals the declaring class's own name; the materializers additionally guard sub <> base. The parser fix stops new self-edges, but a non-wiping refresh leaves the old self-EXTENDS behind (both endpoints are the surviving class node, so node reconciliation never sweeps it — the edge gap item 86 closed for Natural), so a delete-self-inheritance-edges enrichment step reaps any self-EXTENDS/IMPLEMENTS edge project-wide before the inheritance graph is traversed.
    • Infra: a delete-synthetic-inheritance-edges enrichment step reaps all resolvedVia:'INHERITANCE' edges before the three materializers rebuild them, so a non-wiping refresh picks up the new properties/gates (otherwise MERGE ... ON CREATE never updates a pre-existing edge). Query-only fixes for A/B + reap; parser fix for C. IT JavaInheritanceWiringIT (3 tests). No response-shape change — originFile flows through the existing callees sites.callSiteFile.
  • 91. db-accesses/workfile-accesses/sql-statements carry copycode provenance (2026-07-19, third WGEAGB0S deep API audit — manual source pass). A DB or work-file access whose statement lives in an INCLUDEd copycode was reported with a copycode-local lineNo and no file context, so the number read as a line of the host module. Concretely: the DB2 sequence read SELECT … FROM SYSIBM-SYSDUMMY1 lives in USIX043C.cpy at lines 31/39/45/51/57; db-accesses for the 9 including modules (YAPRFMN0, YCUACMN0, YLITEMN0, YMODAMN0, YMTABMN0, YMULTMN0, YPRODMN0, YRAMOMN0, YUGRPMN0) reported those as bare lineNos that land on each host's own comment/DEFINE DATA lines. Root cause: the provenance was already on the READS/WRITES edge (item 66 stamps originFile/viaCopycode/includedAt on every edge in CopycodePreprocessor.remap, persisted via SET r += e.properties), but the DB_ACCESSES/ WORKFILE_ACCESSES/SQL_STATEMENTS queries never returned it — exactly the gap item 85 closed for functions and item 66 for callees/variables. Fix (pure query + DTO, no re-parse of data): db-accesses and workfile-accesses now return sites: [{lineNo, sourceFile, viaCopycode, includedAt}] (new AccessSite record) alongside the kept lineNos; sql-statements gains sourceFile + viaCopycode from the DB_ACCESS node's (remapped) file. includedAt is stored as a string, so the site queries wrap it in toInteger(...). Direct and transitive (?depth>0, *_FOR_MODULES) variants. REST auto-serializes the records; MCP returns the same DTOs; the CLI is a JSON passthrough — all in sync. IT WorkfileAndCopycodeFunctionIT#dbAccessSiteNamesTheCopycodeFileForCopycodeSourcedAccess (host FIND vs copycode FIND → sites[0].sourceFile/viaCopycode distinguish the two).

  • 151. search/identifier?priorityModule= pins the caller's module into the page; deterministic order (renumbered 2026-08-28 — this item had mistakenly also been given the number 91.) (2026-07-26, UI click-to-identify test). Click-to-identify sent search/identifier?name=&limit=25, but the query had no ORDER BY and paginated in incidental index order, so for a name declared in >25 modules (e.g. #I-LINE-LEV, 192 declarations) the open module's own declaration was truncated away and the popover falsely reported "0 in this module". Fix: SEARCH_IDENTIFIER gains $priorityModule — it does not filter (unlike module=) but computes a pinRank (0 for that module's file, else 1) and ORDER BY pinRank, sourceFile, startLine, so the local match survives the limit while the global list is preserved; ordering is now deterministic (it was undefined before). Delivered across REST (priorityModule), MCP search_identifier, CLI --priority-module, and the UI hook. IT IdentifierPriorityModuleIT (four modules sharing one LOCAL field, PRIO_ZZZ_TARGET sorts last: excluded at limit=2 without the pin, first in the page with it, and the full set still returned at a large limit — i.e. no filtering).

  • 90. DELETE no longer mis-parsed as a table write (2026-07-19, second WGEAGB0S deep API audit, Finding 5). Natural DML DELETE [(label)] deletes the current record of the enclosing READ/FIND loop and names no view, and the EXAMINE … DELETE [FIRST] clause is not a DELETE statement at all — but DB_WRITE captured the token after DELETE as a table, producing phantom FROM (108×, from SQL DELETE FROM <table>), (OLD.)/(*)/label refs (19×), and FIRST (7×) — and, for SQL, lost the real table (it sat after FROM). Fix: DELETE removed from DB_WRITE (STORE/UPDATE keep their view operand); a new DB_DELETE_FROM captures the SQL DELETE FROM <table> real table (mode DELETE); DELETE (label)/DELETE FIRST name no table. The same label-reference shape also affects UPDATE (label) (UPDATE (OLD.)/(HOLD-PRIME.)) and a (*) read operand — a view never starts with (, so DB_WRITE/DB_READ now carry a (?!\\() guard that rejects a parenthesized label reference (no phantom (OLD.)/(*) table) while UPDATE <view> still records the real view. Both parsers. Tests in NaturalParserTest (DELETE FROM, DELETE (OLD.), EXAMINE … DELETE FIRST, UPDATE (OLD.)). Follow-up (2026-07-19, final verification): one last (*) phantom survived, from the SQL SELECT parser, not DB_READ. A SELECT column list can contain a hyphenated Natural field whose last segment is literally FROM (YCOMIROW.DAT-CALC-FROM (*)); FROM_VIEW = \bFROM\s+(\S+) treated the hyphen as a word boundary, captured the trailing (*) as a phantom table, and — since the FROM view binds on the first match only — swallowed the real FROM VERSVW_COMISION clause below. Fix: FROM_VIEW now uses a negative lookbehind (?<![-\\w.])FROM (FROM must be a standalone SQL keyword, not an identifier tail) plus the (?!\\() operand guard. Both parsers. Test NaturalParserTest#sqlSelectColumnEndingInFromDoesNotShadowTheRealFromClause.

  • 89. READ WORK <n> (FILE keyword omitted) recognized as work-file I/O (2026-07-19, second WGEAGB0S deep API audit, Finding 4). The item-84 guard only matched READ WORK FILE; the corpus also writes READ WORK 1 ONCE RECORD … (289× project-wide) without FILE, which still fell through to DB_READ and produced a phantom DB_TABLE 'WORK' (154 accesses). Fix: the WORK [FILE] n guard and the WORKFILE_ACCESS/WORKFILE_DEFINE patterns now treat FILE as optional (guarded on a following digit, so a view whose name merely starts with WORK is unaffected). Both parsers; NaturalParserTest (READ WORK 1 ONCE RECORD).

  • 88. Finalize sweep deletes edgeless DB_TABLE/WORKFILE placeholder nodes (2026-07-19, second WGEAGB0S deep API audit). Companion to item 86: reaping a stale access edge left the placeholder node (sourceFile="", never node-swept) behind with degree 0 — invisible to db-accesses (edge-driven) but still surfacing in search_identifier?type=DB_TABLE and the DB-table inventory (observed: orphaned NUMBER/WORK/FIRST after items 84/87). New finalize step delete-orphaned-placeholder-tables (DELETE_ORPHANED_PLACEHOLDER_TABLES) removes any DB_TABLE/ WORKFILE with no relationships; degree-0 only, so a table any file still accesses (or a Java @Entity's MAPS_TO target) is kept. Covered by StaleTableEdgeReapIT.orphanedPlaceholderTableNodeIsDeleted.

  • 87. FIND NUMBER <view> no longer mis-parsed as a phantom DB_TABLE 'NUMBER' (2026-07-19, second WGEAGB0S deep API audit). FIND NUMBER <view> is a count-only FIND (natural-grammar.md §7.1); NUMBER is a statement keyword, not the accessed view — but the DB_READ regex (in both NaturalParser and NaturalCoarseScanner) captured it as the table name, so db-accesses reported a bogus NUMBER table and lost the real view (e.g. CON-DB2-AUTHPROF-USED-IN-AUTHSPC, NEXT-VIEW). Seen on 11 modules of the WGEAGB0S call tree (YAPRFMN0, YCUACMN0, YENTIMN0, YGARAMN0, YLITEMN0, YMODAMN0, YMTABMN0, YPRODMN0, YRAMOMN0, YTABLMN0, YUGRPMN0; 13 accesses; 182 FIND NUMBER occurrences project-wide). Same class as item 84's READ WORK FILE. Fix: DB_READ now skips the FIND options ALL/FIRST/NUMBER/UNIQUE and the RECORDS/IN/FILE noise words before the view (and allows a variable record-limit (operand), not just a literal). Covered by NaturalParserTest (FIND NUMBER MY_VIEW / FIND NUMBER IN FILE OTHER_VIEW). Item 86 reaps the existing NUMBER edges on the next refresh.

  • 86. Re-ingest reaps stale Natural DB_TABLE/WORKFILE access edges (self-healing) (2026-07-19, follow-up to items 84/85). A statement whose access target changed between parses orphaned its old READS/WRITES edge forever: the target is a placeholder (sourceFile="", never node-swept) and the source node survives, so neither the item-58 node sweep nor the target-keyed edge MERGE reaped it. Seen as the WORK db-access that lingered on USIX052N after the item-84 parser fix (a pre-fix READ WORK FILE→DB_TABLE 'WORK' edge), and it applies to any edited view name too. Fix: before re-merging a re-parsed Natural file's edges, DELETE_STALE_NATURAL_TABLE_ACCESS_EDGES drops its READS/WRITES edges to DB_TABLE/WORKFILE placeholders; the fresh parse (which always re-emits them) re-creates the current ones, unchanged ones round-trip identically. Scoped to language:'natural' source nodes — Java DB edges are resolver-built (RESOLVE_JAVA_DB_ACCESS) and untouched. Covered by StaleTableEdgeReapIT (edit READ VERSVW_OLD → READ VERSVW_NEW + READ WORK FILE, assert old view + WORK gone). (A clean re-ingest of upms already cleared the existing stale WORK; item 86 prevents recurrence on incremental refreshes.)

  • 84. Natural work-file access tracking (workfile-accesses) + READ WORK FILE no longer a phantom DB table (2026-07-19, WGEAGB0S deep API audit, Finding 2). READ WORK FILE n <buf> (sequential flat-file I/O) was matched by the (READ|FIND) <view> DB pattern in both NaturalParser and NaturalCoarseScanner, creating a bogus DB_TABLE 'WORK' READS access (seen on USIX052N in the WGEAGB0S call tree). Both DB_READ patterns now negative-lookahead WORK FILE, and READ/WRITE WORK FILE are modelled as first-class WORKFILE + WORKFILE_ACCESS nodes (analogue of DB_TABLE/DB_ACCESS), keyed by work-file number, with the record buffer on the READS/WRITES edge and the DEFINE WORK FILE n '<name>' physical name on the node. New GET /modules/{name}/workfile-accesses → [{workFile, physicalName, mode, recordBuffers, lineNos}], MCP workfile_accesses, CLI ac workfile-accesses. Covered by NaturalParserTest (READ + WRITE) and full-stack WorkfileAndCopycodeFunctionIT; mcp-api-usage/system-prompt docs updated.

  • 85. /functions items carry sourceFile + viaCopycode (copycode-provided subroutine provenance) (2026-07-19, WGEAGB0S deep API audit, Finding 1). A subroutine pulled into a module via INCLUDE was listed with declaredIn=the including module and the copycode's startLine/endLine but no file, so the lines pointed outside the module's own (shorter) file — e.g. ISIN0019 (59-line file) reported GET-FORMAT at 90–120, which actually live in ISIC0010.cpy. The FUNCTION node already stored the right sourceFile; the MODULE_FUNCTIONS/MODULE_FUNCTIONS_OWN/_INHERITED projections just dropped it. Now InheritedFunction/FunctionInfo expose sourceFile (+ derived viaCopycode = f.sourceFile <> m.sourceFile). Covered by WorkfileAndCopycodeFunctionIT; docs updated.

  • Module callers default is external-only; no MODULE self-loop (2026-07-19, WGEAGB0S deep API audit). Two coupled defects in CypherQueries.callers(scope): (1) the top-level main body's PERFORMs originate at the MODULE node, so scope=internal/default reported the module as its own caller (WGEAGB0S → WGEAGB0S), a self-loop callees never mirrors — fixed with AND caller <> m on the internal scope; (2) the default (scope=null) merged external callers with intra-module PERFORM wiring, so a module's own subroutines showed up as its "callers" — the default now maps to external (genuine incoming CALLNAT/inheritance only). context/digest (both call callers(…, null)) inherit the clean view; scope=internal still exposes function→function PERFORM wiring; callees unchanged. Covered by ModuleCallersSelfLoopIT (fails 2/3 before the fix). MCP callers tool description + REST endpoint doc + agent-api-usage-ac-implementation.md updated. (Supersedes the earlier "Not a bug (verified): context.callers includes internal PERFORM callers — noisy but accurate" note.) Ego-graph direction=in for WGEAGB0S now returns its dynamic callers W-LST-N0/W-MNT-N0 (item 75). Payload direction is always REQUEST for PDA-derived contracts (a single interface PDA doesn't encode direction) — a documented limitation, not a bug.

    • Follow-up (2026-07-19, WGEAGB0S call-tree closure re-audit): the "external-only" default was only half-fixed. The default view still returned the calling FUNCTION node (a subroutine/method) whenever the call originated inside a subroutine rather than the main body — ModuleCallersSelfLoopIT missed it because its fixture caller CALLNATs from the main body, where the edge already starts at the MODULE. Measured on upms: 46/56 modules in the WGEAGB0S closure had FUNCTION-typed rows in the default callers, 29/56 had duplicate rows, and hot utilities were unusable (CDRANGE default callers = 500 rows / 3 distinct FUNCTION names / 0 module callers). Fixed by making the external branch of CypherQueries.callers(scope) roll every caller up to its owning MODULE via (callerModule:MODULE)-[:CONTAINS*0..1]->(source)-[r]->(m) and collect(DISTINCT …) — symmetric with how callees anchors its source side. Covered by new CallersRollupIT (caller invokes from inside a subroutine, twice → one rolled-up MODULE row with two aggregated sites, no FUNCTION leak); the old ModuleCallersSelfLoopIT, JavaWiringIT, JavaModulesExtendsFilterIT stay green.
  • 100. DEFINE DATA ... USING <member> binds by level-1 record name, not by member (file) name (found and fixed 2026-07-28, WGEAGB0S deep API audit tier 2 — 19 of 379 USING sites (5.0%) in the WGEAGB0S call-tree closure are wrong or unresolved)

    Symptom, two shapes.

    1. Wrong file. GET /api/projects/upms/modules/WGEAGB0S/data-structures reports
      W-WIF-A2  USING  PDA  fieldCount=10  src/manual/parameter_data_area/old/W-WIF-A7.pda
      
      Ground truth: WGEAGB0S.nat:47 says PARAMETER USING W-WIF-A2, i.e. member W-WIF-A2 = new/W-WIF-A2.pda (5 fields: P-LINE-TYPE/LEVEL/KEY/VALUE) — exactly the fields the module uses at 386, 732–735 and 1286. old/W-WIF-A7.pda is a different member whose level-1 record was copy-pasted as 1W-WIF-A2; it holds P-REST-*, which WGEAGB0S reads from W-WIF-A1 (verified: W-WIF-A1 carries P-REST-FLAG, P-REST-LEVEL-IND, P-REST-POINT-KEY, P-LINE-START, P-LINE-END). So both sourceFile and fieldCount are wrong. Same shape: BGEAGFN0/USIX052N USING YFRAMBL0 → old/ZFRAMBL0.lda instead of new/YFRAMBL0.lda, and USIX052N USING YFRAMBL1 → new/ZFRAMBL1.lda instead of old/YFRAMBL1.lda. Not cosmetic: YFRAMBL0/ZFRAMBL0 and YFRAMBL1/ZFRAMBL1 differ in the browse-array bound V (CONST<13> vs CONST<1000>), so an agent reading the wrong twin gets the wrong page size.
    2. Never resolved at all. USING VLAYERLA, USING USIX020L, USING USIX036L report sourceFile: null, area: UNKNOWN, fieldCount: 0 — although all three .lda files exist inside the project root and are ingested. 15 of the 19 affected sites are this shape (10× VLAYERLA in ISI173N0/VMULTDN1/VMULTGN1/VMULTMN1..4/VMULTON1/VMULTSN2/ZINELEM1, 2× USIX020L in ISI173N0/USIX021N, 3× USIX036L in YGARAMN0/YMODAMN0/YPRODMN0). search/identifier shows only the sourceFile: "" placeholder for each.

    Cause, two cooperating places.

    • NaturalParser.parseDataArea (ac-parser-natural/.../NaturalParser.java, ~line 920) emits the member-named wrapper root only for the single-top-level case:
      if (topNames.size() == 1 && !topNames.get(0).equalsIgnoreCase(areaName)) { … }
      
      A data area with several level-1 records therefore gets no node named after its member, so a USING of it can never resolve. VLAYERLA.lda (constants), USIX020L.lda (1#C-HM-FUNC, …) and USIX036L.lda (1#V-ID-TRAN_TAB_FWD, …) are exactly that case. The existing comment claims "multi-top-group areas keep their existing shape (no wrapper, no name collision)" — that decision is what produces shape 2.
    • CypherQueries.resolvePlaceholderTargets (ac-neo4j-store/.../CypherQueries.java, ~line 2650) matches a placeholder purely on (type, name, project). It already excludes module-owned groups (item 74) but has no preference for the node that is the member root of a data-area file, and no tie-break when two files declare the same level-1 name — so USING W-WIF-A2 matches the level-1 node inside W-WIF-A7.pda just as well as the root of W-WIF-A2.pda. In upms 42 level-1 names are declared in more than one data-area file, so this is not a one-off.

    Fix. (a) In parseDataArea, emit the member-named root whenever no level-1 record already carries the member name (!topNames.contains(areaName)), parenting every level-1 record under it — so each data-area file contributes exactly one node named after its member. (b) In resolvePlaceholderTargets, for DATA_STRUCTURE placeholders prefer a real node that is a data-area member root (real.sourceFile ends .lda/.pda/.gda and its basename equals real.name), falling back to the current global name match only when no such candidate exists — so coverage never regresses for USINGs that have no matching file. Characterization tests: NaturalParserTest case for a multi-top-level .lda (must yield a member-named root containing all top-level records), plus a Testcontainers IT with two fixture PDAs declaring the same level-1 name (the USING must bind to the file whose member name matches).

    Done. Both halves landed as described. NaturalParserTest.multiTopLevelDataAreaStillGetsAMemberNamedRoot

    • …dataAreaWhoseTopLevelAlreadyMatchesTheMemberKeepsItsShape cover the parser; DataAreaMemberResolutionIT covers the graph end to end (DAOTHER.pda carries a copy-pasted 1DAMEMBER record, DAMULTI.lda has several level-1 records and none named after the member). parsesLdaWithMultipleTopLevelStructures was updated: its "top-level structures have no CONTAINS parent" assertion encoded exactly the behaviour this item changes.
  • 101. /data-structures/{name}/fields silently unions homonymous definitions from different files (found and fixed 2026-07-28, WGEAGB0S deep API audit)

    Symptom. GET /api/projects/upms/data-structures/W-WIF-A2/fields returns 15 fields — the union of new/W-WIF-A2.pda (5) and old/W-WIF-A7.pda (10) — with no sourceFile on any row and no parameter to disambiguate. An agent cannot tell that it is looking at two unrelated record layouts merged into one, and will happily "verify" a field that the module it is analysing cannot see.

    Cause. CypherQueries.DATA_STRUCTURE_FIELDS collects every same-named definition into canon and UNWINDs it:

    MATCH (s0:AstNode {type: 'DATA_STRUCTURE', name: $name, project: $project})
    WITH collect(s0) AS defs
    WITH [d IN defs WHERE d.sourceFile <> ''] AS withFile, defs
    WITH CASE WHEN size(withFile) > 0 THEN withFile ELSE defs END AS canon
    UNWIND canon AS s
    

    The Javadoc and agent-api-system-prompt.md both claim the opposite ("scoped to the structure's own definition"), so the contract is documented as something the query does not do.

    Fix. Return sourceFile on every row, accept an optional ?sourceFile= filter, and when the parameter is absent and more than one definition exists prefer the member-root definition (item 100) rather than the union. Covered by an IT with two fixture areas declaring the same level-1 name.

    Done. DataStructureField gained sourceFile (so db-tables/{name}/columns, which shares the record, returns it too); REST ?sourceFile=, MCP data_structure_fields(sourceFile) and CLI ac data-structure-fields --source-file delivered together. Covered by DataAreaMemberResolutionIT.dataStructureFieldsAreNotUnionedAcrossHomonymousDefinitions (fails before the fix: the decoy record's fields are merged in) and …CanBePinnedToOneSourceFile.

  • 102. module_data_structures collapses homonyms into one row with an arbitrary file and a blended fieldCount (found and fixed 2026-07-28, WGEAGB0S deep API audit)

    Symptom. The W-WIF-A2 row on WGEAGB0S reads sourceFile = old/W-WIF-A7.pda, fieldCount = 10 — a combination that is wrong even if you accept either candidate as the intended one, because the file and the count can come from different nodes.

    Cause. CypherQueries.MODULE_DATA_STRUCTURES groups by d.name only and then reports

    WITH d.name AS name, rel, max(fc) AS fieldCount,
         head([sf IN collect(d.sourceFile) WHERE sf <> '']) AS sourceFile
    

    head(collect(…)) is order-dependent (nondeterministic across ingests) and max(fc) is taken over all homonyms, so the reported sourceFile and fieldCount need not describe the same definition.

    Fix. Group by (name, sourceFile) and return one row per resolved definition. With item 100 in place this normally collapses back to a single row; where it does not, the agent sees the ambiguity instead of a fabricated blend. Covered by the same IT as item 101.

    Done. A still-unresolved placeholder is reported only when no resolved definition exists for that (name, relationship), so the sourceFile: null / area: UNKNOWN row keeps its meaning. Covered by DataAreaMemberResolutionIT.moduleDataStructuresReportsOneRowPerUsedDefinition (before the fix: [DAMEMBER.pda, DAOTHER.pda] collapsed to one arbitrary row).

  • 103. db-accesses silently truncates at the default limit=50 — a bare array with no truncated signal (found and fixed 2026-07-28, WGEAGB0S deep API audit)

    Symptom. Measured on upms:

    GET /modules/WGEAGB0S/db-accesses?depth=10             → 50 rows
    GET /modules/WGEAGB0S/db-accesses?depth=10&limit=500   → 64 rows
    

    The 14 dropped rows hide 7 tables entirely — VERSVW_MULTILIN, VERSVW_MULTTABL, VERSVW_PRODUCTO, VERSVW_RAMO, VERSVW_TABLAS, VERSVW_USERGRP, VERSVW_USUARIO_NEW. The response is a bare JSON array: no envelope, no total, no truncated flag, so the caller cannot detect the cut. An agent following the documented Natural playbook (db-accesses?depth=3, no limit) therefore gets a silently incomplete DB footprint — the single most damaging failure mode for a reengineering or impact analysis, because it looks like a complete answer.

    Cause. AnalysisResource.dbAccesses applies effectiveLimit(limit), which defaults to 50. sql-statements on the same closure has no such cap (389 rows returned uncapped), so the two endpoints disagree about the same data. workfile-accesses shares the defaulted limit.

    Fix. Remove the default cap on db-accesses/workfile-accesses so they match sql-statements.

    Done — narrower than first proposed. Only the default cap was removed (AnalysisResource.uncappedLimit

    • the same helper in McpQueryTools); the {items, total, truncated} envelope was not added. Rationale: the defect is silent truncation, i.e. a cut the caller never asked for and cannot detect. An explicit limit is neither — the caller chose it — so wrapping the response would change the array contract for every REST/MCP/CLI/UI consumer to signal something already known. If a truncated flag is wanted anyway, it should be a separate item covering all paginated endpoints, not just these two. Covered by IncludeProvenanceAndAccessLimitIT.dbAccessesAreNotSilentlyTruncatedAtFifty (60-table fixture; returns 50 before the fix) and …explicitLimitIsStillHonoured.
  • 104. includedAt is ambiguous for nested copycode includes — it can point into an intermediate file that the response never names (found and fixed 2026-07-28, WGEAGB0S deep API audit)

    Symptom. GET /modules/ISI173N0/callees reports the YFRAMN04 callee as viaCopycode: 'YFRAMC01', includedAt: 27. Line 27 of ISI173N0.nat is a comment. The real chain is

    ISI173N0.nat:232  INCLUDE USIX050C 'YFRAMMC1' …
    USIX050C.cpy:58     INCLUDE &1&                 (parameterised)
    YFRAMMC1.cpy:27       INCLUDE YFRAMC01
    YFRAMC01.cpy:12         CALLNAT 'YFRAMN04'
    

    includedAt: 27 is a line in YFRAMMC1.cpy — an intermediate file that appears nowhere in the response. In the single-level case (WGEAGB0S → ADLML02, includedAt: 673 = the INCLUDE ISIYESNO statement) the same field is a line in the module's own file, so the field silently means two different things and the caller cannot tell which. (The resolution itself is correct — the parameterised 3-level expansion is followed properly; only the provenance reporting loses the chain.)

    Fix. Report the chain rather than one line: make includedAt always the line in the module's own file (232 here) and add a full includePath: [{sourceFile, lineNo}, …] on the sites entries of callees/callers/db-accesses/workfile-accesses. Covered by an IT over a 2-level include fixture.

    Done. CopycodePreprocessor.LineOrigin now carries the host include line unchanged through every nesting level plus the chain as a compact file:line>file:line edge property (Neo4j properties cannot hold a list of maps); CallSite/AccessSite expose it as List<IncludeStep>. sites is a nested response object, so MCP and the CLI pick the new field up without a signature change. Covered by IncludeProvenanceAndAccessLimitIT.nestedIncludeReportsHostLineAndTheWholeChain (before the fix: includedAt = 2, the line inside the intermediate .cpy — the ISI173N0 shape exactly) and …directCallHasNoIncludeChain.

  • 106. A module's resolved USING edges are never reaped, so an old binding survives every refresh (found and fixed 2026-07-28 while re-verifying item 100 against the live upms graph)

    Symptom. After the item-100 fix had landed and upms had been rebuilt and deep-refreshed, WGEAGB0S USING W-WIF-A2 reported two rows — the correct new/W-WIF-A2.pda (5 fields) and the pre-fix old/W-WIF-A7.pda (10 fields). The fix had not failed: the correct edge was created, the wrong one simply was never removed. Measured across the WGEAGB0S closure: the 19 wrong/unresolved USING sites dropped to 5, and all 5 residuals were this shape — a stale edge sitting next to the right one (BGEAGFN0/USIX052N → YFRAMBL0, USIX052N → YFRAMBL1 ×2, WGEAGB0S → W-WIF-A2).

    Cause. Item 86 reaps a re-parsed Natural file's stale access edges, but only those pointing at a placeholder (sourceFile = ""). A USING edge is resolved onto a real DATA_STRUCTURE node by resolvePlaceholderTargets, and from then on nothing deletes it: MERGE only ever adds. So any binding an older ingest made is permanent. This is not specific to item 100 — plain editing of a module's DEFINE DATA ... USING list leaves the dropped data area attached forever, which is the more common everyday case.

    Fix. DELETE_STALE_NATURAL_USING_EDGES, the INCLUDES counterpart of item 86, run in the same spot (right before the fresh edges are merged). Deleting all of a re-parsed file's USING edges is safe because the fresh parse always re-emits every one of them as a placeholder and the same finalize re-resolves them; unchanged ones round-trip identically. Covered by DataAreaMemberResolutionIT.anEditedUsingDropsTheOldBindingOnRefresh (edit USING DAMEMBER → USING DAOTHER, refresh, assert the old binding is gone — fails before the fix). Ordered last in that class because it mutates the fixture.

    Note: the fix prevents recurrence; it does not retro-clean a graph that already carries such edges — those disappear on the next refresh of each affected file, since the reap runs per re-parsed file. upms still carried the 5 residuals at the time of writing and needs one more refresh.

  • 105. A fan-out query that surfaces a data-area file re-ingests it on every call — search/identifier took ~60-75 s per lookup (found 2026-07-28 WGEAGB0S deep API audit, root-caused and fixed the same day)

    Symptom. GET /search/identifier?name=X on upms: ZFRAMBL0 (1 hit) 56.8 s / 62.2 s / 74.7 s across runs, YFRAMBL1 (9 hits) 58.2 s. Both the system prompt and the Natural playbook recommend this endpoint for orientation and for confirming a candidate is a real module; at that latency it cannot be used in a loop, and a batch of lookups exceeds a 2-minute client timeout.

    Cause — not the query. The first suspicion recorded here (a missing/unusable (project, name) index) was wrong, and the measurement that settled it is worth keeping:

    | call | result | |---|---| | ?name=NOSUCHNAME12345 (0 hits) | 0.90 s | | ?name=ZFRAMBL0 (1 hit, a .lda) | 74.7 s |

    Identical scan work, 80× the latency — so the cost is not the scan. (The scan is a full label scan: PROFILE shows 2,000,905 DbHits, because the name predicate is wrapped in a CASE that strips a leading Natural sigil and so cannot use ast_node_project_name. But that is ~1 s, and it is the same ~1 s in both rows above. Worth its own item if 1 s ever matters; it is not this bug.)

    The real cost is in withFanoutWarm → DeepIngestCoordinator.ensureDeepMany, which deep-ingests the source files a result set surfaced. FULLY_INGESTED_SOURCE_FILES asks for a MODULE node with ingestDepth = 'FULL' — but a Natural data area (.lda/.pda/.gda) produces only DATA_STRUCTURE nodes, never a MODULE. So a data area can never be reported as fully ingested, is treated as pending on every call, and is re-warmed forever. Server log for one lookup:

    07:11:29,595  Ingesting DATA_STRUCTURE ZFRAMBL0, YFRAMBL0 ... [old/ZFRAMBL0.lda]   <- 40 ms
    07:11:29,635  Finalizing project 'upms' (1 files persisted, scoped deep to 0 modules)
    07:11:29,635  Finalize upms (scoped-deep): 45 steps
    07:12:31,175  <next request>                                   <- ~60 s in finalize
    

    The warm itself is trivial; each one drags a whole-project 45-step finalize behind it and then reports "changed", so the caller re-runs its query on top. A repeat lookup re-ingested the same file again — it never converges.

    Fix. DeepIngestCoordinator.ensureDeepMany filters .lda/.pda/.gda out of the warm candidate set (isWarmable). Nothing is lost: a data area has no deep tier — both ingest tiers run the same parseDataArea — so warming one can never add anything to the graph. Applies to every withFanoutWarm caller (search/identifier, callers, callees, call-tree), not just this endpoint.

    Covered by DataAreaMemberResolutionIT.surfacingADataAreaDoesNotReIngestItOnEveryCall, which asserts the node id is stable across two lookups — ids are regenerated on every re-ingest, so a changed id is the re-ingest. It fails before the fix with two different UUIDs. Timing is deliberately not asserted: on a small fixture the finalize is fast, so a wall-clock bound would not reproduce the bug.

    Note: two further latency questions were surfaced by this and left open on purpose, not folded in: the 2M-DbHit label scan above, and why a scoped 1-file finalize runs all 45 project-wide steps (scoped deep to 0 modules). The second is the larger prize and affects every incremental ingest.

Natural parser robustness

(Item 59 — shared field-declaration tokenizer — completed 2026-07-15; items 61 — comments parsed as CALLNAT targets — 62 — data literals as false MODULE call targets — and 63 — CALLNAT matched inside a string literal — completed 2026-07-16. All moved to x-docs/features.md. The unanchored-CALLNAT family (#61 comments / #62 data literals / #63 string literals) is closed, and the shared lexical helpers now live in NaturalLines + NaturalFieldTokenizer so a fix lands in both ingest tiers at once. Items 120 and 121 below reopened this track: both sit in the copycode argument parser, which the earlier work never touched. All three — 120, 121 and the 123 found while validating them — completed 2026-08-07; measured after the deep refresh on 2026-08-09, they recovered 2202 copycode-derived call pairs (7319 → 9521, +30%) with none lost. Item 123 is the one to re-read before writing the next lexical pattern: the grammar doc that patterns are supposed to be written from was itself wrong, so following the convention correctly reproduced the bug.)

  • 120. INCLUDE arguments on continuation lines are never bound — 263 call edges silently absent (found 2026-08-06 in the upms source-vs-API cross-check, done 2026-08-07)

    Symptom. A Natural INCLUDE may spread its positional arguments over several lines. Only the first line is read, so every argument beyond it stays unbound, &3& survives substitution verbatim, and the CALLNAT &3& it feeds is dropped.

    VCOMIN50:1241  INCLUDE YFRAMBC8 '"AGNT-CHG-CMP-SP"'
    VCOMIN50:1242    '"YAGCHBN0"' 'YAGCHKEY' 'YAGCHROW' 'YAGCHPRI'   ← &2& lives here
    
    GET /modules/VCOMIN50/callees                      → YAGCHBN0 absent
    GET /modules/VCOMIN50/reaches?target=YAGCHBN0      → {"reachable": false}
    GET /modules/YAGCHBN0/callers                      → BAGCHFN0 only
    GET /dynamic-calls/unresolved                       → no VCOMIN50 entry either
    

    Scale. Of 25 175 INCLUDE statements in upms, 4697 have continuation lines. Checking every candidate pair against the live API: 263 of 264 module→module call edges are missing, across 154 calling modules and 89 targets. This lands hardest on the browse/access layer, because that is exactly where the idiom is used — YCARPBN1 and YPOLIBN1 report 0 callers, YCOMIBNH reports 1 while 29 files name it.

    Why it is the bad kind of wrong. The call is missing from callees, callers, call-tree and reaches and from /dynamic-calls/unresolved, so nothing anywhere says "not analysed". Same failure class as items 103 and 107: unanalysable reads as "nothing found".

    Root cause. CopycodePreprocessor.java:29 — INCLUDE_STMT is anchored to a single line — and :74, which passes only inc.group(2) to parseArgs. Fix: collect following lines that consist solely of literals/tokens before parsing arguments. Note substitute() deliberately leaves unmatched &n& as-is; once continuation lines bind, a still-unbound ref should be reported (as an unresolved dynamic call), never silently dropped.

    Resolution (2026-08-07). Lookahead added, bounded by the member's own highest &n&, so it stops as soon as the parameters are satisfied and cannot swallow a following statement's continuation; consumed lines are dropped from the emitted source. CALLNAT_DYNAMIC now accepts &n&, so a still-unbound parameter surfaces in /dynamic-calls/unresolved instead of vanishing.

    Correction to the scale figure above. The "263 edges" is the joint effect of this item with 121 and 123 — this item alone recovers exactly 0. Measured by simulating each fix over all 3587 upms modules before implementing: 121 alone +207 copycode-derived CALLNAT pairs, 120 alone +0, 120+121 +453, 120+121+123 +2672, 0 lost in every combination. (A cleaner A/B run after implementation — old parser from HEAD vs new, identical counting rule — measured 7319 → 9521, +2202, 0 lost. The pre-implementation figures used a looser counting rule and a smaller baseline; the honest headline number is +2202.) The example quoted above is itself a 123 case — its target is written '"YAGCHBN0"' — so fixing only what this item describes would have left the motivating symptom untouched. That is why 123 exists; see it below.

    Guard. CopycodePreprocessorTest (new — copycode expansion had no unit test at all, which is how two argument-parsing bugs survived: both are invisible unless you inspect the substituted text, and the existing IT only saw whether some edge came out) plus CopycodeExpansionIT#aBrowseIncludeWithMultiLineEscapedArgumentsResolvesItsCall end-to-end.

  • 121. Natural's doubled-quote escape shifts every positional copycode argument by one (found 2026-08-06 in the upms source-vs-API cross-check, done 2026-08-07)

    Symptom. '''YAGCHBN0''' is one Natural literal ('YAGCHBN0'). The argument pattern splits it into three — '', 'YAGCHBN0', '' — so &1& binds to an empty string and every later position is off by one. The CALLNAT &2& then names whatever &1& was: an ADABAS sort key.

    DAGCHEN0:999  INCLUDE YFRAMBC8 '''AGNT-CHG-CMP-SP''' '''YAGCHBN0'''
    GET /modules/DAGCHEN0/callees  → 'AGNT-CHG-CMP-SP' listed as a MODULE; YAGCHBN0 absent
    

    This also creates a phantom MODULE node named after the sort key (sourceFile: "", unresolved: true).

    Scale. 2308 INCLUDE lines across 260 modules use the escape. 37 of the 115 entries in /dynamic-calls/unresolved are sort keys rather than genuine variables — i.e. a third of that list is this bug, not real dynamic dispatch.

    Mitigating. The phantom is at least flagged unresolved: true, and /dynamic-calls/overrides offers a manual correction path — so nothing here silently claims to be resolved.

    Root cause. CopycodePreprocessor.java:33 — ARG = '[^']*'|\S+. Needs '(?:''|[^'])*', plus unescaping '' → ' in dequote(). Shares a root cause with item 120; both are the copycode argument parser and are best fixed and tested together.

    Resolution (2026-08-07). Exactly as diagnosed. In isolation this recovers +207 pairs (see the measurement table under item 120). The phantom MODULE nodes named after sort keys are gone, which also removes 37 of the 115 entries from /dynamic-calls/unresolved — that list is now dynamic dispatch rather than a third parser artefact.

  • 123. CALLNAT "X" — the double-quote string delimiter was never accepted, hiding 2219 calls (found 2026-08-07 while validating items 120/121, done 2026-08-07)

    Symptom. Natural delimits a string with ' or ". The CALLNAT patterns in both ingest tiers matched only ', so a call whose target arrives as a double-quoted literal was not a call at all — and, unlike a dynamic call, it was not reported as unresolved either.

    How it was found — and why it matters procedurally. It was not found by reading code. Before implementing 120/121 the fixes were simulated over all 3587 upms modules, which showed item 120 contributing a delta of exactly 0. Item 120's own headline example passes its target as '"YAGCHBN0"' — the double-quote-inside-single-quote idiom, 7232 uses in upms against 2521 for '''X'''. Without that check, both items would have been implemented, a long deep refresh run, and the motivating symptom found still broken afterwards.

    Scale. With 120 and 121 in place, accepting " is what turns a +453 recovery into the full one: measured A/B over all 3589 upms modules, 7319 → 9521 copycode-derived CALLNAT pairs (+2202, +30%), 0 lost. Independently confirmed as a genuine Natural delimiter three ways: 3817 direct = "SRO"-style literals in hand-written modules; the browse idiom's 7232 uses; and NaturalLines.java:55, whose existing javadoc already documented both the " delimiter and the doubled-delimiter escape.

    The real lesson. Everything needed to write these patterns correctly was already in the repo, in two places, and both were bypassed. NaturalLines.java:55 documented the delimiter and the escape. natural-grammar.md §21.3 gives the correct EBNF — both delimiters, doubled-delimiter escape, repetition count. What misled was natural-grammar.md:118, a one-line simplification (character-string = "'" { any-character } "'") carrying only a quiet "refined in Section 21.3" pointer. ARG and CALLNAT were written from the summary line, not the refinement. The durable fix is therefore not new grammar text but making the summary line impossible to use by accident: it now names both delimiters and the escape inline and says to read §21.3 first.

    This is worth remembering as a research failure rather than a coding one: the authoritative answer existed, was correct, and was one cross-reference away.

    Fix. One shared NaturalLines.CALLNAT_LITERAL + literalTarget(Matcher), used by NaturalCoarseScanner and NaturalParser, so the tiers cannot drift — the same consolidation the #61/#62/#63 family got. Guarded by aCallnatTargetMayUseEitherStringDelimiter in both tiers' tests and by aCallnatInsideAStringLiteralIsStillNotACall, which keeps the widened delimiter from re-opening bug #63.

  • 124. A full deep refresh does not reap placeholder MODULE nodes the parser no longer produces — fixed bugs keep answering from the graph (found 2026-08-09 while verifying 120/121/123, done 2026-08-09)

    Symptom. After the deep refresh that shipped items 120/121/123, DAGCHEN0/callees still lists the item-121 phantom AGNT-CHG-CMP-SP as a MODULE — alongside the now-correct YAGCHBN0, from the same call site.

    GET /modules/DAGCHEN0/callees
      YAGCHBN0         CALLNAT          YFRAMBC8.cpy:27  includedAt 999   ← correct, new
      AGNT-CHG-CMP-SP  CALLNAT_DYNAMIC  YFRAMBC8.cpy:27  includedAt 999   ← stale, unresolved:true
    

    Proof it is stale, not re-created. All 5 include sites feeding that copycode are byte-identical (INCLUDE YFRAMBC8 '''AGNT-CHG-CMP-SP''' '''YAGCHBN0''' + continuation), and running the current CopycodePreprocessor over the real DAGCHEN0.nat emits CALLNAT 'YAGCHBN0' YAGCHKEY and zero lines mentioning AGNT-CHG-CMP-SP. The parser cannot produce this edge any more; a full refresh (6311 files persisted) left it standing.

    Scale. 29 unresolved placeholder MODULE nodes carry 511 CALLS edges whose originFile is a .cpy. Most are legitimate — NDBERR/NDBNOERR (427 edges) and USR1009N/USR1023N are real externals absent from the corpus. The bug's own residue is the sort-key family: 19 names, 45 edges, including one whose name still carries its quotes ('YCOTBMN0'). Small, but it is exactly the wrong 45: they are the artefacts of a bug that is now fixed, and they outlive the fix.

    Why it matters beyond the count. This is a meta-defect: it means fixing a parser bug does not fully take effect until someone notices the residue and rebuilds from scratch. Every parser fix from here on inherits it, and the residue is indistinguishable from a genuine unresolved call, so nothing flags it. Note the reap must not be naive — &2&/&3& placeholders (11 edges) are the intended item-120 output and must survive.

    Where to look. The finalize pass has 48 steps and several reap steps (delete-resolved-field- contains, the item-62 data-literal cleanup); none of them appears to drop a placeholder MODULE that no longer has any producing call site in the re-parsed source.

    Resolution (2026-08-09). The gap was structural and had a precedent: GraphRepository already reaps a re-parsed Natural file's READS/WRITES to access placeholders (item 86) and its USING edges (item 106) before merging the fresh ones. Nobody had done it for CALLS. Item 86's own javadoc states the general cause — "the target placeholder is never swept and the source node survives, so neither DELETE_STALE_FILE_NODES nor the edge MERGE ever reaped it".

    Two parts: DELETE_STALE_NATURAL_CALL_EDGES (per-file, gated on reconcile) and DELETE_ORPHANED_PLACEHOLDER_MODULES (finalize, the MODULE counterpart of item 88's table sweep).

    Two boundaries the tests forced, both found by running rather than reasoning:

    1. Duplicate markers must survive the node sweep. Item 114 records a duplicate identity as a placeholder, and an unreferenced one is a degree-0 placeholder — exactly the shape the sweep targets. MARK_DUPLICATE_IDENTITIES runs before finalize, so the sweep deleted the marker just written and the endpoints fell back to 404 instead of 409 DUPLICATE_IDENTITY. Guarded on duplicatePaths IS NULL. Caught by DuplicateIdentityIT (3 failures).
    2. Resolver-built dynamic edges must not be reaped. The proposal said "reap all of the file's CALLS". That is wrong: 486 CALLNAT_DYNAMIC edges point at real modules and 392 of them carry no marker at all (no folded, no resolvedBy), so nothing but the target's file distinguishes enrichment output from parser output. Reaping them broke the item-37a path-warm — a scoped finalize does not reliably re-resolve an edge whose far side is outside its scope, so the deletion was permanent and a dataflow trace stopped at the dispatch boundary. The predicate is therefore t.sourceFile = "" OR r.callKind <> 'CALLNAT_DYNAMIC': everything the parser re-emits is reaped, the resolvers' own output is not. Caught by AnalysisResourceIT#flowForwardPathWarmCrossesIntoDynamicallyDispatchedCallee.

    Scope limit (deliberate). Keyed on the source node's file, so it misses the 605 call edges whose source subroutine is defined inside a copycode — 14 of them stale. Those nodes are MERGEd per (type, name, sourceFile) and are therefore shared by every including module, so reaping them during one module's refresh would delete edges other modules contributed and never re-create them. That needs a per-module identity for copycode-resident nodes and is a separate item. The same hole exists in items 86 and 106, which use the same key.

    Guard. StaleCallEdgeReapIT — three cases, two sabotage-verified. The fixture changes only the CALLNAT target and keeps the enclosing subroutine, which is precisely why the bug hid behind a green suite: DerivedCallsModuleRefreshIT deletes the whole subroutine, so there the FUNCTION node disappears and DETACH DELETE takes the edge along.

Agent API / MCP tooling gaps

(Done items 52, 53, 54, 56 moved to x-docs/features.md.)

  • 122. dispatch-table rows carry a bare lineNo with no provenance — copycode-derived rows point into the wrong file (found 2026-08-06 in the upms source-vs-API cross-check, done 2026-08-07)

    Symptom. callees, db-accesses, workfile-accesses and functions all carry sourceFile / viaCopycode / includedAt / includePath. dispatch-table carries none of it — just lineNo — so a row spliced in from a copycode reports the copycode's local line number as if it were a line of the host module.

    GET /modules/VCOMIN50/dispatch-table  → 26 of 44 rows report lineNo 18 or 20
    VCOMIN50:18   * #05 20.10.2011 SAGAPI  Erweiterung der Felder ...   ← a change-history comment
    real sites:   ISICINDE.cpy:18  and  ISICINDI.cpy:20  (included 13× each)
    

    An agent that follows the line number lands on a comment in the module header and finds nothing — quietly, with no signal that the pointer was resolved against the wrong file.

    Fix. Reuse the existing site/provenance shape rather than inventing a second one; the preprocessor already tracks LineOrigin for every expanded line, so the data is present and only needs carrying through to the response DTO. Cheap, and it makes the endpoint's rows checkable.

    Unrelated to item 108 in cause (that one is about which dispatcher idioms are recognised at all), but both are dispatch-table and worth touching in one pass.

    Resolution (2026-08-07). DispatchEntry gains the same quartet the other site-bearing endpoints already carry — sourceFile, viaCopycode, includedAt, includePath — with DISPATCH_TABLE using the established coalesce(w.originFile, m.sourceFile) idiom, so a host-local row is unchanged and a copycode-derived one names the .cpy. Additive: every existing field keeps its meaning. MigrationDossier.tsx gains a "from" column and a file-aware line link, so clicking a copycode row opens the copycode rather than the wrong line of the host. ac-cli needed no change — DispatchTableCommand prints the body verbatim (verified, not assumed). Guarded by three cases in NestedDispatchGuardIT, including one asserting the guard chain spans the INCLUDE boundary.

  • 109. variables/{name}/writes gives the location but not the written value, so resolving a dispatch needs the source anyway (found 2026-08-02, upms webservice-layer audit; done 2026-08-06)

    The data was already in the graph. 452 553 of 466 843 Natural WRITES edges (97%) carry r.value; the query behind the endpoint simply never selected it. So the fix is one line of Cypher plus three DTO fields — the audit's workaround (open the file at the six line numbers the endpoint had just returned) was never necessary, only unreachable.

    Two corrections to this item's own fix text. It asked for the value "when the right-hand side is a literal, null otherwise". The property does not work that way: it holds the right-hand side as written — 'YVLOGBN0' (literal), *PROGRAM (system variable), #DISPLAY(1) (indexed), #SELECTED-KEY.NUM-CIS-GC (qualified reference). Filtering to literals would discard the majority, so it is returned raw, exactly as dispatch-table already reports the same field. And the indexed case is not an array index but a substring window (item 83), surfaced honestly as assignedSubstrPos/assignedSubstrLen rather than under an invented name.

    Java carries no value at all — 1125 WRITES edges in pur, 0 with one, because the Java parser never captures the right-hand side. assignedValue is therefore always null there, and a bare null reads as "nothing is assigned" rather than "not captured for this language". Documented on the DTO and in the usage guide; the alternative would have been a new silent falsehood of exactly the kind items 107/114 exist to remove.

    Symptom. Working around item 108:

    GET /variables/%23WT-OBJ-PROG/writes?module=WPOLIX0S
    → [{"function":"INIT-OBJECT-TABLE","sourceFile":"…/WPOLIX0S.nat","lineNo":772, …}, … 6 rows]
    

    Six correct write sites, and not one of the six assigned values. Answering "what does this dispatcher dispatch to" therefore requires opening the file and reading lines 772/775/777/780/783/785 — the API narrows the search to the right lines and then stops one step short. dispatch-table already returns assignedValue for the DECIDE idiom, so the concept and the field name exist.

    Fix. Add assignedValue (and, where the write is indexed, assignedIndex) to the writes rows, populated when the right-hand side is a literal, null otherwise. Cheap next to item 108 and useful far beyond it: "which constants does this module put into field X" is a routine question in a reengineering pass.

  • 110. No reachability query — "can A reach B?" has to be hand-rolled as ~100 callers calls (found 2026-08-02, upms webservice-layer audit; done 2026-08-06)

    GET /modules/{name}/reaches?target=A,B,C&direction=up|down&depth=N → {reachable, paths, truncated}, plus ac reaches. The audit's question — does any W* module reach the commission calculation — now runs as one bounded query:

    MATCH p = shortestPath((a)-[:CALLS_MODULE*1..6]->(b))   → 25 Pfade, 1,7 s
    

    Previously ~100 HTTP round-trips and a hand-written path reconstruction.

    The edge type is a correctness decision, not an optimization. The obvious -[:CALLS*1..n]-> is wrong: Cypher cannot constrain the intermediate nodes of a variable-length pattern, so the path would route through FUNCTION nodes and report module reachability where there is none. The materialized module-to-module CALLS_MODULE (item 68) is module-level by construction.

    The guard differs by direction, which is item 107 applied one level deeper: downward the answer comes from this module's own calls, so an un-ingested placeholder must not answer "not reachable" — that would be empty for want of data. Upward the routes are made of the callers' source and are genuine even when the target itself was never parsed.

    What reachable: false does not mean. CALLS_MODULE is built only over resolved calls, so a route through an unresolved dynamic CALLNAT (item 82) is invisible. The answer is "no path over known edges" — stated on the DTO, because the unqualified reading is exactly the false-negative this codebase keeps having to remove.

    Bounded on purpose (item 75: 22 self-loops, 162 two-cycles in one project); ReachabilityIT covers the witness path, the absent path, both depth sides of the bound, several targets in one request, the upward direction, a cycle, and the empty-target refusal.

    Symptom. The question was "does any W* module reach the commission calculation (ISINCOMI / VCOMIN00 / VVERAN50 / VCOMIN55 / VCOMIN57 / VCOMIN50)?" — a yes/no with a witness path. There is no endpoint for it. call-tree goes downward from one root and returns a flat closure without paths, so it answers "what does A reach", never "who reaches B", and never "how". The workaround was a client-side breadth-first search upward over /callers, six seeds, depth 6: ~100 HTTP round-trips, 101 modules visited, and the path reconstruction written by hand.

    Fix (proposal). GET /modules/{name}/reaches?target=<name>&direction=up|down&depth=N returning {reachable: bool, paths: [[module, …], …], truncated: bool} — or, more useful for this shape of question, a filtered variant of callers/call-tree that accepts a set of targets and returns only the witnesses. In Cypher this is one bounded shortestPath/variable-length match; done client-side it is 100 requests and an easy place to introduce a bug. Note the traversal must be bounded — see item 75 on the CONTAINS cycles.

    Why it matters. "Who can trigger X" is the recurring question in legacy reengineering: which entry points reach a calculation, a table write, an external interface. It is the natural counterpart to call-tree and currently the biggest hole in the query surface for that work.

  • 82. Manual override for unresolvable dynamic CALLNAT targets (human/agent-settable) (proposed + implemented 2026-07-19, from the WGEAGB0S deep-API audit; REST + MCP + ac CLI + Testcontainers ITs green). The dynamic-CALLNAT resolvers cannot follow every name-assembly pattern — e.g. YGEAGGNH.nat:443 CALLNAT #GETSHORT-MODUL where the name is built via MOVE 'YGEAGKEY' TO #GETSHORT-MODUL + MOVE 'GN0' TO SUBSTR(#GETSHORT-MODUL,6,3) → YGEAGGN0 (a real, ingested module). Such a call site leaves an unresolved placeholder (type=MODULE, sourceFile="", name = the variable). Add a REST + MCP + ac-CLI capability to list unresolved dynamic call sites and manually resolve a call site to one or more target modules (multiple targets = deliberate branches, each materialised as a real CALLS callKind=CALLNAT_DYNAMIC edge, provenance manual). Decisions (2026-07-19): callsite key = originFile + lineNo; overrides are persistent (own node type the refresh never deletes) and auto re-applied by an enrichment step after every refresh/deep-refresh; a reset endpoint clears manual overrides (one call site, or all). Related: this is the actionable counterpart to the Bug B consistency gap — callees/digest should also surface the unresolved flag that graph already exposes.

  • 83. Auto-resolver for string-assembled dynamic CALLNAT targets (SUBSTR/MOVE constant-folding) (proposed 2026-07-19, from the WGEAGB0S deep-API audit follow-up; implemented 2026-07-27). Many unresolved dynamic call sites are in fact statically foldable: the target name is built from literals only, e.g. the Y…GNH "GetShort" family (~44 modules) — MOVE 'YxxxxKEY' TO #GETSHORT-MODUL + MOVE 'GN0' TO SUBSTR(#GETSHORT-MODUL,6,3) → YxxxxGN0 — plus similar families (PXFRAA01.ST-PGM across the P…MP0 programs, the YFRAMBCK.cpy:27 group). Add an enricher that constant-folds a chain of MOVE <literal> and MOVE <literal> TO SUBSTR(var,pos,len) assignments feeding a CALLNAT var into the effective target, and resolves the edge automatically when that target is a real ingested module — so item 82's manual override is only needed for genuinely runtime-dependent names, not the deterministic string idioms. Must respect precedence: a manual override (item 82) still wins over an auto-fold. Deliver with characterization ITs (a minimal fixture per idiom) + the usual REST/MCP/CLI-visible effect (fewer unresolved sites; resolved callees/callers). Done: NaturalParser records MOVE '<lit>' TO SUBSTR(var,pos,len) as a WRITES on the base var carrying substrPos/substrLen; new RESOLVE_DYNAMIC_CALLNAT_FOLD (+_SCOPED) folds base literal + ordered overlays (reduce/left/substring) → resolved CALLS edge tagged folded=true; the direct-literal/indirect/cross resolvers now skip partial-slice writes (substrPos IS NULL). Precedence honoured by guarding the fold against :DynamicCallOverride sites plus a delete-folded-overridden-dynamic-callnat step before apply-manual. Characterization ITs in DynamicCallnatFoldIT (fold resolves YABALKEY+GN0@6/3 → YABALGN0; manual override wins); full dynamic-callnat regression 89/89 green. Docs in agent-api-usage-ac-implementation.md.

  • 26. MCP session reliability — RESOLVED BY REMOVAL (2026-08-04) (investigated 2026-07-07, reproduced 2026-08-02, never fixed). mcp__agenticcode__* calls intermittently — and in the 2026-08-02 session, from the very first call — failed with "the first message from the client must be initialize: tools/call", forcing every playbook step to be re-expressed as curl. The evidence pointed at the MCP client's reconnect handling rather than a server-side bug this codebase's config could fix, and the REST endpoints answered normally throughout. Decision 2026-08-04: the MCP server surface was removed entirely rather than debugged — see "MCP surface removed" in x-docs/features.md. REST + ac CLI are now the only access paths; all 40 former tools had a REST twin, so no capability was lost.

Comments as graph data (item 141) — 2026-08-28

  • 141. Comments and commented-out code are not indexed — so the documented way to find a Java counterpart returns a confident false negative

    Symptom, part one: the counterpart lookup. The consuming project's convention (its CLAUDE.md §3) is that a reengineered web service carries its Natural origin in a Javadoc block, and that the way to go Natural→Java is to search the program name in pur, where "an empty result means not yet reengineered". Measured:

    GET /pur/search/value?value=JX0034N0&contains=true   → 1 hit, MultiTableImportJob.java:33
    GET /pur/search/value?value=WPARTX0S&contains=true   → []
    GET /pur/search/value?value=WEXKYX0S&contains=true   → []
    

    The batch job is found because PROGRAM_IDENTIFIER = "JX0034N0" is a string literal. WPARTX0S and WEXKYX0S are fully reengineered — PartnerController and its logic classes have existed for months — but their origin is recorded only in a Javadoc block:

    /**
     * ServiceEndpoint: partner.update.partnercs.Update
     * UPMSFunction: com.uniqagroup.upms.partner.svc.esp.UpmsPartnerCsUpdate
     * UpmsObject: PartnerCs, UpmsAdapter: Update
     */
    

    Comments are not nodes, so the search cannot see it. The API therefore answers "not yet reengineered" — the exact wording the consuming project derives from an empty result — for every web service that has already been reengineered. That is not a missing answer; it is the wrong one, delivered with no signal, on the single question every reengineering job opens with. The only reason it has not yet caused a duplicate implementation is that the agent happened to distrust it.

    Symptom, part two: the semantics of upms live in the comments. Natural source in this corpus carries its change history, its business caveats and its disabled logic as comment text:

    GET /upms/search/value?value=Bug%20266&contains=true → []
    

    while WAGNTX0S.nat:21 reads * #01 09.05.07 VOVBJ03 Bug 266, and lines 22-23 record two regenerations with their ticket numbers. The #01…#05 change markers correlate to --> #04 / <-- #04 blocks that delimit which statements a given change introduced — often the only record of why a branch exists. None of it is queryable. Every such question falls back to reading the file, which in the consuming project is explicitly the exception path and, for the Java side, blocked by a hook.

    Proposed shape. A COMMENT node per contiguous comment block, carrying sourceFile, startLine, endLine, text, and an edge to the nearest following declaration (module, function, field) — so /modules/{name}/comments answers "what does this module's header say" and search/value reaches comment text like any other content. Three deliberate points:

    • search/value must keep the two kinds separable. A comment hit and a code hit are not the same evidence. Suggest kind: "COMMENT" on the existing row shape (it already discriminates ASSIGNMENT/NODE) rather than a fourth search endpoint, plus a switch for callers that want today's behaviour. Folding comments into the default result set silently would move every existing completeness count — item 131's lesson.
    • Cost must be measured before it is defaulted on. Natural in this corpus is comment-dense (WAGNTX0S.nat is ~25% comment lines in its header alone), and item 128 already showed a 3-5× deep-refresh cost for one new edge family. Measure persist time and store growth on upms before deciding whether this belongs in Tier 1 or in the deep tier only.
    • Javadoc is structured, plain comments are not. The ServiceEndpoint: / UPMSFunction: / UpmsObject: block above is a key-value list. Parsing it into properties is item 143's business; 141 only has to make the text reachable, and 143 should not be blocked waiting for it.

    Implemented 2026-08-28. Comment blocks are graph nodes, opt-in everywhere.

    • NodeType.COMMENT + EdgeType.DOCUMENTS (ac-parser-core). One node per contiguous block, text in value, properties commentKind / lineCount / truncated (text cut at 4 000 chars). Exactly one DOCUMENTS edge per block, to the declaration immediately below it, else the one enclosing it, else the module — so a header banner documents the MODULE and a Natural /* on a field's own line documents that field. CommentBlocks/CommentProperties hold the shared rule so the two parsers cannot drift.
    • Deliberately not CONTAINS. That edge is walked by the functions listing, the SEARCH_BY_VALUE assignment arm, the ego-graph and the stale sweep; comments hung off it would leak into queries that never asked for them (item 128's MENTIONS vs REFERENCES reasoning).
    • The name carries no prose. A comment node is named comment@<startLine>; SEARCH_IDENTIFIER_CORE matches every node's name with no type filter, so text there would have turned every contains=true identifier lookup into a full-text search. It additionally excludes type = 'COMMENT' outright.
    • search/value?includeComments=true (CLI --include-comments) adds comment hits, labelled kind: "COMMENT". Default output is byte-identical to before — no existing completeness count moves.
    • GET /modules/{name}/comments?kind= (CLI ac comments) lists a module's blocks with their target declaration. **SAG directives are a separate kind and excluded unless asked for: they are generator metadata, and extractDescription already mines them.
    • Deep-gated. Comments come from the full parse, not the Tier-1 coarse scan, so the endpoint deep-ingests on demand and answers 409 NOT_DEEPLY_INGESTED rather than [] — an empty list that means "not analysed" is the very failure this item was filed about.
    • Natural specifics. Full-line * runs group into one block; a trailing /* comment is its own single-line block and is detected quote-aware, so MOVE 'A/*B' TO #X is not a comment (the older NaturalLines.stripInlineComment is still not quote-aware — a separate, pre-existing precision bug in the strip path, filed as item 152). Comments are read from the module's own file, never from copycode-expanded lines: a .cpy's comments belong to the copycode's own module, once, with correct line numbers (items 75-B/75-C).
    • Tests: CommentIndexIT (both languages, the opt-in guards, the SAG exclusion), plus parser unit tests for block grouping, adjacency-based targeting and the quote-aware inline scan.

    Measured cost (2026-08-28, deep refresh of upms + pur on the real corpus). The item asked for this before defaulting the feature on, and the answer is: comments are not cheap.

    before after delta
    upms nodes 508 670 938 561 +429 886 (+84 %), of which every one is a COMMENT
    upms deep refresh 977 s 1 727 s +77 %
    pur nodes 60 690 (call-graph depth) 83 801 (deep) +23 102 comment nodes
    Neo4j store (all projects) 2.0 G 2.2 G +10 %

    By kind: upms 242 454 NATURAL_INLINE / 136 480 NATURAL_BANNER / 50 952 SAG (17.6 M chars total); pur 15 421 LINE / 7 473 JAVADOC / 208 BLOCK (2.8 M chars). 452 989 DOCUMENTS edges, zero orphans, zero endLine < startLine, zero empty texts, 157 blocks truncated at the 4 000-char cap.

    Consequences, stated rather than buried:

    • Comments are 46 % of all upms nodes. Every unindexed substring scan (search/identifier, search/value) walks them even when it then filters them out — the MATCH is over :AstNode, the exclusion is a predicate. Measured on upms: a contains name scan is 2.8 s in Cypher (2.5 s with the exclusion), ~10 s through the paged endpoint. An index or a separate label for comment nodes would fix the scan, and is the obvious follow-up if this becomes the complaint.
    • The deep tier only — the Tier-1 coarse scan emits no comments, so a plain refresh pays none of this. That is why /modules/{name}/comments is deep-gated rather than answering [].
    • NATURAL_INLINE alone is 242 k nodes (56 % of upms's comment nodes) for 5 M chars — the per-field /* description idiom. If the cost has to come down, dropping or coalescing those is the first lever, and the one that loses the least: field descriptions are also reachable via the field's own row.

    Verified against the real corpus — the three probes the item was filed with:

    GET /pur/search/value?value=WPARTX0S&contains=true                      -> 0   (code only, unchanged)
    GET /pur/search/value?value=WPARTX0S&contains=true&includeComments=true -> 3   (PartnerController's
                                                                                   two origin Javadocs
                                                                                   + PartnerCsUpdateLogic)
    GET /upms/search/value?value=Bug%20266&…&includeComments=true           -> 6
    GET /pur/search/value?value=WEXKYX0S&…&includeComments=true             -> 8
    GET /pur/search/identifier?name=UpmsObject&contains=true                -> 2   (two real Java fields,
                                                                                   no comment leakage)
    GET /upms/modules/WAGNTX0S/comments                                     -> the header banner, the
                                                                              `#01 … Bug 266` change log,
                                                                              and per-field inline notes
    

Persist phase: instrumentation and the placeholder sweep — 2026-09-05

  • 153. The persist phase was a single opaque number (2026-09-05)

    The finalize phase has always logged per step (runEnrichment: duration, nodes and edges created/deleted). Persist had nothing of the kind — a deep refresh of upms spent 724 s of 1 256 s there, spread over 32 batch lines without any breakdown. An extrapolation from the raw MERGE rate explained only ~234 s of it; ~460 s were unattributable.

    GraphRepository.mergeResults runs 5-9 statements per batch in one transaction (node merges, up to 17 edge merges per EdgeType, five stale sweeps), none of which were consumed — so neither timings nor SummaryCounters were available.

    Built: PersistStats measures every statement (via .consume(), which also yields the counters) plus the Java-side preparation, and logs one aggregated line per batch:

    Persist deep 200 files: total 8740 ms = prepare 108, merge-nodes 1630, merge-positional-nodes 506,
      reap-table-access-edges 310, reap-using-edges 161, reap-call-edges 175, merge-edges 1479 (6 types),
      sweep-stale-file-nodes 132, sweep-stale-placeholders 1887, sweep-resolved-field-edges 278, commit 20
    

    Deliberately one line rather than one per statement: per statement it would be ~800 lines per refresh, burying the 51 finalize lines. The residual commit = total − sum of labels is intentional — an unnamed residual is exactly what this instrumentation is meant to eliminate. The single-file path logs at DEBUG, because a fan-out warm runs through it hundreds of times.

    Result of the first measurement (upms, deep, 2026-09-05): parsing 18 s, persist 639 s, finalize 567 s. sweep-stale-placeholders 344.9 s = 54 % of the persist phase and 28 % of the whole run. The Java-side preparation, which I had suspected as a possible hidden cost block: 1.5 s.

  • 154. DELETE_STALE_PLACEHOLDER_NODES ran once per owner instead of once per batch (2026-09-05, found via item 153)

    The sweep (item 76) did UNWIND $owners AS o MATCH (n {project, sourceFile: "", ownerModule: o}). No index covers ownerModule, so every seek went through ast_node_project_sourcefile — and since all placeholders share sourceFile = "", every seek returned all 14 551 placeholders of the project and discarded all but a handful. At ~200 owners per batch x 32 batches that is ~6 400 full passes through the same bucket.

    Fix: one pass, filtered against the list (WHERE n.ownerModule IN $owners). Semantically identical (same node set; IN deduplicates repeated owners, which is inconsequential for a DELETE), one line of Cypher, no new index.

    before after
    sweep-stale-placeholders 344.9 s 3.6 s (−99 %)
    persist phase total 639 s 332 s (−48 %)
    deep refresh upms total 1 225 s 956 s (−22 %)

    Two wrong turns the measurement saved — both plausible, both false:

    1. For resolve-bare-included the plan looked like "expand first, filter later". The index-first rebuild (seek on (project, name), then check containment) was slower and was still running after 10 minutes — field names are not selective enough project-wide.
    2. For this sweep the obvious suggestion was an index on (project, ownerModule). The reformulation counter-measured during validation was 3x faster without an index — and thus without a write surcharge on each of the 940 k node merges. An index remains open as the next lever, but has to prove itself against this shape.

    The success prediction was wrong too, though in the favourable direction: ~100 s of savings were estimated (assumption: the DETACH DELETE is irreducible), and it became ~341 s. The cost was almost entirely in the repeated scanning, not in the deleting.

    Next lever (measured, open): the finalize phase is now the larger item at 607 s, of which resolve-bare-included READS+WRITES is 331 s and link-args-to-params 105 s (0 edges created). Together 436 s = 72 % of the finalize phase.

  • 155. Enrichment steps can be profiled (refresh?profile=true) (2026-09-05)

    The persist instrumentation (item 153) answered "which statement", not "on what within a statement". For the five expensive field steps it therefore remained open whether their time sits in matching or in writing — and a read-only reconstruction cannot answer it, because after the refresh it only sees the leftovers (the fallacy "the step finds nothing" lay ready exactly there).

    runEnrichment optionally prefixes each step with PROFILE and logs, for every step from 5 s upwards, the five operators with the most db-hits (8 of the 51 steps rather than all). Opt-in per run via POST /refresh?profile=true and ac refresh --profile, both explicitly marked DIAGNOSTIC and listed in the agent guide only in the performance section, not in the endpoint table: a profiling switch is a tool for developers, not for agents.

    Measured overhead: none (908 s profiled against 905 s unprofiled). My warning about 10-30 % was not borne out — the numbers are directly comparable.

    Result (upms, deep, 2026-09-05). Every expensive step shows the same pattern: expand wide, then throw almost everything away. Not a single write operator appears in the top 5.

    step time dominant operator ratio
    resolve-bare-included READS/WRITES 2x ~165 s Filter 273 320 729 hits -> 43 753 rows 1 : 6 246
    VarLengthExpand(All) 127 180 800 hits -> 63 017 519 rows
    link-args-to-params 105 s NodeIndexSeek 65 711 041 hits -> 64 653 592 rows, filtered to 18 371 1 : 3 518
    resolve-field-placeholder READS/WRITES 2x ~50 s Filter 50-52 M hits -> ~220 000 rows 1 : 229
    resolve-view-alias-* (3 steps) 3x ~12 s Expand(All) 37-43 M hits

    For link-args-to-params the individual seek is not the problem — it yields 11 estimated rows. It is merely executed millions of times, because the query forms a cross product over the include files of both sides times the argument list (UNWIND cvfiles x UNWIND pvfiles). An index therefore does not help there; the candidate set has to get smaller.

  • 157. link-args-to-params looked the parameter up once per caller variable instead of once per argument (2026-09-05, found via item 155)

    The step cost 105 s and created zero edges on a re-refresh; not a single write operator appeared among the five most expensive in the profile. The fan-out measurement explained it:

    call sites with arguments 27 240
    argument slots 91 954
    seeks caller side (cv) 2 509 135
    seeks parameter side (pv) 28 868 877

    The parameter at position i depends only on (callee, i), but was looked up inside the caller loop — 28.9 M seeks for 18 371 result pairs.

    Fix: parameter lookup moved into an aggregating CALL subquery per (callee, i), caller side afterwards. In both variants (LINK_ARGS_TO_PARAMS and ..._SCOPED, the latter in the interactive deep-ingest path).

    before after
    step link-args-to-params 105 s 66.8 s (-36 %)
    deep refresh upms total 905 s 831 s

    The read-only advance measurement had predicted 67.2 s — the implementation hit it to within 0.4 s. Of the 74 s total saving, however, only 38 s are attributable to the step; the rest lies in the run-to-run variance of ~10 % that was observable throughout this measurement series.

    Two details that carry the reformulation:

    • The subquery aggregates (RETURN collect(pv)). It therefore always returns exactly one row, and an empty list when the parameter is missing, which the guard size(pvs) > 0 discards. A non-aggregating subquery would have swallowed the row — the same result today, but for the wrong reason, and wrong at the next change.
    • The scoped variant dropped cm after the first WITH. A copy-paste of the whole-root version would have been unable to build the caller file list there — in the interactive path, where a silent null result is hard to notice.

    Not done, with reasons: deduplicating the caller side as well. 91 954 argument slots stand against 50 610 distinct (file, argument name) pairs — a factor of 1.8 for a considerably more complicated query with a re-join. The floor of today's structure is 56.5 s (caller side measured on its own), and the rebuild at 66.8 s is ten seconds away from it.

    Balance of the performance work (items 153-157): deep refresh upms 1 225 s -> 831 s (-32 %), achieved with two Cypher reformulations. Remaining distribution: persist ~310 s, finalize ~520 s, of which resolve-bare-included READS+WRITES alone is 298 s.

  • 158. resolve-bare-included expanded every data area once per placeholder (2026-09-05, found via item 155)

    The most expensive enrichment step (147 s + 151 s for READS/WRITES) ran per (module, placeholder) and expanded the complete subtree of every included data area down to depth 10. The subtree of a data area was thus traversed again once per placeholder of the including module. The profile showed 273 M db-hits for 43 753 result rows and 63 M intermediate rows from the VarLengthExpand.

    subtree expansions before 380 639
    after (once per (module, include)) 41 377

    Fix: the query is driven by the INCLUDE rather than by the placeholder — expand once per (module, include) first, then match the placeholders onto it by name. In both variants (whole-root and scoped).

    before after
    resolve-bare-included READS 147.4 s 75.8 s (-49 %)
    resolve-bare-included WRITES 150.6 s 76.8 s (-49 %)
    deep refresh upms total 831 s 695 s

    Edges created identical (+23 788 / +68 936 as before), graph unchanged at 940 592 nodes and 2 186 496 edges — the equality evidence on the real corpus, in addition to 411 green ITs (among them GroupQualifiedLeafResolveIT for the qualified path from item 81).

    Why it is twice as fast, structurally and not by accident: ph is now bound with all four properties of the index (project, sourceFile, type, name). A composite index only applies when every property is bound — the old ordering had no name at that point and could therefore never use the index.

    The price, knowingly paid: modules with includes but without placeholders now expand in vain — 1 359 of the 3 489 including modules in upms, ~11 618 additional expansions. Against 339 262 saved ones that is a good trade, and the share shrinks precisely when many placeholders are unresolved, i.e. when the step has real work to do.

    The advance measurement had estimated 80-120 s of savings; it became 145 s. As with item 154 the estimate was too low, because it treated the write share as irreducible.

    Balance of the performance work (items 153-158): deep refresh upms 1 225 s -> 695 s (-43 %), achieved with three Cypher reformulations without a single new index.

  • 159. resolve-field-placeholder expanded the field subtree once per reference (2026-09-05, found via item 155)

    After item 158 the largest remaining cost in phase D. The qualified field reference (CDPDA-M.SORT-KEY) was resolved reference-driven: for every one of the ~220 000 (ph, phv, src, m) rows the module's INCLUDES list was scanned for the structure name and the whole field subtree of real re-expanded to depth 10 — ~50 M Filter hits for ~220 000 result rows.

    Fix: the resolution (ph, phv) -> real -> realv runs first, over the project's 485 placeholder fields; the reference edges join on afterwards. real comes from the (project, type, name) index. The INCLUDES edge still decides which data area counts — it just no longer finds real, it checks it.

    | | | |---|---| | subtree expansions before | ~220 000 (one per row) | | after (once per (phv, real)) | 520 |

    | | before | after | |---|---|---| | resolve-field-placeholder READS | ~50 s | 18.2 s | | resolve-field-placeholder WRITES | ~50 s | 22.5 s | | deep refresh upms total | 695 s | 628 s |

    Two clauses are load-bearing, not cosmetic — both verified with EXPLAIN, because the first attempt without them had no effect at all: the WITH DISTINCT is a planner barrier (without it the planner reverts to the reference-driven order and the hoist is undone), and the INCLUDES test is written as EXISTS { (m)-[:INCLUDES]->(real) } rather than a MATCH (as a MATCH the planner sources m from real's includers and builds an m x src cartesian — measurably worse than the starting point).

    Equivalence evidence: 411 green ITs (among them QualifiedFieldResolveIT, QualifiedGroupTargetResolveIT, QualifiedWriteReconcileIT, PlaceholderResolveNullLineIT) and, more directly, the old query form finds 0 remaining rows for both READS and WRITES on the graph the new one produced: the new order resolves exactly the same set. The by-name fallback (steps 26/27) stayed at +1/-1 and +260/-220, so it picked up no leftovers either.

    The precondition that can tip: the order only pays while a placeholder's structure name stays selective enough that the index seek does not out-fan the include list it replaces — on upms at most 4 real structures share a placeholder name (mean 1.45). The result stays correct either way, since the INCLUDES test still filters; only the cost flips back.

    The _SCOPED variant is deliberately unchanged — it starts from $names and processes a handful of modules, where there is nothing to hoist. Whole-root and scoped therefore have different query shapes for the same result.

    Balance of the performance work (items 153-159): deep refresh upms 1 225 s -> 628 s (-49 %), achieved with four Cypher reformulations and no new index.

  • 160. The five per-file reap/sweep statements seeked once per (file, owner) pair (2026-09-05, found by reading the item-153 persist instrumentation)

    With finalize down to ~316 s, persist (290 s) became the larger half and was measured for the first time. One batch stood out: the last 111 files cost 45.9 s, of which 32.4 s (71 %) went into the five reap/sweep statements — for 1 123 new nodes and 7 467 edges. Batch 3, with 48 154 edges, spent 1.4 s on the same statements.

    Cause — the item-154 pattern again, one level up. All five ran as UNWIND $files AS p MATCH (n {project, sourceFile: p.f, ownerModule: p.o}), but the index is (project, sourceFile). A copycode node carries its expansion site as ownerModule, so one file holds many owners' nodes and every seek returned all of them:

    | | | |---|---| | copycode files (ownerModule <> "") | 264 | | (sourceFile, ownerModule) pairs they carry | 20 343 | | owners per file: mean / max | 77 / 1 993 (YFRAMEC2.cpy) | | nodes in those files | 53 740 | | rows the seeks requested to reach them | 23 255 522 (1 : 433) |

    Fix: $files carries {f, os: [ownerModule, ...]} — one entry per file — and the query filters n.ownerModule IN p.os. One seek per file instead of one per pair. No new index: the obvious alternative, (project, sourceFile, ownerModule), would tax the write side of all 940 k node merges, and merge-nodes is the largest persist item at 69.5 s.

    Measured read-only on YFRAMEC2.cpy with its full owner list before implementing: 7 944 098 db-hits pair-driven against 3 986 grouped, both matching the same 3 986 nodes. The objection this shape had to survive was whether IN re-scans the list per node — it does not, Neo4j hashes it.

    | | before | after | |---|---|---| | reap-table-access-edges | 17.9 s | 3.5 s | | sweep-resolved-field-edges | 17.5 s | 4.3 s | | sweep-stale-file-nodes | 16.4 s | 1.4 s | | reap-using-edges | 16.3 s | 2.3 s | | reap-call-edges | 16.2 s | 2.3 s | | the five together | 84.3 s | 13.8 s (-84 %) | | persist phase | 289.8 s | 217.9 s | | deep refresh upms total | 628 s | 552 s |

    The 45.9 s outlier batch is gone; the most expensive batch is now 20.6 s and is dominated by merge-nodes and merge-edges.

    Equivalence evidence: every finalize step that consumes what these statements leave behind reports byte-identical counts to the previous run — resolve-field-placeholder +167 536/-123 279 and +206 414/-168 671, resolve-bare-included +23 788/-23 788 and +68 936/-68 936, delete-resolved-field-contains -157 667, delete-resolved-placeholders -178 371. Had the reaps deleted too much or too little, these would move. Plus 411 green ITs.

    A test that asserted nothing was fixed on the way: StaleFileSweepIndexIT passed {sourceFile, ids} — keys the query has never read — so it checked the plan of a query whose parameters were meaningless. It now passes the real shape and covers all five statements instead of one.

    Estimate vs. outcome: predicted 45-70 s, measured 72 s in persist / 76 s end to end. Third time in a row the estimate came in low.

    Balance of the performance work (items 153-160): deep refresh upms 1 225 s -> 552 s (-55 %), achieved with five Cypher reformulations and no new index.

  • 162. The view-alias resolvers searched from the access, not from the alias (2026-09-06)

    Three finalize steps (resolve-view-alias-tables READS/WRITES and resolve-view-alias-access-nodes) cost ~32 s to make 355 edge changes — the worst ratio of any step. PROFILE showed why: the planner entered at the 52 762 INCLUDES edges and re-expanded each module's whole CONTAINS+READS subtree once per include.

    | operator | rows | db-hits | |---|---|---| | start at INCLUDES | 52 762 | 158 286 | | (m)-[:CONTAINS*0..1]->(f) | 3 979 711 | 3 979 711 | | (f)-[r:READS]->(alias) | 11 206 268 | 37 650 118 | | Filter alias.type='DB_TABLE' | 111 566 | 11 429 400 |

    ~53 M db-hits, of which the type filter discarded 99 %. The corpus holds only 100 alias declarations (100 distinct names, at most one each), so the fix is to start there and join the accesses on afterwards — the same inversion as items 158-160. The module scoping is untouched: the INCLUDES edge still decides which data area counts, it is now checked rather than searched.

    | | before | after | |---|---|---| | resolve-view-alias-tables READS | 10.9 s | 0.7 s | | resolve-view-alias-tables WRITES | 10.7 s | 0.9 s | | resolve-view-alias-access-nodes | 10.3 s | 0.8 s | | the three together | ~32 s | 2.4 s (-93 %) | | finalize | 298.8 s | 262.9 s | | deep refresh upms total | 517 s | 492 s |

    Equivalence evidence: all three steps report edge and property counts identical to the previous run — +45/-45 props 135, +117/-117 props 675, +193/-193 props 193 — plus 411 green ITs including NaturalCrossFileViewAliasIT and NaturalViewAliasDbAccessIT.

    Same two load-bearing clauses as item 159, verified with EXPLAIN for all three queries: WITH DISTINCT as a planner barrier, and the INCLUDES test as an EXISTS predicate rather than a MATCH.

    An honest limit on the pre-measurement: on the resolved graph the new form short-circuits (no alias DB_TABLE survives finalize), so the read-only 11.6 s vs 0.7 s comparison only proved that the old form pays full price even for zero hits. The real proof was the refresh.

    An inherited claim that could not be reproduced: the original comment justified module scoping with "NEXT-VIEW is declared over 100 different tables across the corpus". Measured on upms: 100 declarations over 100 distinct names, none repeated. The scoping was kept exactly as it was precisely because the justification could not be re-verified.

    Balance of the performance work (items 153-162): deep refresh upms 1 225 s -> 492 s (-60 %), achieved with six Cypher reformulations and no new index.

  • 163. resolve-bare-included searched from the include side, not from the placeholder (2026-09-06)

    The step was the largest single item left in finalize at 135.8 s, 52 % of it. Item 158 had already turned it once — from placeholder-driven to include-driven — which halved it. The remaining cost was structural and on the other side of the same trade-off:

    | measurement on upms | value | |---|---| | INCLUDES edges | 52 762 | | distinct data areas behind them | 2 583 | | subtree walks per data area | 20.4x | | rows out of (s)-[:CONTAINS*1..10]->(realv) | 3 264 895 (61 167 paths exist) | | rows surviving the realv filter, each costing a ph index seek | 2 485 384 | | placeholders on the other side | 21 080, 1 023 distinct names, only 635 with a real field |

    So the join runs over 635 names but was driven from the 2.49 M side. Turned around: realv is bound by the (project, name) index, and realv.sourceFile IN includedFiles prunes the name hits to 6 973 before any subtree walk.

    Correction (item 167, same day): this entry first said "prunes 944 746 name hits". That was the count after the file filter. The name seek actually returns 9 363 762 rows, ~8.4M of them other modules' placeholders (sourceFile ""), which the (project, name) index cannot exclude. The prune is a factor of 1 343, not 135 — and reading it as 944 746 is precisely why the file-driven seek of item 167 was not tried here straight away.

    | | before | after | |---|---|---| | resolve-bare-included READS | 67.0 s | 43.7 s | | resolve-bare-included WRITES | 68.8 s | 46.5 s | | the step together | 135.8 s | 90.2 s (-34 %) | | finalize | 262.9 s | 225.7 s | | deep refresh upms total | 492 s | 457 s |

    Equivalence evidence: edge counts identical to the previous run in both verification runs — +23 788/-23 788 and +68 936/-68 936 — plus BareFieldModuleScopeIT (3) and CopycodeExpansionIT (7) green.

    The prefilter is semantics-bearing, not tuning. It assumes a field lives in its data area's file. Verified in the sharp form rather than by its consequence: upms has zero cross-file CONTAINS edges out of a DATA_STRUCTURE, project-wide. Reachable (46 057 nodes) is a strict subset of same-file (47 362), so it is a correct prefilter and the exact EXISTS stays — a pure file swap would have over-matched by 1 305 nodes. If the assumption ever breaks, this step resolves less and reports no error, which is why the check is written down here.

    This approach was recorded as falsified and it was worth re-testing. Item 156 stated that seeking the candidate by (project, name) first was slower ("still running after 10 minutes, aborted"). That observation is correct and was reproduced today. What was too broad was the conclusion: the name alone is unselective, the name plus the module's include files is not. Two clauses make the difference, both verified with EXPLAIN — the WITH DISTINCT planner barrier and the file prefilter. Without the barrier Neo4j plans the filter behind the SemiApply and the old failure reappears exactly.

    Two measurement mistakes, recorded because they cost a wrong prediction. The read-only pre-measurement said 67 s -> 11 s; the truth is 67 s -> 43.7 s. On an already resolved graph the query short-circuits (6 973 surviving rows against ~92 700 during finalize), so it only ever proved that the old form pays full price for zero hits — the same caveat that applies to items 159 and 162. And the first verification refresh reported 114.8 s for the step and 554 s overall, worse than the 492 s baseline, because the host was loaded: persist, which this change cannot touch, was 65 s slower in it. The delta only became readable after a second run and after checking a step that had not been changed.

    Balance of the performance work (items 153-163): deep refresh upms 1 225 s -> 457 s (-63 %), achieved with seven Cypher reformulations and no new index.

  • 164. link-args-to-params had no index for the callee parameter, and lost every copycode call site ( 2026-09-06)

    The largest item left in finalize after item 163. PROFILE showed two operators carrying almost all of it, and only one of them was suspected beforehand:

    | operator | rows | db-hits | result | |---|---|---|---| | NodeIndexSeek pv (project, sourceFile) then Filter paramPosition | 62 048 947 | 63 068 522 | 29 484 | | NodeIndexSeek cm (project, sourceFile) then Filter cm.type = 'MODULE' | 11 601 107 | 11 629 527 | 27 225 |

    The index. paramPosition was in no index, so every candidate file was read whole (~140 nodes) and filtered afterwards, once per argument slot per included file (94 379 x ~15.8). A (project, sourceFile, paramPosition) index fixes it, and it is nearly free to maintain because a composite index only holds nodes carrying every property: 1 988 of 938 746 nodes (0.2 %).

    The caller module. cm was found by matching a MODULE in src.sourceFile. That threw away 99.8 % of what it read, and it silently skipped every call site inside a copycode: a .cpy has no MODULE of its own, so the lookup found nothing for 1 195 call sites. The javadoc claimed "a MODULE is 1:1 with its sourceFile" — true for .nat, false for copycode. Reaching cm through (cm)-[:CONTAINS*0..1]->(src) fixes both.

    Result, proven by A/B on the same code, graph and machine minutes apart:

    | | with index | without index | |---|---|---| | link-args-to-params | 25.2 s | 66.6 s | | finalize | 211.7 s | 256.6 s | | deep refresh upms | 458 s | 499 s | | resolve-bare-included (control, untouched) | 53.6 / 57.0 s | 54.2 / 57.2 s |

    The control step pairs the two runs, so the difference is the index and nothing else. The 66.6 s without it reproduces the original 66.7 s baseline to a tenth of a second. The whole gain is the index; the cm change contributes nothing to runtime — a first verification run accidentally shipped the cm fix without the index (the comment was written, the CREATE INDEX line forgotten) and came in at 69.5 s.

    Correctness. ARG_TO_PARAM on upms goes 17 813 -> 17 871, matching a count computed from the graph before the change. 90 ITs green (AnalysisResourceIT 85, AutoInvalidationIT 3, JavaCrossClassFlowIT 2). The size(cms) = 1 guard is the conservative bound: a copycode FUNCTION can be one node contained by several modules (item 76 counts 132 in upms), and resolving an argument against the includes of each of them is the cross-module misattribution item 77 had to fix elsewhere. None of upms's 15 120 call-site nodes is ambiguous today, so the guard drops nothing — it is there so a future parser change cannot turn this into silent bad data.

    Item 156's note that "no index helps here" was wrong and is corrected there. It holds for the caller side, which already uses the 4-property index correctly; it never applied to the callee side. That is now the second blanket "already tried, does not work" entry this run of work has had to qualify rather than trust — see item 163 for the first.

    Two mistakes worth recording. The prediction query bucketed by origin and put that bucket in the DISTINCT key, so an edge reachable both ways was counted twice and predicted 17 872 instead of 17 871 — the code was right and the prediction wrong. And a patch script opened the target file with mode "w" before encoding its content; an unencodable character then left CypherQueries.java at zero bytes. Encode first, write to a temporary file, os.replace last.

    Balance of the performance work (items 153-164): deep refresh upms 1 225 s -> 458 s (-63 %), with eight Cypher reformulations and exactly one new index.

  • 165. The node merge key carried the copycode expansion site for every node, including the 76 % that never needed it (2026-09-06)

    The persist phase was the largest remaining block, and the roadmap listed its three steps as never examined. MERGE_NODES and MERGE_POSITIONAL_NODES keyed on (type, name, sourceFile, project, ownerModule) while the widest index has four properties, so the seek stopped at (project, sourceFile, type, name) and a Filter discarded the rest. Measured on L4NCOPY.cpy: 75 076 rows read to keep 548.

    The cost is very unevenly distributed, which is what the previous attempt missed:

    | non-positional nodes | keys | nodes | max per key | |---|---|---|---| | module-own | 712 104 | 712 104 | 1 | | copycode-resident | 1 305 | 17 198 | 916 |

    So ownerModule is redundant in the key for 712 104 nodes and load-bearing for 17 198. The split follows that line: module-own nodes merge on the four-property key (exactly the existing index), copycode-resident ones on a new copySite property that is written nowhere else, so its index holds 66 445 of 938 746 nodes instead of all of them.

    Result:

    | | before | after | |---|---|---| | merge-nodes + merge-positional-nodes | 140.1 s | 54.3 s (-61 %) | | deep refresh upms | 458 s | 403 s |

    Correctness — the point that mattered most here. A mistake in a merge key does not drop an edge, it fuses two distinct nodes or duplicates them. Verified before the change: zero (project, sourceFile, type, name) keys carry both a module-own and a copycode node, and the own-key is unique (712 104/712 104 non-positional, 160 197/160 197 positional with startLine). The invariant is in GraphRepository.nodeOwner(). Verified after: node count, copycode node count and edge count identical to baseline (938 746 / 66 445 / 2 703 356) in both verification runs, and 411 ITs green.

    A migration was required and is easy to miss. The existing graph had no copySite, so the first persist would have missed every pre-existing copycode node on its key and created a duplicate beside it. BACKFILL_COPY_SITE runs in ensureSchema() before the first persist. Two traps found while verifying it: ensureSchema() is subscribed asynchronously, so Quarkus logs "started" while the backfill is still running, and CALL { ... } IN TRANSACTIONS does not propagate its inner transactions' counters, so propertiesSet() reads 0 on a run that just wrote 66 445 values. The log line now reports the re-counted result instead.

    This index is not free, unlike item 164's. A/B on the same code, graph and machine:

    | | with index | without | |---|---|---| | merge-nodes-copy | 7.7 s | 47.4 s | | merge-positional-nodes-copy | 2.7 s | 39.1 s | | merge-edges | 55.0 s | 40.2 s | | commit | 36.8 s | 26.5 s | | deep refresh | 403 s | 460 s |

    Re-measured on a quiet host against that same no-index run — controls 49.2/52.7 s against 48.4/52.9 s, so the two pair tightly — the picture sharpens: node merges 120.8 s -> 48.8 s, merge-edges 40.2 -> 47.1 s, commit 26.5 -> 32.8 s, refresh 460 s -> 374 s. The index costs ~13 s of write maintenance and saves ~72 s, net -86 s. The ~35 s first recorded here came from a loaded run and overstated the cost. It is still measurably not free — 66 445 indexed nodes are 34x item 164's 1 988 — so "a narrow index costs nothing" does not generalise.

    The cost is write maintenance, not page-cache pressure. The store had grown 2.4 -> 2.9 GB against an unchanged 2 GB cache, which made cache thrashing the obvious suspect. Measured instead of assumed: 11.2 MB read over an entire refresh, against ~120 MB per run when item 161 sized the cache at a 2.4 GB store. The cache is not the constraint and raising it to 3 GB would have taken a GB from the host for nothing.

    Item 156's experiment is now explained rather than merely recorded. The five-property merge-key index made these steps much faster and everything else ~33 % slower. The mechanism: every node carries ownerModule ("" when module-own), so any index over it spans the whole graph. Splitting the key first is what makes the index affordable.

    A prediction that was wrong, recorded because it was wrong in a new way. From db-hits alone I predicted the split without the index would be worth ~4 % (1-3 s). Measured: 140.1 s -> 120.8 s, i.e. 19 s, or ~10 s once normalised for load. Counting rows read underestimates MERGE, which also pays locking and comparison per candidate — db-hits are exact but they are not the whole cost model.

    Balance of the performance work (items 153-165): deep refresh upms 1 225 s -> 374 s (-69 %), with nine Cypher reformulations and two narrow indexes.

  • 167. resolve-bare-included sought the field by name alone, reading 9.4 M rows to keep 6 973 (2026-09-06)

    Item 163 made the placeholder name the driver, seeking realv over (project, name). That entry recorded "944 746 name hits" — which was the count after the file filter. The seek actually returns 9 363 762 rows, about 8.4 M of them other modules' placeholders (sourceFile ""), which a name index cannot exclude. Misreading that number is why the obvious next step was not taken sooner.

    The module's include files are already known at that point, so realv can be sought per file over the existing (project, sourceFile, type, name) index instead:

    | | item 163 | item 167 | |---|---|---| | realv seek | 9 363 756 db-hits | 6 973 | | following filter | 9 373 474 db-hits | gone | | all operators | ~21.4 M | ~3.6 M | | candidates | 6 973 | 6 973 |

    It trades 21 074 wide seeks for 660 816 exact ones; the empty ones are nearly free.

    | | before | after | |---|---|---| | resolve-bare-included READS | 49.2 s | 24.7 s | | resolve-bare-included WRITES | 52.7 s | 28.5 s | | the step together | 102.0 s | 53.1 s (-48 %) | | deep refresh upms | 374 s | 336 s |

    Equivalence: edge counts unchanged (+23 788/-23 788, +68 936/-68 936), node and edge totals identical (938 746 / 2 703 356), full IT suite green. Control step link-args-to-params 25.2 s against 25.9 s, so the two runs pair despite a loaded start.

    Prediction wrong again, this time by two. 75-85 s was predicted, with "anything below 70 s is unlikely" stated explicitly; the result was 53.1 s. The error was assuming the ~636 000 property writes formed a large fixed block — the lookup dominated the old runtime as well.

    A caveat that belongs with the change: the type list now drives the number of seeks rather than filtering a result, and the seek count is placeholders x include files x types, where one module can have 109 includes. A project with a wide include fan-out and few real hits could prefer the old shape.

    Balance of the performance work (items 153-167): deep refresh upms 1 225 s -> 336 s (-73 %).

  • 169. After item 167 the whole read cost was the sheer number of seeks, and they were 6x redundant (2026-09-06)

    With the seek made exact by item 167, measuring where the rest of the time went gave an unusually clean answer: full read side 8 045 ms, prefix alone 7 997 ms — the EXISTS containment check, the qualifierGroup check and the aggregation cost 48 ms together. Everything was in the seeks.

    And the seeks repeated: 330 408 (module, included file, placeholder name) triples over only 54 630 distinct (file, name) pairs, a factor of 6.05, because many modules include the same copycode and carry equally named placeholders. Resolving once per pair and joining the modules back on afterwards — the same inversion as items 159, 162 and 163, one level up:

    | | before | after | |---|---|---| | read side | 7.5 s | 3.5 s | | db-hits | 6 234 128 | 4 160 044 | | resolve-bare-included READS | 24.7 s | 18.4 s | | resolve-bare-included WRITES | 28.5 s | 21.7 s | | the step together | 53.1 s | 40.1 s (-24 %) | | finalize | 157.4 s | 146.0 s | | deep refresh upms | 336 s | 329 s |

    Equivalence was proven as set equality, not as a count. 6 839 (m, ph, realv) triples on both sides with zero difference in either direction — a count alone would not have caught a swap. Then the refresh: edge counts +23 788/-23 788 and +68 936/-68 936, totals 938 746 / 2 703 356, 411 ITs green. Control step link-args-to-params 25.9 s against 26.2 s.

    db-hits understate this kind of saving. They fall by a third while the clock halves, because what is saved is mostly per-seek overhead. Item 165 made the same mistake in the other direction, predicting ~4 % from db-hits where the real figure was ~14 %.

    The prediction landed at its pessimistic edge, and the stated caveat is why. 30-40 s was predicted by scaling the read side linearly with the placeholder count, with that assumption flagged as unverified. Result: 40.1 s. The read side halves on a resolved graph but only quarters during finalize.

    Two constraints this shape carries, both in the code comment: the type list drives the seek count rather than filtering a result, and collect materialises the hit list (~6 839 here, ~17 000 during finalize), so phase one must finish before phase two starts — a memory bound where there was none.

    Balance of the performance work (items 153-169): deep refresh upms 1 225 s -> 329 s (-73 %).

  • 170. commit in the persist log was a residual, not a measurement (2026-09-06)

    The roadmap listed commit (~40 s, the third-largest item) as an "irreducible floor, probably". It was in fact Math.max(0, totalMs - attributed) — everything the labelled statements did not account for, which includes the transaction open and the driver overhead. A guess about a number that did not measure what its name said.

    Two System.nanoTime() marks inside the transaction lambda split it three ways, keeping the managed transaction's retry semantics untouched. Result:

    | | | |---|---| | commit | 31.6 s | | unattributed | 0.0 s | | tx-open | 0.0 s |

    A representative batch line reads tx-open 1, unattributed 0, commit 275. The residual was genuine commit all along, so the roadmap's guess was right — it is now measured rather than assumed, and there is no lever here. Third negative result in a row after items 166 and 168.

    A wrong claim, retracted before it reached the code. Mid-analysis this entry was going to say "two thirds of the residual are unexplained", based on a bench transaction that committed ~270 ms for 60 000 relationships. That bench wrote ~6 MB where a real batch writes ~61 MB (1 955 MB per refresh over ~32 batches) — a factor of ten missed. At the ~49 MB/s the real numbers imply, the residual is exactly what commit I/O should cost. The measurement above confirms it.

    A real bias found on the way and fixed: PersistStats truncated every individual measurement to whole milliseconds before summing, biasing all ~570 measurements per refresh downwards and pushing ~285 ms into the residual. It now sums in nanoseconds and rounds only when printing.

    Historical commit values in these docs are not comparable with the ones printed from here on — they were the sum of all three parts. Noted in the javadoc as well.

Styling: theme tokens and the style inventory — item 196 (2026-09-22)

  • 196. Styling: theme-token usage and sx/styled inventory

    MUI v6 + Emotion: 244 sx={}, 27 styled(), 11 className, one theme in pur-ui-common/src/theme.ts, three index.css (fonts + body reset). Decided scope: theme-token usage and per-component inline inventory; plain .css only as MODULE with a STYLE per selector. theme.ts → DATA_STRUCTURE with a FIELD per token (palette.primary.dark, spacing, …); each sx/styled/style block → NodeType.STYLE under the component with its CSS property keys and REFERENCES to the tokens it uses; hard-coded literals (#005CA9, 16px) recorded as literals. Answers: where is a token used, which tokens are dead, which components bypass the theme, which components override height/zIndex. GET /modules/{name}/styles, GET /projects/{p}/styles/theme-usage?token=; ac styles, ac theme-usage. Static only — no cascade or rendered-layout claims.

    Implemented 2026-09-22. As planned, with the VALIDATE adjustments: three theme-root forms (Theme-typed values, { theme } styled parameters, the theme object imported under any name); uses counts project references only and unused is documented as "no project reference", never "dead" (MUI consumes tokens itself); tokens the code reads that no theme declares keep their placeholder and are listed with declared=false (MUI defaults, spacing, typos); sx={props.sx} is a dynamic block; nested selectors flatten to &:hover.color; compound values yield their literal parts (1px solid #D2D2D2 → 1px, #D2D2D2). Endpoints GET /theme, GET /theme/{token}/usages, GET /styles; CLI ac theme, ac theme-usages, ac styles; CSS rules from the Tier-1 scanner (STYLE per rule). Sidecar contract version 4 (themeTokens, styles, tokenRefs). Sidecar dry run on pur-ui-common: 129 tokens (111 paths, 18 constants), 169 style blocks (87 sx, 55 style, 27 styled; 66 with literals, 42 reading tokens), 89 token reads outside blocks. Verified in StylesIT and on purfe (server 326, recreate + deep refresh): 129 declared tokens, 12 undeclared ones the code reads, 304 style blocks (109 with hard-coded literals, 75 reading tokens), palette.primary.dark read 26 times; no placeholder left except the undeclared tokens. The item-195 inherited-field fix is confirmed on the same run (its 8 placeholders are gone).

Project rename — item 202 (2026-09-23)

  • 202. A project cannot be renamed

    POST /api/projects/{name}/rename with {"newName": "..."} (CLI ac project rename <old> <new>) rewrites the project key everywhere it lives: project on every AstNode (batched IN TRANSACTIONS like the delete, implicit transaction), on the DynamicCallOverrides, the entries of other projects' counterparts lists, and finally the Project shell's name. Edges carry no project. 400 INVALID_REQUEST for a blank or unchanged name, 404 for an unknown project, 409 PROJECT_EXISTS for a taken target. The metadata cache is invalidated for both names. Because the nodes move first and the shell last, an interrupted rename is finished by re-running it (the shell still answers to the old name until then). Test: ProjectRenameIT (nodes, callees, override and a peer's counterpart reference follow; old name 404; the three refusals).

    Found on the way: DELETE /api/projects/{p} left the project's DynamicCallOverride nodes behind (10 orphans after deleting upms2). The full delete now removes them; the recreate path (item 78) keeps them on purpose, they are configuration. Asserted at the end of ProjectRenameIT.

    Measured: renaming upms (938 746 nodes, 10 overrides) to upms_alt took 97 s on server 333.

Override apply rebuilds CALLS_MODULE; single-INCLUDE programs — items 200, 201 (2026-09-23)

  • 200. A dynamic-CALLNAT override is applied to CALLS at once, but the derived CALLS_MODULE edges are not rebuilt

    Found while replaying the 10 manual overrides of upms into a freshly ingested upms2: the pinned CALLS edges appeared at once, but JMIGRUN0 kept 1 CALLS_MODULE edge instead of 21 until the next refresh, so reaches and field-flow (the item-68 consumers of the derived edge) did not see the pinned targets. Now upsertDynamicCallOverride and the reset collect the calling modules of the site (CALLER_MODULES_AT_SITE, same includer fan-out as the apply) and run the scoped DELETE_CALLS_MODULE_SCOPED + BUILD_CALLS_MODULE_SCOPED in the same transaction; a project-wide reset collects every module owning a manual edge before deleting them. Test: DynamicCallOverrideIT.overrideRebuildsTheDerivedModuleEdgesWithoutARefresh (reaches false → true after the override, false again after the reset, no refresh in between).

  • 201. A program that consists of a single INCLUDE yields a second MODULE node named after the program with the copycode as sourceFile

    CopycodePreprocessor.remap mapped every host node, the module node included, onto the origin of its first expanded line; for ZDTSTBP6.nat (INCLUDE ZDTSTBC6 / END) that is the copycode, so the deep parse produced MODULE ZDTSTBP6 with the .cpy as sourceFile next to the Tier-1 shell keyed on the program file. The module node now always keeps the host file, line 1 to the host's own line count, and no viaCopycode tag. Test: NaturalParserTest.aHostThatStartsWithAnIncludeKeepsItsOwnFileOnTheModuleNode. A graph ingested before the fix keeps its stray .cpy-sourced module until the project is recreated (the item-58 sweep is keyed on pairs the fresh parse produces, and it no longer produces this one); the doc gives the one-line cleanup.

Styling robustness — item 199 (2026-09-22)

  • 199. Styling review findings: several themes, repeated token reads, CSS scanner edge cases

    From the code review of item 196. (1) The sidecar read only the first createTheme per file and the resolver demanded exactly one declaring token, so a light/dark pair in one file lost the dark tokens and a pair of theme files left every shared token an unresolved placeholder with ?unused=true reporting used tokens as unused. Now every call is read (one fact per token and file, first value wins, variants counts the themes), and the placeholder resolver redirects a read onto every real token of the name; theme/{token}/usages and styles[].tokens deduplicate the fan-out. (2) Two reads of one token on one line of a style block merged into one edge that kept only the last key; the parser now groups them and property is the comma list. (3) The CSS rule scanner: quotes are tracked so a content: "{" no longer unbalances the depth counter and drops every following rule; a ; at depth zero ends a block-less at-statement so @import no longer leaks into the next selector; a nested rule head inside an at-rule body is stripped before declaration matching (a:hover { was read as property a); a selector repeated on one line (minified CSS) gets a :col suffix instead of collapsing.

    Tests. TypeScriptCoarseScannerTest.cssRulesSurviveAtStatementsStringsNestingAndMinification, the dark theme and the doubled PRIMARY read in the parser fixture (facts regenerated, contract still 4 — variants is optional), StylesIT with a second theme file asserting both rows count the read and neither is unused.

Function-level callers across modules — item 197 (2026-09-22)

  • 197. functions/{fn}/callers cannot see cross-module calls — Java and TypeScript alike

    A cross-module call (Java cross-class, TypeScript import + call, an item-193 endpoint call from a thunk) is a MODULE -CALLS-> MODULE edge carrying callerFn and calleeMethod; the function-level callers query followed only direct FUNCTION -CALLS-> FUNCTION edges (Natural PERFORM, same-class Java), so …/AgstammLogic/functions/handleMerge/callers answered [] although the module-level callers listed AgstammController. (The roadmap's own Java example, a REST controller method, has no Java callers because it is the HTTP entry point; the gap showed on the logic class it calls.)

    Fix. Query only, no enrichment: FUNCTION_CALLERS gained a second UNION branch that joins the module edges into the target module on calleeMethod = callee.name and resolves callerFn to the FUNCTION of the calling module, honouring manualHidden; the same-module branch is untouched. Rows keep the CallRefResponse shape (edgeKind = the edge's callKind, sites from lineNo + the caller module's file). Name matching over-approximates overloads, and a call from top-level code with no enclosing function has no row here (the module-level callers still shows it). REST path and ac function-callers unchanged.

    Test. FunctionCallersCrossModuleIT: a Java method called from its own class and from another class lists both callers with their lines; a TypeScript function called from a component in another module lists the component. AnalysisResourceIT (Natural PERFORM callers, zero-caller case) stays green.

    Verified on pur/purfe (server 328, no re-ingest): AgstammLogic.handleMerge → mergeBroker at line 98; GeneralAgreementUiControllerEndpoint.createNew → the thunk in generalAgreementSlice at line 92 — both empty before.

Stale parsed edges are reaped on a deep refresh — item 198 (2026-09-22)

  • 198. Stale edges to placeholder modules survive a re-parse for Java and TypeScript

    Found while verifying item 193 on purfe: after the sidecar fix that maps pur-ui-common/dist/x to …/src/x, a deep refresh still showed 1 022 CALLS/REFERENCES edges into 46 dist placeholders next to the fresh src edges. The per-file reconcile (item 58) sweeps stale nodes of a re-parsed file, but the stale-edge reaps were Natural-only, so an edge from a surviving module that the new parse no longer produces lived forever — and its placeholder, having an edge, escaped the placeholder sweep. Only recreating the project cleared it.

    Fix. mergeEdgesBatch stamps r.ingestGen = $ingestGen on every parser-emitted edge (a re-emitted edge is re-stamped through its MERGE key). A new language-agnostic step reap-stale-parsed-edges (CypherQueries.DELETE_STALE_PARSED_EDGES) runs after merge-edges and before sweep-stale-file-nodes, keyed on the item-160 (sourceFile, ownerModule) pairs of the re-parsed files, and deletes every edge from those nodes whose stamp is older than the run's. Deep only (reconcile), like the node sweep. Finalize-built edges carry no stamp unless a resolver copied it from a parser edge, and the deep finalize that follows rebuilds those. The three Natural reaps stay (they run before the merge and gate on statement kinds). Pre-existing edges without a stamp are never reaped: the first deep refresh after the upgrade stamps, the second reaps — no project recreation needed any more; the usage doc's "recreate after a parser change" note is retired.

    Test. StaleParsedEdgeReapIT: a TypeScript component retargeted from b to c loses app/src/b from callees while an untouched file keeps it; a retarget onto a missing ./missing/d mints a placeholder that disappears once the import is retargeted again; the same for a Java class switching its call target from B to C. Existing reap ITs (StaleCallEdgeReapIT, StaleTableEdgeReapIT, RefreshReconciliationIT) and the whole server IT suite stay green.

DTO field bindings — item 195 (2026-09-22)

  • 195. DTO field binding: which component reads/writes which backend field

    No mappers exist: *UseCase DTOs sit verbatim in Redux as SvcResult<T>. Bindings are typed path expressions (<SmartInput field={AgstammUseCaseField.broker.ebene}/>, generated Fields classes) resolved by the sidecar to the dotted path broker.ebene and linked to the FIELD of the DTO interface; pur-r-vbuch binds lodash paths into the whole state. SmartInput → WRITES, SmartOutput and plain reads → READS, from the component FUNCTION to the FIELD. With 193's COUNTERPART_OF on the DATA_STRUCTURE this answers "which page edits AgstammUseCase.broker.ebene" across the frontend/backend boundary. GET /data-structures/{name}/fields gains boundBy counts; ac data-structure-fields follows.

    Implemented 2026-09-22. As planned, with these decisions from VALIDATE: the target is the generated interface's FIELD (declaring DTO Broker, not the root), reached through a binding=true placeholder resolved exactly by module → structure → field (the generic resolver never resolves module-owned structures, item 74); every hop is typed by the checker (XFields<TRoot, TSelf>), list hops are the call's result type; a prop-rooted expression is partial and its carrier prop is not part of the path; kind=prefix records a handed-on sub-object as a read of the container field; WRITES for tags matching Input$|Dropzone$|Editor$. One endpoint GET /bindings (with the field's COUNTERPART_OF columns) instead of per-field endpoints; data-structures/{dto}/fields gained boundReads/boundWrites and — a bug found on the way — now returns TypeScript FIELDs at all (the query filtered them out; item 193's doc claim was wrong). A second 193 flaw surfaced on real data: interface FIELDs were named by the bare member, and the node identity is type + name + file, so vid was ONE node under six interfaces of the generated file (six COUNTERPART_OF twins, six-fold binding rows). Fields are now <Interface>.<member> with props field/owner; data-structures/{dto}/fields and bindings report the bare member, counterparts the qualified name. Sidecar contract version 3 (bindings). Sidecar dry run on pur-ui: 204 bindings (174 field, 30 prefix; 50 partial) over 22 DTOs — 61 SmartInput, 64 SmartOutput, 23 fieldTermForRowData. Verified in BindingsIT (frontend + backend, counterpart columns, counts, no placeholder left) and on purfe (server 322, recreate + deep refresh): 257 sites over 23 DTOs, all linked to pur; counterparts?kind=field 744 fields with exactly one twin each. A leaf inherited from a base interface (datStart on AbstractHistorizedDO) resolves to the declaring interface — the single-hop case was fixed after that run and is covered by the next deploy.

The Redux store — item 194 (2026-09-22)

  • 194. Store: STORE_SLICE + FIELD, READS/WRITES/CALLS from reducers, selectors, dispatch

    Redux Toolkit: 16 createSlice/createAppSlice files, thunks via createAppAsyncThunk named <slice>/<op>, status via isSlicePending/Fulfilled/Rejected matchers, three-hop access slice → facade hook (useAgstamm, useAgstammSelector) → component; pur-r-vbuch selects by lodash path into the whole state. New NodeType.STORE_SLICE per slice with FIELD children named slice-qualified (agstamm.agstammUseCaseSvcResult) so /variables/{name}/reads|writes and flow-forward work unchanged. Reducer assignments → WRITES, selectors → READS, dispatch(action) → CALLS to the reducer/thunk FUNCTION. The store field is not flattened into the DTO: it holds SvcResult<X> and USES_TYPE the DTO DATA_STRUCTURE. GET /projects/{p}/store, GET /projects/{p}/store/{slice}/fields/{field}/reads|writes; ac store, ac store-reads, ac store-writes. context extended with store reads/writes.

    Implemented 2026-09-22. Deviations from the plan above, decided while reading the real store: the slice is named by its reducer key (gruppenprovision), not the RTK name (generalAgreement) — the key is what every selector path starts with; the store mapping is traced by the sidecar through xReducer = xSlice.reducer / export default, across workspaces. Reducers are FUNCTIONs named by the action type they handle (schluesseltabelle/updateX, schluesseltabelle/suche/fulfilled, …/matcher:isSlicePending(sliceName)), so a dispatched action's calleeMethod is the reducer. Reads come from three forms: selector arrows (incl. destructured results), wrapper hooks (useSchluesseltabelleSelector, resolved to their base path) and getState() chains (81 sites on pur-ui). Instead of one endpoint per field, one GET /store (slices with fields and counts) and one GET /store/{slice}/accesses?field=&mode=; CLI ac store, ac store-accesses. USES_TYPE field → DTO deferred to 195 (the field's dataType already says SvcResult<X>); context is not extended (the accesses endpoint and variables/<slice>.<field>/reads|writes answer). Sidecar facts contract bumped to version 2 (slices, store, stateAccesses, calls[].actionType). Cross-file reads are placeholders <key>.<field> (store=true) resolved by a new cheap finalize step in every mode; the item-74 stale-edge sweep covers store targets on a deep re-ingest. Verified on the fixtures (sidecar, parser, reader tests) and end-to-end in StoreIT; on purfe (server 318, recreate + deep refresh, 29 s): 9 slices, 43 reducers, 216 reads / 82 writes, no placeholder left; the reducer key ≠ slice name case (gruppenprovision / generalAgreement) resolves correctly. Not modelled: a slice created inside a factory function (filetransferSlice), which is also not mounted in the store.

  • 193. Webservice calls: outbound rest-endpoints on the frontend and COUNTERPART_OF to pur

    Two generated client generations exist: legacy generated/*endpoints.ts classes (one per backend controller, baseUrl + get/post members, AgstammControllerEndpoint.saveBroker.post(...)) and, in pur-r-vbuch, a hey-api sdk.gen.ts (URL literal per function, consumed via TanStack Query hooks and thunks). Both become FUNCTIONs carrying restPath (composed from baseUrl + member path), httpMethod, outbound=true, requestType, responseType — the same properties the Java parser writes (item 130), so GET /rest-endpoints lists the frontend's outbound calls with no new endpoint. Thunks and components reach them through ordinary CALLS.

    New EdgeType.COUNTERPART_OF and a counterpart enrichment step (Cypher in CypherQueries, driven like the other steps — there is no GraphEnricher SPI): outbound frontend FUNCTION → pur handler matched on httpMethod + restPath; generated DATA_STRUCTURE → Java class by simple name; FIELD → FIELD by name. Per-file reconcile DETACH DELETEs incoming edges, so the step re-runs at the end of every refresh of either project; a project setting counterparts: [..] names the partner. First cross-project MERGE in the codebase. GET /projects/{p}/counterparts? module=&limit=&offset= + ac counterparts. Same edge serves item 143.

    Implemented 2026-09-22. Sidecar (extract.mjs, contract v1 + additive endpoints and declarations[].members): the legacy generator's Endpoint classes (baseUrl + get/post members, URL template reconstructed from build<Backend>URL(…), {param} for substitutions, request/response/params types from the member's type arguments) and hey-api sdk functions (url: literal, verb from the client call). Parser: TypeScriptRestPaths splits an application base (/<name>/v<n>) off and strips the query string; endpoint FUNCTIONs carry restPath/httpMethod/outbound exactly like Java handlers, so rest-endpoints lists them with no query change beyond outbound = f.outbound = 'true'; interface members become FIELDs; a member call on an Endpoint instance is retargeted to the endpoint function (calleeMethod = AgstammControllerEndpoint.saveBroker on the module-to-module edge; note that functions/{fn}/callers does not follow such edges for Java either — item 197). Store: EdgeType.COUNTERPART_OF, project property counterparts (create/update/get/list, REST ProjectRequest.counterparts, CLI --counterpart, self-reference → 400 COUNTERPART_SELF), enrichment steps delete-counterpart-edges / link-counterparts-rest (path shape key: {param} → {}, class + method @Path composed as REST_ENDPOINTS does — keep in step) / link-counterparts-dto / link-counterparts-field, run after every finalize for the project and for every project listing it (COUNTERPART_HOLDERS), the first cross-project MERGE in the codebase. GET /projects/{p}/counterparts + ac counterparts (--kind, --unmatched, --count-only, paging). CounterpartsIT builds a Java backend and a TypeScript frontend as two projects and pins: outbound rows, call → handler incl. {vermnr} shape, unmatched = the one call nothing serves, DTO and field links, survival of a backend refresh, the setting and the two 400s. Scope note: the registered frontend is pur-ui + pur-ui-common (backend pur, base ''), so the base-stripping matters only for the excluded vstamm/vbuch clients. Verified on the deployed server (version 312) 2026-09-22: purfe deep refresh 29 s (sidecar 6.1 s + 8.6 s for the two workspaces), 285 files, 0 failures; rest-endpoints 51 outbound rows, 51 of 51 linked to pur handlers, 243 DTOs and 632 fields linked; the 52 unmatched DTOs are mirrors of JDK/framework types (Class, Comparable, Annotation, …) with no class in pur. Two defects found on real data and fixed the same day: imports of pur-ui-common resolved into its dist typings (now mapped to the src twin) and transitive packages (immer, redux) became placeholders (now external imports; Tier-1 reads node_modules directory names as externals). After the fix (server 314, project recreated because of item 198): 0 edges into dist, 7 placeholders left (from the Tier-1 pass before the node_modules rule), counterparts unchanged 51/51.

TypeScript / React analysis — item 192 (2026-09-22)

  • 192. ac-parser-typescript: language wiring, coarse scan, sidecar, import/call graph, LoC

    New Maven module ac-parser-typescript implementing LanguageParser, CoarseScanner and LineCounter for language typescript (.ts/.tsx) plus css (.css). Frontend registered as project purfe. Delivers MODULE per file, FUNCTION per exported function / React component / hook (property kind), CALLS and REFERENCES from imports and call sites, DATA_STRUCTURE + FIELD per exported interface/type (generated ones carry generated=true and javaCounterpart=<simpleName>), and LoC/SLoC so GET /loc?language=typescript works.

    Validated design decisions (do not re-derive):

    • Sidecar = ac-parser-typescript/sidecar/ (own package.json, pinned typescript, built in a node:24-slim image stage and copied into Dockerfile.jvm together with the node binary). The JVM starts it per workspace with ProcessBuilder, a timeout and --max-old-space-size=1024, reads a per-file JSON facts document from stdout, and the process ends. noEmit, no incremental (the project root is mounted :ro). excludeDirs (node_modules, dist by default) applies to the file walk only; the sidecar still reads the project's node_modules for library typings.
    • Project context like CopycodeLibrary: a TypeScriptFacts built once per ingest and passed into every per-file parse(); the Tier-1 coarse scanner is pure Java regex and needs no Node.
    • SourceFiles.Language gains TYPESCRIPT and CSS, and every == JAVA / "else Natural" branch (AstIngestService.parse/coarseScan/count, ProjectIngestService ~223 / ~1054, the seven language: 'natural'|'java' literals in CypherQueries) becomes an exhaustive switch. ProjectResource.SUPPORTED_LANGUAGES gains typescript.
    • Anonymous nodes (a sx block, a store write) use the startLine MERGE variant like DB_ACCESS, named <owner>@<line>.
    • docker-compose.yml: mem_limit 4g → 5g, frontend root mounted :ro.
    • Fixtures under ac-parser-typescript/src/test/resources/fixtures/typescript/; the JSON facts format is the tested contract between the two halves, so Java unit tests need no Node.
    • Accepted: a changedOnly refresh re-runs the sidecar over the whole workspace but re-persists only the changed files; unchanged files' facts may lag one refresh (same class as item 46a).

    Implemented 2026-09-22. New module ac-parser-typescript (TypeScriptCoarseScanner, TypeScriptParser, TypeScriptLineCounter, CssLineCounter, TypeScriptModuleNames, TypeScriptProject, TypeScriptFacts + TypeScriptFactsReader, TypeScriptSidecar) and the sidecar ac-parser-typescript/sidecar/extract.mjs (pinned typescript 5.9.3, npm ci; the JSON facts contract v1 is documented at the top of the script and pinned by the checked-in facts-pur-r-vstamm.json, which TypeScriptSidecarTest regenerates live and compares). Server: SourceFiles.Language gained TYPESCRIPT/CSS with exhaustive switches in AstIngestService (parse/coarseScan/count); SourceFiles.ingestedBy gates TypeScript/CSS to typescript projects; DEFAULT_EXCLUDE_DIRS = target, node_modules, dist; TypeScriptSidecarService builds the per-ingest context (Tier-1: package.json only; deep: sidecar per workspace, failures reported as sidecar:<workspace> pseudo-paths); the copycode stand-down of item 129 is now Natural-only by name. ProjectResource.SUPPORTED_LANGUAGES and ac project create --language accept typescript. Dockerfile.jvm is a two-stage build (node + sidecar copied from node:24-slim), the compose build context moved to the repository root with a root .dockerignore, mem_limit 4g → 5g, the frontend root mounted :ro. Config agenticcode.typescript.* in application.properties (%prod points into the image). Measured on the real frontend: 413 files, 15 s for all four workspaces, imports resolved 99.5 % (the rest: index.css, two deep moment locale paths). Decision taken while implementing: a CSS module keeps its .css in the identity, because index.css next to index.ts would otherwise collide on …/src/index and be skipped as a duplicate. 35 unit tests in the module, 232 across the build.

Closing the performance campaign (items 173-175) — 2026-09-06

  • 173. Where the deep refresh stands after items 153-172, and what is left (written 2026-09-06)

    Deep refresh of upms: 1 225 s -> 301 s (-75 %) (item 175, two clean runs: 301 s / 290 s). Nine reformulations landed, two narrow indexes were kept, five investigations produced no lever and are recorded so nobody repeats them.

    Where the time sits now. Re-measured 2026-09-06 (item 175) with two full rebuild-and-refresh.sh upms cycles, no instrumentation attached, both figures from the runs' own log lines: run 1 = 301 s, run 2 = 290 s (spread 3.7 %; run 2 is the faster one because it starts against an already resolved graph). Both columns below are run 1 / run 2.

    | persist (148.7 / 139.0 s) | | finalize (136.8 / 135.1 s) | | |---|---|---|---| | merge-edges | 47.6 / 45.9 s | resolve-field-placeholder WRITES | 23.8 / 23.9 s | | commit | 34.1 / 29.2 s | link-args-to-params | 23.7 / 23.2 s | | merge-nodes | 24.0 / 22.9 s | resolve-bare-included WRITES | 20.7 / 20.2 s | | merge-positional-nodes | 13.3 / 14.2 s | resolve-field-placeholder READS | 19.6 / 19.3 s | | merge-nodes-copy | 7.8 / 7.0 s | resolve-bare-included READS | 17.2 / 17.0 s | | sweep-* (3 steps) | 9.6 / 9.1 s | delete-resolved-placeholders | 7.1 / 7.1 s | | reap-*-edges (3 steps) | 8.3 / 7.4 s | stamp-unresolved-placeholders | 4.3 / 4.2 s | | prepare + tx-open | 1.2 / 1.1 s | remaining 44 steps | < 2.6 s each |

    Parsing is only ~9 s: parse and persist are interleaved, so the refresh is essentially persist plus finalize in equal halves. The estimates this table previously carried held up well at step level (every one within ~1 s of the measurement); only the two totals were off, ~162 s / ~140 s against a measured 148.7 s / 136.8 s.

    Already investigated without a lever — do not re-open without a new idea: merge-edges (166, the Eager costs ~5 %, removing it would freeze the property model), resolve-field-placeholder (168, write-bound: 2.6 M property writes in 45 s), commit (170, measured as genuine commit I/O once the residual was split), merge-nodes (171, two suspicions both refuted), link-args-to-params (172, seek-count-bound; the one index that helps costs 15 s to save 10).

    Candidates for a next attempt, honestly ranked by what is actually known:

    1. Fewer, larger persist batches. Measured and rejected 2026-09-06 (item 174) --- see below. The fixed per-batch cost is ~70 ms, so halving the batch count would save ~1.1 s out of 331 s.
    2. delete-resolved-placeholders (7.2 s) and the other sub-5 s finalize steps. Small, and the effort-to-payoff ratio is now clearly worse than it was at item 153.
    3. Nothing on the read side. After items 163/167/169 the field resolvers are seek-bound, and item 172 showed the only index that would cut seeks costs more than it saves.

    What is NOT worth doing, so the next person does not spend a run finding out: raising the page cache (item 166 measured 11.2 MB read over an entire refresh against a 2.9 GB store — the cache is not the constraint), and any further whole-graph index (item 172: write cost depends on how many written nodes touch the index, not on index size).

    Method notes worth keeping, each of them learned the hard way today: db-hits systematically mislead on seek-heavy and MERGE-heavy work — measure the clock as well; a single refresh is not a verdict, always check a step you did not change before believing a delta; and a read-only measurement against an already resolved graph short-circuits the field resolvers and flatters every prediction.

    Housekeeping still open: commit d6924a1 is a pure deletion of CypherQueries.java (the file was destroyed by a patch script and restored from 56314d5), and CypherQueries.java / GraphRepository.java contain mojibake bytes that make grep treat them as binary, so recursive searches silently skip them. (Two further points listed here on 2026-09-06 — the half-staged CypherQueries.java and item 161's "not yet measured" title — were resolved the same day.)

  • 174. Larger persist batches --- measured, no lever (2026-09-06)

    Item 173 listed "fewer, larger persist batches" as the only remaining idea with a plausible mechanism. It was measured before it was attempted, and the mechanism does not exist. No code change; agenticcode.ingest.batch-size stays at 200.

    How it was measured. One deep refresh of upms at batch size 200 (32 batches, 6 311 files), with the PersistStats instrumentation from item 170 and a sampler polling SHOW TRANSACTIONS YIELD estimatedUsedHeapMemory on the Neo4j side every ~3.8 s. The run took 331 s rather than the 302 s baseline; the sampler spawns a cypher-shell JVM per sample and accounts for that ~10 %. Timings below are from the run's own log lines, not from the wall clock.

    The fixed per-batch cost is ~64 ms. Summed over all 32 batches: prepare 1.2 s, tx-open 0.008 s, unattributed 0.000 s --- 37 ms per batch of genuinely fixed work. The only other fixed component is the commit floor, visible on the batches that wrote nothing at all (commit 25-28 ms). Going from 200 to 400 files removes 16 batches, i.e. ~1.0 s of ~300 s (0.3 %) --- inside run-to-run noise, which item 175 measured at 3.7 %. Everything else in persist scales with the data written; per-label figures are in item 175's table.

    Correction (item 175). The persist figures first written here came from the sampler run and were inflated by it: persist 159.3 s instead of 148.7 s (+7 %), and the persist/finalize split was given as 159 s / 163 s when it is really ~149 s / ~137 s --- the sampler cost finalize ~19 %. The per-label numbers above were withdrawn for the same reason. The conclusion is unaffected and in fact slightly stronger: the fixed per-batch cost is 64 ms, not 70 ms.

    The earlier batch-400 failure was blamed on the wrong limit. Item 173 recorded that the attempt "failed on dbms.memory.transaction.total.max". That attribution was unfounded and is withdrawn: the measured peak Neo4j transaction heap at batch size 200 is 313 MB against a 1.40 GiB limit (default 70 % of the 2 GiB Neo4j heap), i.e. 4.5x headroom --- Neo4j was never the binding constraint. The likelier constraint is the server JVM: JDK_JAVA_OPTIONS=-Xmx2500m, and the run logs show it sitting at 2 219 / 2 500 MB (89 %) throughout, with a whole batch of parsed ASTs held in memory before persisting. The original failure log is gone (the server container has since restarted), so this is the likely cause, not a proven one. Either way the upside does not justify finding out.

    Where persist time actually goes. Persist is ~149 s and finalize ~137 s (item 175) --- parsing is only ~9 s, because parse and persist are interleaved. Persist is therefore still half the refresh, but all of it is per-row write work in merge-edges / merge-nodes / commit, all three of which are already recorded as investigated without a lever (items 166, 171, 170).

  • 175. Closing measurement of the performance campaign (2026-09-06)

    Every figure in item 173's table was an estimate carried over from mid-campaign runs, and item 174 had published persist figures taken from a sampler-perturbed run. Both are now replaced by measurements from two full rebuild-and-refresh.sh upms cycles run back to back with nothing else touching the stack. Docs only, no code change.

    | | run 1 | run 2 | |---|---|---| | deep refresh, end to end | 301 s | 290 s | | persist (32 batches) | 148.7 s | 139.0 s | | finalize (51 steps) | 136.8 s | 135.1 s |

    Run-to-run spread is 3.7 %, and it is not noise: run 2 starts against a graph the preceding run already resolved, so the field resolvers find less to do. Any future A/B has to compare a first run with a first run. Take 301 s as the campaign's closing figure, since that is the condition every earlier baseline was measured under.

    The old estimates were better than expected at step level --- every single step in item 173's table came within ~1 s of its measurement (merge-edges 46.1 estimated vs 47.6 measured, link-args-to-params 23.6 vs 23.7, resolve-bare-included R+W 38.1 vs 37.9). Only the two totals were wrong, and the sampler run's figures in item 174 were wrong by more (+7 % persist, +19 % finalize). Instrumentation that polls the database distorts finalize far more than persist --- worth remembering before attaching a sampler to a run whose timings are meant to be quoted.