435 KiB
AgenticCode — Implemented Features
Completed work, moved out of x-docs/roadmap.md (which now tracks only open
items). Each entry records what was built; IDs are preserved from the roadmap
(some IDs recur across sections — they are kept as-is for traceability).
The MCP surface no longer exists. It was removed on 2026-08-04 (roadmap item 26) after intermittent session failures that were not fixable from this codebase; all 40 tools had a REST twin, so no capability was lost. REST and the
acCLI are the only access paths.Entries below still name MCP tools (
module_payload,McpQueryTools, …). Those are kept as the historical record of what each feature shipped with — they are not a description of anything callable today. Where an entry reads as present tense, read "REST + CLI".
Metrics
-
47. Generated vs. user-exit LoC/SLoC split (2026-07-13) — a project can declare a source
language(required at creation; attribute only — ingest still classifies by extension) and ageneratedDir/userExitDirpair (directory names, matched as path components likeexcludeDirs; both-or-neither,400 LANGUAGE_REQUIRED/GENERATED_USEREXIT_PAIRotherwise). A generated module already contains its hand-written user-exit twin inline, so at ingest a module undergeneratedDirwhose name also occurs underuserExitDiris annotated with that twin's LoC/SLoC (userExitLoc/userExitSloc); user-exit files are not ingested as standalone modules (the walk skipsuserExitDir, avoiding name collisions). NewUserExitMetrics.scancomputes the twins with the same per-languageLineCounter;ProjectIngestService.withUserExitMetricsstamps the generated nodes in both the coarse (Tier-1) and deep paths (correct at any depth).GET /loc(project_locMCP,ac loc) now reports, per language row and in the total:loc/sloc(total, incl. exit),userExitLoc/userExitSloc, andgeneratedExclusiveLoc/generatedExclusiveSloc(= total − exit, clamped ≥0 per file). Fields added toProjectInfo/ProjectRequest(REST + CLIproject create/updatewith-l/-g/-u; project create has no MCP tool). Tests:SourceFilesTest,ProjectResourceIT(validation),UserExitLocIT(rollup split against the counter). -
46. Deterministic per-file LoC / SLoC, per language (2026-07-12) — every file-level node (
MODULE, orDATA_STRUCTUREfor a Natural.lda/.pda) is stamped at ingest withloc(physical lines) andsloc(source lines: non-blank, non-comment). SLOC is computed per language by a newLineCounterSPI inac-parser-core(LocMetricsrecord +loc/slocproperty keys):NaturalLineCounterfollowsnatural-grammar.md§22.6.1 (full-line*/**//*, inline/*);JavaLineCounteris a char-state scan handling//,/* … */and string/char/text-block literals (so//inside a string stays code). Both the Tier-1 coarse scan and the Tier-2 deep parse call the same counter, so a module's metrics are identical at any ingest depth (numbers are reproducible and summable). Stamped inNaturalCoarseScanner/JavaCoarseScannerand, on the deep path, inProjectIngestService.withShellMetrics(viaAstIngestService.count). Surfaced onlist_modulesandmodule_context(newloc/slocfields), oninspect_node(raw properties), and via a new rollup:GET /api/projects/{p}/loc?language=&sourceFile=(PROJECT_LOCcypher →GraphRepository.projectLoc→ProjectLocDTO) — per-languagefileCount/loc/slocplus a project total, each source file counted once (collapses multi-node files viamaxper(language, sourceFile)before summing). Delivered as REST + MCP (project_loc) + CLI (ac loc) together. Tests:LocMetricsTest,NaturalLineCounterTest,JavaLineCounterTest, andAnalysisResourceIT(metrics match the counter; rollup totals/per-file consistent).
Natural ingest fidelity (WSUBPX0S findings)
Surfaced 2026-07-12 while re-analysing WSUBPX0S (project upms) via the API vs. the
grep-based analyses. Two classes of information the API could not recover; both now fixed.
Executable acceptance tests: ac-code-server/.../api/NaturalIncludeMacroAndXmlPayloadIT.java
(the _target tests are green; fixtures under src/test/resources/fixtures/natural/framework-gaps/).
-
44. Framework-mediated DB access via
INCLUDEmacros (2026-07-12) — table access performed through the generic table-access framework (INCLUDE YFRAMGC0 … '"<accessor>"' …) was invisible: theCALLNATto the generic accessor lives in the copycode member, not the including module, socallees/db-accesseswere empty.NaturalParserandNaturalCoarseScannernow recognise the statement-level framework macro (a newINCLUDE_MACROmatcher, gated on the declarativeFrameworkMacrosregistry that maps a macro name → the positional index of the accessor argument), de-quote the accessor name ('"FRELEMG0"'→FRELEMG0, args collected across continuation lines), and emit aCALLSedge taggedcallKind=INCLUDE_MACRO(newCallKindconstant; surfaced asedgeKindon callers/callees). Because it's a normalCALLS, the accessor's table access surfaces transitively via the existingdb-accesses?depth=query withvia=accessor. Scope: targeted recogniser only (roadmap option 1); general.nsccopycode expansion and direct READS/WRITES mode tagging were deliberately left out (unnecessary for the acceptance criteria). -
45. XML payload / interface schema (2026-07-12) — XML wrapper subprograms map data-area fields to XML tags via the
ADD-XML-LINEemit idiom (#W-TAG := '<tag>'/#W-VALUE := <field>/PERFORM ADD-XML-LINE, whose subroutineCOMPRESSes'<' #W-TAG '>' #W-VALUE); that tag↔field↔direction contract was not captured. The deep parser now detects the emit subroutine and its tag/value variable pair, walks the emit sequences, and models each triple as a newPAYLOAD_FIELDnode ({tag, field, direction=REQUEST}, qualifier stripped to the field name) contained by the module. Exposed as RESTGET /modules/{name}/payload+ MCPmodule_payload+ CLIac payload(PAYLOADcypher →GraphRepository.payload→PayloadFieldDTO). Scope: the emit (REQUEST) idiom with literal tags; parse-side (RESPONSE) and derived-tag normalisation remain open. -
46a. General copycode (
.cpy) expansion (2026-07-13) — statement-levelINCLUDE <member> <args>is now expanded (deep + coarse) before parsing, so a copycode'sCALLNAT/PERFORM, DB access and dataflow surface on the including module (previously invisible).CopycodePreprocessorsplices the resolved copycode body with positional&1&…substitution and records a per-line origin table; after parsing, node/edge line numbers are remapped back to their real file positions — host statements to the host, copycode statements into the.cpy(taggedviaCopycode/includedAt) — so line-based navigation stays correct despite the splice. Copycodes are resolved by a per-ingestCopycodeLibrary(.cpyname→path, read + cached); they are not ingestible as standalone modules. Framework macros (item 44), data-areaUSING, unknown members and copycodes declaringDEFINE DATAare excluded; recursion is cycle-guarded. Threaded via a newCopycodeResolverintoNaturalParser.parse/NaturalCoarseScanner.scan(Java unaffected). Acceptance:CopycodeExpansionIT(a host whose only DB access + call live in a.cpy). -
46b. XML payload derived from the interface PDA (2026-07-13) — real production XML wrappers (
WNAUTD0S-style) are generic, runtime-driven serializers (YFRAMN07tag-builder,ADD-XML-LINE/ADD-XML-ACT) with no static tag list in source, so item-45's idiom scan finds nothing. Such a module is now flaggedxmlWrapperwith itsPARAMETER USINGinterfacePda, andGET /payloadderives the contract from that PDA when no static idiom exists: each field → a payload field, wire tag = field name normalised (#→_), directionREQUEST,source=PDA(vssource=IDIOM). Idiom fields take precedence. Acceptance:NaturalIncludeMacroAndXmlPayloadIT.genericWrapperPayloadIsDerivedFromInterfacePda. The static idiom additionally handles derived tags (EXAMINE #W-TAG FOR '#' REPLACE '_'→#→_) and both directions:ADD-XML-LINE-style emit sub →REQUEST,GET-XML-LINE-style parse sub (reversefield := #W-VALUEbinding) →RESPONSE. -
46c. Ingest Global Data Areas (
.gda) (2026-07-13) —.gdafiles now classify as NaturalDATA_STRUCTUREs (like.lda/.pda) andDEFINE DATA GLOBAL USING <gda>resolves to them (theUSINGrecogniser now acceptsGLOBAL). Was: GDAs never ingested,GLOBAL USINGunresolved.
Lazy / deferred three-tier ingest
Reworked ingest from eager whole-project parsing into a lazy, on-demand model.
Three tiers: Tier 1 = cheap eager reference index (per file: nodes,
identifiers, coarse call/DB references — no deep bodies); Tier 2 = lazy deep
ingest (control flow, statement-level dataflow, precise reads/writes) triggered
on demand; Tier 3 = source served from the filesystem, no longer stored on
nodes. Reverse queries (callers, search_identifier, flow_backward) stay
answerable because Tier 1 pre-indexes coarse references globally. (Item 43,
automatic invalidation, remains open in the roadmap.)
-
36. Tier 1 reference index + tri-state ingest status (2026-07-11) — on project create, scan every source file and create nodes (no deep edges) plus the identifier index and coarse call/DB references. Each node carries a tri-state status
not-ingested/ingesting/ingestedwith a lock to serialise concurrent deep-ingest of the same node. Coarse references include Natural dynamic call targets (CALLNAT PGM-VAR), copycode/INCLUDE, continuation lines — computed by the lexer, not grep.- The Tier-1 coarse reference scan on project create: new
CoarseScannerSPI (ac-parser-core);NaturalCoarseScanneris a lexer-level single pass (module /function/data-structure shells, declared-field identifier index, coarseCALLS/READS/WRITES/INCLUDESincl.PERFORM,CALLNAT '...', dynamicCALLNAT PGM-VAR,PARAMETER/LOCAL USINGcopybooks — attributed to the enclosing subroutine exactly asNaturalParserdoes, so coarse edges merge cleanly with a later deep ingest);JavaCoarseScannerreusesJavaParserbut persists only the coarse projection (class/method shells + call/type edges), since a symbol-free Java scan would mis-resolve. Each shell carries asourceHash(feeds item 41).AstIngestService.coarseScan+ProjectIngestService.scanTier1walk the root, persist the coarse graph and run the cheap CALL_GRAPH enrichment (so cross-filecallers/calleesresolve immediately) — modules landCALL_GRAPH/NOT_INGESTED. Wired intoPOST /api/projects/{p}create (configagenticcode.tier1.scan-on-create, default true; non-fatal on an unscannable root). Covered byTier1IndexIT,NaturalCoarseScannerTest,JavaCoarseScannerTest,SourceHashTest. - The durable tri-state status + lock: real
MODULEnodes carry a durableingestStatus(NOT_INGESTED/INGESTING/INGESTED,IngestStatusenum) with aningestStatusAtstamp, kept separate fromingestDepth(enrichment tier).GraphRepository.claimIngesting/clearIngesting(CypherQueries.CLAIM_INGESTING/CLEAR_INGESTING) implement a best-effort DB claim — Neo4j locks only atSET, so it is not a hard mutex; correctness rests oningestModulebeing idempotent, with true intra-JVM exclusion still from the in-process monitor.DeepIngestCoordinatorsetsINGESTINGat the start of a by-name deep ingest (markIngestDepth(FULL)flips the tree toINGESTED; failure rolls back toNOT_INGESTED), and a cross-process loser coalesces by pollingmoduleIngestStateuntil the winner reachesFULL, re-claiming if the claim is released or goes stale (agenticcode.deep-ingest.ingesting-ttl-seconds, default 1800; poll boundclaim-wait-seconds, default 120). Covered byDeepIngestStatusITandCoalescingIT(concurrent deep ingests converge to a singleFULLnode).
- The Tier-1 coarse reference scan on project create: new
-
37. Tier 2 lazy deep-ingest on every call (2026-07-11) — every MCP/API call deep-ingests the nodes it touches, transitioning them to
ingested. Decision: this includes fan-out result sets (e.g.search_identifier,callers), not just the explicitly named node — so the node budget (item 39) is the primary cost bound and its default must be a real, tuned number. Traversal queries (call_tree,flow_*) deep-ingest along the path as they walk it.- Module-gated field-level queries auto-trigger:
DeepIngestCoordinator.ensureDeepruns a scopedingestModule(with a per-(project,module)coalescing lock) from thewithDeepModulechoke point in bothMcpQueryToolsandAnalysisResource;flow-forward/backward,field-flowetc. auto-deep-ingest instead of returning409, falling back to the hint only when the module does not resolve. - Fan-out result-set warm for
callers,search_identifier, andcall_tree: blocking-then-rerun — each query runs against the graph as-is,DeepIngestCoordinator.ensureDeepManydeep-ingests the surfaced result-set files (by path, via the multi-rootProjectIngestService.ingestFiles) under a fan-out node budget (agenticcode.deep-ingest.fanout-nodes, default 50), and the query re-runs only if the warm deepened something. Wired at thewithFanoutWarmchoke point in both surfaces. Note:callerswarms already-surfaced callers (downstream precision) but cannot reveal a caller invisible at the coarse tier. - Cross-module
flow_*path-ingest (37a) forflow-forward,flow-backward, andfield-flow: an ingest-and-re-traverse fixpoint at thewithFlowPathWarmchoke point — afterensureDeep(startModule)it runs the trace and, each round, deep-ingests the frontier (surfaced modules seeded with the start module, plus their direct callee modules) in one scope (DeepIngestCoordinator.ensureFlowFrontier→GraphRepository.flowFrontierSourceFiles→ProjectIngestService.ingestFiles), then re-traverses. Ingesting the whole frontier (callers included, even ifFULL) re-links each caller into the freshly-ingested callee params, letting a trace cross into a dynamically-dispatched callee. Bounded byagenticcode.deep-ingest.flow-rounds(default 3) + thefanout-nodesbudget; stops early at the fixpoint. Covered byAnalysisResourceIT.flowForwardPathWarmCrossesIntoDynamicallyDispatchedCallee. - Global concurrency cap: all deep-ingest/warm work in
DeepIngestCoordinator(by-nameensureDeep, fan-outensureDeepMany, flowensureFlowFrontier) is gated by a fairSemaphore(agenticcode.deep-ingest.max-concurrent-warms, default 2), acquired viatryAcquireonly around the actual ingest (never while a coalesce-waiter waits); on timeout (warm-acquire-timeout-seconds, default 10) the warm is skipped and the query returns its Tier-1 answer.
- Module-gated field-level queries auto-trigger:
-
38. Deep-ingest depth cap (2026-07-10) — the by-name deep-ingest BFS (
ProjectIngestService.ingestModule) is bounded bymaxDepthhops from the named module: default 5 (agenticcode.deep-ingest.default-depth), clamped to the ceiling 20 (agenticcode.deep-ingest.max-depth). Hitting the bound truncates + reports (not a hard error) — the reached modules are markedFULLand the response carries anIngestSummary.Truncation. Override via?maxDepth=(REST, now onrefresh/{name}), therefreshMCP tool'smaxDeptharg, andac refresh --max-depth(CLI). -
39. Deep-ingest node budget (2026-07-10) — alongside the depth cap, the BFS ingests at most
maxNodesfiles (default 300,agenticcode.deep-ingest.default-nodes); overflow truncates + reports the same way (reasonNODES/NODES_AND_DEPTH). Override via?maxNodes=(REST), therefreshMCP tool'smaxNodesarg, andac refresh --max-nodes(CLI). Defaults are@ConfigPropertyand should still be tuned against the realacproject. -
40. Unresolved-reference nodes (2026-07-12) — when Tier 1 sees a variable / dynamic call target it cannot resolve, store it as a deduped placeholder node/edge flagged
unresolved, so reverse queries (e.g.callers) surface dynamic call sites while still distinguishing them from resolved edges. Resolution pass runs during Tier 2.- Placeholder nodes (
sourceFile = "") were already deduped (MERGE ontype,name,sourceFile,project); this adds theunresolvedflag. New enrichment stepSTAMP_UNRESOLVED_PLACEHOLDERS(runs last infinalizeProject/finalizeProjectScoped, project-wide, cheap, idempotent, every mode) setsunresolved = (no real definition of that (type,name) is ingested)— the resolution pass: a reference readsunresolved=falseonce its target file is ingested (a real twin exists),truewhile genuinely dangling. Surfaced explicitly onsearch_identifier(IdentifierMatch.unresolved) and, for free, oninspect_node(nodes/{id}returns the whole node). Reverse call queries (callers/callees) already surface these targets as entries with a blanksourceFile. Covered byUnresolvedRefIT(dangling →true; ingest target →false).
- Placeholder nodes (
-
41. Tier 3: stop storing source on nodes (2026-07-12) — drop stored source text; serve
module_source/node_sourceby openingsourceFile(project-relative path + content hash) and slicingstartLine..endLine. Stale check: on read, compare the stored content hash against the current file; on mismatch, return a structuredSTALE_SOURCEerror ("run refresh") instead of slicing — never serve current lines against old line numbers.- Source text was already not stored on nodes (served off disk by
SourceSnippetServicesince item 28); this adds the mandatory stale check. SharedSourceHash(SHA-256) is stamped on each file'sMODULE/DATA_STRUCTUREshell at ingest — by the Tier-1 scanners and (for robustness) the deep full-parse paths (ProjectIngestService.withSourceHash), surviving re-ingest viaSET +=. On asourceread,GraphRepository.sourceHash(project, sourceFile)supplies the stored hash andSourceSnippetService.readrecomputes the current file's hash; on mismatch it throwsStaleSourceException, which both REST (409 STALE_SOURCE) and MCP (STALE_SOURCEtool error) return with a "re-ingest (refresh)" message. A legacy file with no stored hash skips the check (serves unchecked). Covered byStaleSourceIT(fresh→200, edit→409, re-ingest→200). Copycode/INCLUDE slices remain raw pre-expansion file text. Note: a fan-out warm query (e.g.search_identifier) re-ingests and re-hashes a surfaced module, clearing its staleness as a side effect.
- Source text was already not stored on nodes (served off disk by
-
42.
refreshtool (manual invalidation) + removeingest_all/ingest_module(2026-07-12) — remove the eager bulk/one-module ingest tools across all three surfaces (MCPMcpIngestTools, RESTAnalysisResource,acCLI) and replace with arefreshoperation that re-runs Tier 1 + deep ingest. Interim manual invalidation until item 43.refreshis the single (re-)ingest surface across REST + MCP + CLI;ingest_all/ingest_module/ingest-call-graphwere removed from all three. RESTPOST /refresh(whole root, call-graph tier;?deep=truefor field-level) andPOST /refresh/{name}?maxDepth=&maxNodes=(deep re-ingest a module tree); MCP onerefresh(project, module?, deep?, maxDepth?, maxNodes?)tool; CLIac refresh [name] [--deep] [--max-depth --max-nodes].ProjectIngestService.refreshProject/refreshModuledelegate to the existing (now internal) ingest primitives — a whole-root refresh runs the full-parse call-graph pass (a superset of the create-time coarse Tier-1 scan, so intra-module dynamicCALLNATstill resolves), and is the remedy for aSTALE_SOURCEread (re-stampssourceHash). It MERGEs current files (does not wipe; deletions await item 43). All ~20 ITs bootstrapping via the old endpoints were migrated (behavior- preserving URL swaps);CLAUDE.md+mcp-api-usagedogfooding steps now sayrefresh. Covered byRefreshIT(whole-project + module refresh work; removed endpoints 404).
-
43. Automatic invalidation (hash-based) + deleted-file sweep (2026-07-15) — the graph now self-heals when source files change or are deleted on disk, without a manual
refresh. Chosen approach: hash-based, lazy (on-access) — not a filesystem watcher (no background thread; reuses the item-41sourceHash). Two parts, gated byagenticcode.auto-invalidate.enabled(defaulttrue).- Part A — change detection → auto re-ingest.
DeepIngestCoordinator.ensureDeep/ensureDeepMany(the two deep-ingest choke points every field-level query routes through) no longer short-circuit aFULLmodule blindly: they compare the file's currentSourceHashto the stored one (same compare as the/sourcestale check). AFULL-but-changed module is invalidated (GraphRepository.invalidateModule→ shell reset toCALL_GRAPH/NOT_INGESTED, soisFull()is false and the coalesce loop re-ingests) and re-parsed — with the item-58 reconciliation purging its stale nodes. Conservativefalse(never thrash aFULLmodule) when the flag is off, no source file, no stored hash (legacy node), or the file is unreadable/missing (a deleted file is Part B's job). Per-module granularity: only the queried module's own file is checked, not its whole dependency tree (a changed dependency is caught when it is queried by name). - Part B — deleted-file sweep. A whole-project
refresh(ProjectIngestService.refreshProject) now reconciles against the filesystem:GraphRepository.distinctSourceFiles→ any realsourceFileno longer present on disk →deleteNodesForSourceFiles(DETACH DELETE). Closes the item-58 gap (a refresh MERGEd current files but left nodes for deleted files). Reconciles against disk existence, not the walked-candidate list, so on-demand dependency files (PDAs/LDAs/.cpy) that exist but aren't top-level candidates are never wrongly swept. - No new REST/MCP/CLI surface — Part A is internal to the query path; Part B changes
refreshsemantics (now also drops deleted-file orphans). Both are documented behaviour changes, so the three surfaces stay in sync by construction. Covered byAutoInvalidationIT(edit aFULLmodule on disk → a field query auto-re-ingests,409 STALE_SOURCE→200, no manual refresh) andDeletedFileSweepIT(delete a file →refreshremoves its module + field nodes, survivors kept).PlaceholderResolveNullLineIT(which seeds phantom nodes for files it never writes) opts out via a@TestProfiledisabling the sweep.
- Part A — change detection → auto re-ingest.
Layered / performance-aware ingest
-
P2-c.
ingest-alldefers field resolution AND dataflow by default (fast whole-root) (2026-06-21) — even with P2-b's scoped per-program resolution, a whole-rootingest-allstill ran the global field-placeholder resolution AND the argument→parameter dataflow — measured live onupms: the field WRITES pass took ~26 min andlink-args-to-params~25 min (each one transaction over ~100k+ edges). So/ingest-allnow defaults to the fast call-graph pass (EnrichmentLevel.CALL_GRAPH): placeholder resolution (CALLS/INCLUDES/USES_TYPE/EXTENDS/IMPLEMENTS) + intra-module dynamicCALLNAT+ polymorphic (CHA) fan-out. It defers the expensive field-placeholder resolution and the dataflow + cross-module dynamicCALLNATdispatch to a scoped per-program deep ingest (/ingest/{name}, P2-b) — which fans out top-down from the analyzed program (the user's insight: dynamic dispatch is a top-down concern, not needed globally bottom-up). The old full whole-root behaviour is available viaPOST /ingest-all?deep=true. Implementation:EnrichmentLevel { CALL_GRAPH, FULL }(each carryingdataflow/resolveFieldsgates),finalizeProject(project, EnrichmentLevel),ingestAll(project, deep),AnalysisResource.ingestAllgains?deep. The default tags modulesCALL_GRAPH, so field-level endpoints return the409 NOT_DEEPLY_INGESTEDdeep-ingest hint (P2-a). Covered byAnalysisResourceIT(ingestAllDefaultDefersDataflowAndFieldResolutionToDeepIngest: fast ingest → intra-dynamic resolved, cross-dynamic absent +409on field-flow →/ingest/{name}→ cross-dynamic resolved200); the IT'singestAll()helper uses?deep=trueso existing field/dataflow assertions still exercise the full path. An earlier attempt kept dataflow+cross-callnat in aCALL_GRAPH_DATAFLOWmiddle tier; liveupmstesting showedlink-args-to-paramsalone was ~25 min, so it was dropped.
-
P2-b. Scope deep field resolution to the program tree (2026-06-20) — in the accumulating one-project model (call-graph everything, then deep-ingest programs over time), a by-name deep ingest previously ran the unscoped field-placeholder resolution over the whole project (~110k edges → ~28 min once the full call graph is loaded). Added
$names-scoped variants of the field-resolution and dataflow queries (resolvePlaceholderFieldTargetsScoped,…ByNameScoped,resolveBareIncludedFieldTargetsScoped,LINK_ARGS_TO_PARAMS_SCOPED,…_JAVA_SCOPED) that start from the just-ingested program tree's modules instead of scanning all placeholder edges.GraphRepository.finalizeProjectScoped(project, moduleNames)runs them (call-graph resolution, placeholder cleanup, CHA fan-out stay project-wide/idempotent);ingestModulenow calls it with the BFS-collected module names. So deep ingest stays fast (≈ the small-project 8.5 s case) even when the project already holds the whole call graph. Correctness covered byIngestModuleITfield-flow/dataflow tests (now exercising the scoped path); all ingest/analysis ITs green. Live perf re-validated onupms(2026-06-21): whole-codebaseingest-call-graph= 6309 modules in 67 s (placeholders left intact), then a by-name deep ingest ofKDWWIFN0(56-module tree) with the full call graph already loaded ran in 5.2 s — confirming the scoped resolution avoids the old unscoped whole-graph path (~28 min). Depth guard verified end-to-end:flow-forwardreturns 200 for theFULLmodule,409 NOT_DEEPLY_INGESTEDfor a call-graph-only module (ACCNPE01), and409 NOT_INGESTEDfor an un-ingested name (JE999), each with an actionablenextAction. -
P2-a. Layered ingest: call-graph mode + ingest-depth guard (2026-06-20) — field-placeholder resolution is volume-bound (whole-codebase ≈110k edges → not viable even cycle-safe/APOC node-global; measured), while per-program deep ingest is fast (~8.5 s for
KDWWIFN0's 56-file tree). So ingest is now layered: (1)POST /api/projects/{p}/ingest-call-graph— whole-root, persists everything but runs only the cheap call-graph/module-level enrichment (CALLS/INCLUDES/USES_TYPE/EXTENDS/IMPLEMENTS + CHA fan-out), skips field-placeholder resolution, and leaves placeholders intact for later deep ingest. Tags modulesingestDepth=CALL_GRAPH. (2)POST /ingest/{name}(by-name) andingest-allrun full enrichment and tag the ingested modulesingestDepth=FULL(accumulates over time in one project). (3) Field-level endpoints (flow-forward/flow-backward/field-flow) check the target module viaGraphRepository.moduleIngestStateand, when notFULL, return an agent-actionable409 { status: NOT_DEEPLY_INGESTED|NOT_INGESTED, module, detail, nextAction:{method,path} }instead of a misleading empty result. NewIngestDepth/ModuleIngestState,finalizeProject(full)splittingenrichmentSteps(full)into call-graph vs field groups, depth-marking Cypher, per-step finalize logging (label/duration/rows/JVM heap). New ITcallGraphIngestEnablesCallGraphAndGuardsFieldFlow(409→200 after deep ingest); fullAnalysisResourceIT+ingest ITs green.
Parser / model
- 1. Persisted per-language source/module kind (done 2026-07-07,
scoped down after investigation) — add an
AstNodeproperty capturing each module's finer kind beyond the coarseNodeType. Delivered: JavamoduleKind=CLASS/INTERFACE(exact, fromtype.isInterface()); NaturalmoduleKind= a best-effortPROGRAM/SUBPROGRAMguess (has a top-levelPARAMETERsection →SUBPROGRAM).GET /modulesgained a?moduleKind=filter. Not attempted (found to be a separate, larger feature, or genuinely unrecoverable from source text alone): Javaenum/record—JavaParserdoesn't parse these asMODULEnodes at all today (onlyClassOrInterfaceDeclaration), so this needs a new node-family traversal, not just a property; NaturalCOPYCODE/MAP/GDA— pernatural-grammar.md§22.2,PROGRAM/SUBPROGRAM/SUBROUTINEare real SAG Natural catalog metadata, andmain-program/external-subprogram/subprogramare syntactically identical productions — no reliable source-level signal exists for these without the original catalog. Seex-docs/agent-api-usage-ac-implementation.mdstep 0 for the heuristic's limits.
Foundation (schema, health, errors, ingest CLI)
- 1. Neo4j indexes & constraints — index
(:AstNode) (project, name)and(project, type, name), unique constraint on(:Project) (name). Created on application startup viaSchemaInitializer(idempotentIF NOT EXISTS). - 2. Health checks —
Neo4jHealthCheck(@Readiness) verifies Neo4j connectivity viaDriver.verifyConnectivity(). - 3. Structured error responses for
AnalysisResource— sharedErrorResponserecord ({error, code, details}), reused byProjectResource. - 4. Project-existence validation — ingest and all project-scoped query
endpoints return
404 PROJECT_NOT_FOUNDif the project hasn't been created. - P0-a. Uniqueness constraint/index on
AstNode.id—mergeEdgematches both endpoints byid(MATCH (a:AstNode {id: $sourceId}), (b:AstNode {id: $targetId})), butidwas not indexed (only(project, name)and(project, type, name)were). Every edge merge did a full label scan over allAstNodenodes accumulated so far, so ingest got progressively slower as more files/projects were ingested. AddedCREATE CONSTRAINT ast_node_id_unique IF NOT EXISTS FOR (n:AstNode) REQUIRE n.id IS UNIQUEtoSCHEMA_STATEMENTS, giving edge merges O(1) lookups regardless of graph size. - P0-b. Batch ingest writes with
UNWIND—GraphRepository.save()previously issued onetx.run()per node and per edge (N+M round trips per ingested file). Replaced withCypherQueries.MERGE_NODES(UNWIND $nodes AS n MERGE ...), one round trip for all nodes of a file, andCypherQueries.mergeEdgesBatch(EdgeType)(UNWIND $edges AS e MATCH ... MERGE (a)-[:<TYPE>]->(b)), one round trip per distinctEdgeTypepresent in the file (edges grouped viaCollectors.groupingBy(AstEdge::type)) — Cypher relationship types can't be parameterized, so true O(1) isn't possible, but this cuts round trips from N+M to 1 + (number of distinct edge types, usually 1-3). Same merge semantics, low risk. - P0-c.
--exclude-diroption foringestCLI command — recursive ingest now skips any file whose path (relative to the ingest root) contains a component matching one of the repeatable-x/--exclude-dir <name>values (case-insensitive), e.g.--exclude-dir user_exit --exclude-dir testto skip subprograms already represented ingenerated_sources. Excluded files are printed asSKIP (excluded) <file>and counted in a newexcludedsummary counter alongsideingested/skipped/failed. - P0-d.
ingest-module <folder> <moduleName>— dependency-driven ingest (2026-06-15) — new CLI command that ingests a single module and its transitiveCALLNAT/PERFORM-external/EXTENDS/IMPLEMENTStargets andINCLUDE/USING'd data areas, located by filename stem (case-insensitive) within a given folder.POST .../ingest/{java,natural}now returns202 {"dependencies": [{"name":..., "type": "MODULE"|"DATA_STRUCTURE"}]}(newAstIngestService/AnalysisResource.IngestResponse/DependencyRef), derived from placeholder nodes (sourceFile="") created by the parsers. The CLI does a BFS over these names: found files are ingested and their dependencies enqueued; names not found in the folder are printed asUNRESOLVED <name>and counted separately (not a failure).--exclude-dirapplies to the file index, so dependencies under excluded folders are reported as unresolved. Shared file-walk helpers (readSourceFile,endpointFor,isExcluded) extracted fromIngestCommandintoIngestSupport. New ITnaturalIngestReturnsUnresolvedDependencies. - P0-e.
.lda/.pda(Natural data area) parsing (2026-06-15) —NaturalParsernow recognizes.lda/.pdafiles (viaIngestSupport.endpointFor) and parses their wrapper-less field-export format (<TYPE><LENGTH><LEVELDIGIT><NAME> ... CONST<...>|INIT<...>, glued type/level/name prefixes,R <level><name>REDEFINE markers,*-comment lines, multiple top-level (level 1)DATA_STRUCTUREroots per file) into the sameDATA_STRUCTURE/VARIABLE/CONSTANTnode shape asDEFINE DATAfields, sodata-structure-fieldsworks unchanged. Combined with item 6 (placeholder resolution), an ingested LDA/PDA now links up with theDATA_STRUCTUREplaceholder created byLOCAL/PARAMETER USINGin referencing.natprograms. New unit tests (YFRAML01_SAMPLE.lda,WGEAGL01_SAMPLE.pda) and IT (ingestingLdaResolvesIncludePlaceholder). - P0-f.
GET /variables/{name}/writesand/reads(2026-06-15) — queries returning everyFUNCTION/MODULEthat has aWRITES/READSedge to aVARIABLE/CONSTANTwith the given name, withsourceFileand containingmodule(CypherQueries.variableReads(maxDepth)/variableWrites(maxDepth),GraphRepository.variableWrites/variableReads,VariableAccessLocationrecord). Optional?module=...&depth=Nquery params restrict results to that module or anything in its transitiveCALLStree (up todepthhops, sameagenticcode.call-tree.default-depth/max-depthconfig and clamping ascall-tree); the Cypher computes each writer/reader's containing module viaCONTAINS*0..and checksowner.name = $module OR EXISTS { (module)-[:CALLS*1..depth]->(owner) }. New CLIvariable-writes/variable-reads <name> [-m/--module <name>] [--depth N].
Ingest performance
- 24. Index the stale-file sweep — persist was a label scan per file (2026-07-16) — the
sibling defect to P1-m below, one statement further on: P1-m indexed the node MERGE key, but
DELETE_STALE_FILE_NODES(the item-58 reconcile sweep, which runs in the samemergeResultstransaction) matches on(project, sourceFile)— and a Neo4j composite index only applies when every one of its properties is constrained, so the 4-property(project, sourceFile, type, name)index does not cover that 2-property prefix. The planner therefore fell back toNodeByLabelScan+Filter, and because theUNWIND $filesbatch drives it through anApply, it scanned once per file: measured 564,451 db hits per file on a 282k-node graph, i.e. ~56M node reads per 200-file batch. The scan covers the wholeAstNodelabel — all projects — which is why per-batch cost grew with total graph size (2m29s early vs 2m48s late). Fix is one line inSCHEMA_STATEMENTS:ast_node_project_sourcefileon(project, sourceFile). The plan becomes aNodeIndexSeekat 137 db hits per file (~4000×). Verified that the new, less selective index does not displace the 4-property one for the MERGE keys (still 1–2 db hits). Also servesSOURCE_HASH(item 43 auto-invalidation) and thesourceFile:""placeholder sweeps, which had the same blocked shape. Index-only;ensureSchemaadds itIF NOT EXISTS, so existing deployments pick it up on the next restart — no migration. Measured live on a fullupmscall-graph refresh (6311 files, reconcile on): whole refresh 179 s end-to-end — parse ~63 s, persist 104.5 s across 32 batches (median 1.7 s/batch, min 0.2 s, max 10.7 s), finalize ~12 s — against ~2.5–2.8 min per batch logged before the fix, and with no per-batch growth left. A second refresh took 95 s and reproduced the graph exactly (454,255 nodes / 1,495,146 rels, no duplicate MERGE keys), i.e. idempotent. Note the pre-fix graph held only 216,492 upms nodes: the slow run evidently never completed, so the fix also produced the first complete upms graph rather than merely a faster one. Regression testStaleFileSweepIndexITasserts the cause — itEXPLAINs the realDELETE_STALE_FILE_NODESconstant and requires aNodeIndexSeekonAstNode(project, sourceFile)(the planner names indexes by properties, not by index name), afterCALL db.awaitIndexes()since aPOPULATINGindex is invisible to the planner and would make the assertion race startup. Vacuity-checked: with the schema line removed the plan falls back to the filtered label scan and the test fails. - 25.
ingest-allperformance: batch persist transactions — already implemented (closed 2026-07-16) — the roadmap item claimed "persiststill opens one transaction/session per file"; that was stale.ProjectIngestServicechunks files intopersistBatchcalls sized byagenticcode.ingest.batch-size(default 200), andGraphRepository.mergeResultsaggregates each chunk into batchedUNWINDMERGEs (one per node class, one per edge type) in a single transaction — delivered earlier as P1-k below. Closed as done; it was item 24 above, not batching, that actually held persist back. - P1-m. Index the node MERGE key to kill O(N²) persist (2026-06-20) —
on a whole-root ingest of
upms(~6309 files), per-batch persist time grew with the graph (batch 1 ≈ 1 s, batch 8 ≈ 4.5 min) — quadratic.MERGE_NODES/MERGE_POSITIONAL_NODESmatch on(type, name, sourceFile, project)but the only supporting index was(project, type, name); for low-selectivity names that recur project-wide (CONTROL_FLOWIF/FOR/REPEAT/DECIDE,VARIABLE/DATA_STRUCTUREFILLER/#CODE/SQLERR) that bucket grows with the graph and MERGE filteredsourceFile/startLinein memory over it. Added composite indexast_node_project_sourcefile_type_nameon(project, sourceFile, type, name)so MERGE seeks per-file (a handful of nodes). Index-only;ensureSchemaadds itIF NOT EXISTS. - P1-l. Split enrichment into per-statement transactions (2026-06-20) —
finalizeProjectran all ~15 project-wide enrichment statements in one transaction; on the fullupmsgraph that single transaction blew Neo4j'sdbms.memory.transaction.total.max(5.4 GiB) and aborted, so placeholders never resolved.finalizeProjectnow runs each ordered, idempotent statement (enrichmentStatements()) in its ownexecuteWriteWithoutResult, bounding peak per-transaction memory. Same graph result (ITs green); a lone over-large step would still needCALL { … } IN TRANSACTIONS(not yet required). - P1-k. Batch persistence in
ingest all(2026-06-19) —ingestAllpersisted one file per transaction (fresh session + ~8 statements + commit), sequentially — fine for a small slice but slow once the P1-j dedup fix made it ingest thousands of files (each round trip is fixed overhead × N files). NewGraphRepository.persistBatch/AstIngestService.persistBatchpersist a chunk of files in one transaction, aggregating all files' nodes/edges into batchedUNWINDwrites (one statement per node label / edge type per chunk instead of per file).ingestAllnow chunks the non-conflicting results byagenticcode.ingest.batch-size(default 200) and runs onefinalizeProjectafter. Correctness subtlety: different files emit the same placeholder (CALLNAT/USING target,sourceFile="") with distinct UUIDs but one MERGE key; persisting all nodes before all edges collapses them (last id wins) and would orphan earlier files' edges.mergeResultsdedupes nodes by MERGE key (canonical = first id) and remaps every edge endpoint onto the canonical id — the per-file transactions previously masked this. Caught byAnalysisResourceIT(interface fan-out + shared-PDA field-flow) during implementation; full ingest/analysis ITs green. - P1-j. Fix over-aggressive duplicate detection in
ingest all(2026-06-19) —ProjectIngestService.ingestAllflagged cross-file duplicates by scanning every realMODULE/DATA_STRUCTUREnode, so two unrelated Natural programs that each contain an inlineDEFINE DATAgroup with a common name (FILLER,SQLERR,#ERROR-GROUP, …) were treated as duplicates and both whole files skipped. On theupmscodebase this silently dropped 4379 files (3894 of 3895 "duplicate" identities were inlineDATA_STRUCTUREnames; only 1 was a realMODULE), so e.g.W-MNT-N0/KDWWIFN0never ingested and their callers/callees couldn't resolve. NewfileIdentities(Parsed)keys duplicate detection on the file's own identity only: realMODULEnodes (Java FQN-aware) for program/class files, and a single(DATA_STRUCTURE, stem)for.lda/.pdadata-area files — inline group nodes are excluded. RealMODULEand data-area-file duplicates are still reported. In-memory only (no schema change/re-ingest of node shape). New ITsharedInlineDataStructureNamesDoNotSkipModules; fullAnalysisResourceIT+IngestModuleITgreen. A persisted per-language source/module kind is tracked separately as future work. - P1-i. Defer enrichment in by-name ingest (2026-06-19) —
POST /ingest/{name}(ProjectIngestService.ingestModule, the CLIingest <module>path) calledsave()per file in its dependency BFS, which re-ran the full ~15-statement project-wide enrichment (placeholder resolution + dataflow + polymorphic fan-out) after every file — ≈O(N²) over a growing graph, the dominant cost for Natural modules that fan out throughCALLNAT/PERFORM/USING(observed oningest KDWWIFN0). Now mirrorsingestAll:persist()per file (merge only), then onefinalizeProject()after the BFS. Enrichment is idempotent/order-independent, so the end graph is identical (verified byIngestModuleIT, which exercises cross-module dataflow that depends on enrichment). Removed the now-unusedAstIngestService.save()/GraphRepository.save(). - P1-h. Run DB endpoints off the event loop (2026-06-19) — the Neo4j
blocking driver is wrapped in
Uni.createFrom().item(...)in everyGraphRepositorymethod, whose supplier runs on the subscribing thread. TheUni-returning JAX-RS query methods carried no@Blocking, so for those the driver ran on (and blocked) the Vert.x event loop — latent until a slow query tripped the 2 sBlockedThreadChecker(observed onDELETE /api/projects=CLEAR_ALLfull-graphDETACH DELETE, blocked ~3.3 s). Fixed by moving@Blockingto class level onAnalysisResourceandProjectResourceso all endpoints dispatch to a worker thread (the two now-redundant method-level@Blockingon the ingest endpoints were removed). Verified via access log (executor-thread-*instead ofvert.x-eventloop-thread-*);ProjectResourceIT+AnalysisResourceITgreen. - 23.
ingest-allperformance: enrich once, not per file (2026-06-19) —ingest-allran the project-wide enrichment block (placeholder resolution for allRESOLVABLE_EDGE_TYPES, field-target resolution, bare-field redirection, placeholder deletes,LINK_ARGS_TO_PARAMS/_JAVA,LINK_CALLS_TO_IMPLEMENTATIONS) inside everyGraphRepository.save, i.e. once per file — each pass scans the whole project graph, so cost grew ~O(N²) in the file count and dominated large Java scans. Split persistence from enrichment:GraphRepositoryrefactored intomergeResult/runEnrichmenthelpers, exposingpersist(project, result)(MERGE only) andfinalizeProject(project)(enrichment once), withsaveretained asmerge + enrichin one tx for single-file/by-name ingest.ProjectIngestService.ingestAllnowpersists each survivor then callsfinalizeProjectonce (skipped when nothing persisted). The enrichment is idempotent and order-independent, so the final graph is identical; verified by the existingAnalysisResourceIT(ingest-all-based) suite.ingestModuleunchanged (stillsaveper file).
Call graph, edges & provenance
- P1-g. Language-aware
edgeKind(2026-06-19) —edgeKindincallers/calleeswas always derived from the target node type as Natural'sCALLNAT/PERFORM, mislabelling Java calls (e.g. anew Foo()constructor showed asCALLNAT). Each parser now stamps acallKindproperty on theCALLSedge using the newCallKindenum (ac-parser-core): Natural →CALLNAT/PERFORM; Java →METHOD_CALL(intra- and cross-class invocations) /CONSTRUCTOR(new). Stamping at parse time lets Java distinguish constructors from method calls, which the previous target-type-only logic could not.CypherQueries.callers/calleesnow returncoalesce(r.callKind, <legacy CASE>)so graphs ingested before the stamp fall back to the old labels (no forced re-ingest; re-ingest needed for Java precision). Synthetic inheritance edges (LINK_CALLS_TO_IMPLEMENTATIONS) copys.callKind = r.callKind. New IT assertions for Java (METHOD_CALL/CONSTRUCTOR) and Natural (PERFORM/CALLNAT). - 6. Cross-file placeholder resolution (2026-06-15) — decided against a
generic
EnrichmentPipeline.run()(no second use case identified yet;enrichment/stays empty). Instead,GraphRepository.save()runsCypherQueries.resolvePlaceholderTargets(EdgeType)for each ofCALLS/INCLUDES/USES_TYPE/EXTENDS/IMPLEMENTS, thenDELETE_RESOLVED_PLACEHOLDERS, in the same transaction as the node/edge merges. This redirects edges pointing at an unresolved placeholder (MODULE/DATA_STRUCTURE,sourceFile="") onto a real node sharing(type, name, project)once one exists, and deletes the now-orphaned placeholder — works regardless of ingest order.
Natural parser — statements, dataflow, dispatch
- P1-a. SQL/ADABAS statement extraction for agentic Panache generation —
new
NodeType.DB_ACCESSnode per SQL/ADABAS statement occurrence (table, access mode, raw statement text, line span),CONTAINSedge from the containing function/module,USES_TYPEedge to the accessedDB_TABLEand (forSELECT ... INTO VIEW) the associatedDATA_STRUCTURE. NewGET /api/projects/{project}/modules/{name}/sql-statementsendpoint + CLI command, so an agent can read table/view/mode/statement text and the call-graph context and write the equivalent Panache query itself. Existing function ->DB_TABLEREADS/WRITESedges and/db-accessesstay unchanged. - P1-b.
MOVE/ASSIGN/COMPUTEvariable read/write edges — Natural construct #7 (CLAUDE.md priority list). DeclaredVARIABLE/CONSTANTnames fromDEFINE DATAare tracked case-insensitively; forMOVE <src> TO <target>{...}and(COMPUTE|ASSIGN) <target> (=|:=) <expr>,READS/WRITESedges are added from the current function to known variables/constants referenced as source/target/expression operands. ExoticMOVEvariants (BY NAME/BY POSITION,EDITED,SUBSTRING(...),ALL,ENCODED,NORMALIZED,*JUSTIFIED) are skipped. No new placeholder nodes for unknown identifiers/literals. - P1-c. Bare
:=assignment and qualified-target field lookup (2026-06-15) —NaturalParserpreviously required aCOMPUTE/ASSIGNkeyword to recognize an assignment; Natural's implicit-assignment form<target> := <expr>(no keyword) is now matched via a newBARE_ASSIGNpattern (fallback afterASSIGN_COMPUTE, restricted to:=to avoid colliding withIF/comparison=).lookupVariablenow also strips a<QUALIFIER>.prefix (e.g.CDBRPDA.SORT-KEY->SORT-KEY) when the fully qualified name isn't a known local variable, so qualified field references resolve to the declared field name. New unit testsbareAssignWithoutKeywordIsRecognized/qualifiedAssignTargetMatchesUnqualifiedDeclaredField. NewIngestModuleITingests theWGEAGB0Sfixture tree (100 files). - P1-d.
ADD/SUBTRACT/MULTIPLY/DIVIDEread/write edges (2026-06-15) — Natural's arithmetic statements read and write their operands but were not recognized at all byNaturalParser. New patternsADD_STATEMENT,SUBTRACT_STATEMENT,MULTIPLY_STATEMENT,DIVIDE_STATEMENTaddREADSedges for every identifier in the operand-list/expression operands (via newaddExpressionReadshelper) andREADS/WRITESedges for the accumulator/target (addOperandAccesshelper):ADD ... TO x→ reads+writesx;ADD ... GIVING x→ writes onlyx;SUBTRACT ... FROM x [GIVING y];MULTIPLY x BY y [GIVING z];DIVIDE x INTO y [GIVING z] [REMAINDER r]. New unit tests covering all four statements and theirGIVING/no-GIVINGvariants. - P1-e.
lineNoonAstEdge(2026-06-15) —AstEdgegained anint lineNofield (the source line of the relationship/statement), threaded through both parsers'edge()helpers and all call sites. Persisted as a relationship property:buildMergeEdgeQueries()now doesMERGE (a)-[:%s {lineNo: e.lineNo}]->(b)(so the same edge type between the same nodes at different lines becomes distinct relationships), andbuildResolvePlaceholderTargetQueries()preservesr.lineNowhen redirecting placeholder edges. Exposed vialineNoinVariableAccessLocationandCallReference(callers/callees now return one row per call site). - P1-f.
DECIDE FOR/DECIDE ONasCONTROL_FLOW(2026-06-15) — per the construct-coverage review,DECIDE FOR/DECIDE ON(Natural's switch/case) appears in 8 of 26WGEAGB0Sfixtures but was not recognized. NewDECIDE_STATEMENT/END_DECIDEpatterns add aCONTROL_FLOWnode (dataType="DECIDE",value= the full statement text) spanning toEND-DECIDE, mirroringIF/FOR/REPEAT. Statements inside the block still get normalREADS/WRITESedges. New unit testdecideForIsRecognizedAsControlFlowBlock. - P1-g. Resolve qualified
PARAMETER/LOCAL USINGfield references (2026-06-15) —MOVE/ASSIGN/etc. targets of the formSTRUCT.FIELD(e.g.CDBRPDA.SORT-KEY) whereSTRUCTis aPARAMETER USING/LOCAL USINGinclude now resolve to a placeholderVARIABLEnodeFIELD,CONTAINS-child of the placeholderSTRUCTDATA_STRUCTURE. Previously such references either fell back to an unrelated same-named local variable or produced no edge at all. New Neo4j post-ingest step (resolvePlaceholderFieldTargetsforREADS/WRITES,DELETE_RESOLVED_PLACEHOLDER_FIELD_CONTAINS) redirects these placeholder fields onto the real field of the same name in the resolved structure. New unit testqualifiedFieldOfIncludedDataAreaResolvesToPlaceholderUnderThatStructure. Known limitation: two different included structures defining a same-named field converge (merge key has no parent reference); not an issue in current fixtures. - 18. Cross-file bare-field resolution (
MOVE/ASSIGN/etc.) (2026-06-17) — closes the remaining half of the P1-c/P1-g gap: an unqualified reference (MOVE #X TO SORT-KEYwhereSORT-KEYis declared only in aLOCAL/PARAMETER USINGarea).lookupVariablegained anallowIncludePlaceholderflag (true only at explicit MOVE/ASSIGN/COMPUTE/ arithmetic operand sites). When set and the bare name is not a same-file variable, is identifier-shaped, and the module hasUSINGincludes, a module-level placeholderVARIABLE(sourceFile="") is created. NewresolveBareIncludedFieldTargetsredirects it onto a real included field only on a unique single match acrossm-[:INCLUDES]->(:DATA_STRUCTURE{sourceFile<>""})- [:CONTAINS*1..]->; orphan cleanup reaps unresolved placeholders so no spurious edge survives. NewNaturalParserTestcases + ITbareReferenceToIncludedFieldResolvesToRealField. - 19. Field-level dataflow for shared PDAs (
field-flow) (2026-06-17) — a shared PDA's fields are single nodes, so a field written in the caller and read in a transitively-called module is already connected through that one field node plus theCALLSedge — no new edge type needed. NewGET /api/projects/{project}/variables/{field}/field-flow?module=X&depth=Ncorrelates writers and downstream readers of a shared field across the call graph (wmod <> rmod, reachable viaCALLS*1..depth, clamped 10), returning "field F is produced inMOD-A:120and consumed in downstreamMOD-B:45". NewFieldFlowrecord + CLIfield-flow. Caveat: reachability-based, not order-precise. FixturesFF_SHARED.lda+FF_PROD.natFF_CONS.nat; ITfieldFlowTracesSharedPdaFieldAcrossCall.
- 15. Dataflow analysis through the call hierarchy (2026-06-16, scoped) —
captures which variables are passed at each call site. Parser (Phase 1):
CALLNAT 'MOD' ARG1 ARG2captures the positional argument list onto theCALLSedgeargsproperty; top-levelPARAMETERfields tagged withparamPosition. Enricher (Phase 2):LINK_ARGS_TO_PARAMSruns after placeholder resolution, positionally joiningsplit(r.args,',')[i]to the callee param withparamPosition = toString(i)andMERGE-ing anARG_TO_PARAM {callSite, position}edge (newEdgeType.ARG_TO_PARAM). Query (Phase 3):GET /variables/{name}/flow-forward|flow-backward?module=&depth=followARG_TO_PARAM*1..depth, returning{variable, variableType, module, depth}(DataflowStep); CLIflow-forward/flow-backward. Placeholder-target resolution carries edge properties viaSET r2 += properties(r)soargssurvive redirection. Deferred:PERFORM USING, whole-areaPARAMETER USINGposition mapping, Java method-argument dataflow. FixturesDF_CALLER.natDF_CALLEE.nat.
Java parser
-
J6. Ingest noise: JDK/framework filter + target/ exclusion (2026-07-07) — a by-name ingest (
POST /ingest/{name}) follows every referenced type; JDK/stdlib/framework types never resolve to a project file, so they dominated theunresolvedlist and ballooned the dependency fan-out (one job pulled ~1158 files). NewExternalTypesdenylist (uppercased simple names:java.lang/util/time/io/nio/math/concurrent/stream- CDI/JPA/Panache/Mutiny types) is consulted in the dependency BFS — matching
MODULErefs are neither chased nor reported asunresolved, so only genuine gaps remain. The file walk now excludestarget/by default (merged with the project'sexcludeDirs) so build-output generated sources don't create duplicate module/entity definitions. Documented heuristic (a project class named like a JDK type would also be skipped — vanishingly rare). NewJavaIngestNoiseIT(JDK types filtered while a genuine missing type is still reported;target/copy not flagged as a duplicate); all 80 server ITs green.
- CDI/JPA/Panache/Mutiny types) is consulted in the dependency BFS — matching
-
J5. Cross-class Java dataflow (precise arg→param) (2026-07-07) —
flow-forward/flow-backwardacross classes. Investigation found the genericLINK_ARGS_TO_PARAMSlinker already matched Java cross-class (MODULE→MODULE) calls but imprecisely —(callee)-[:CONTAINS*1..]->(param at position i)linked an argument to the position-i parameter of every method in the callee class (no method disambiguation). J5 makes it precise:JavaParserstampscallerFn(calling method) andcalleeMethod(invoked method) on cross-classMETHOD_CALLedges; new dataflow-gated stepsLINK_ARGS_TO_PARAMS_JAVA_CROSS(+_SCOPED) buildARG_TO_PARAMfrom each caller-scope argument to the named callee method's parameter at that position; and the generic linker (+ its scoped variant) is guarded withr.calleeMethod IS NULLso it no longer cross-matches Java.flow-forward/flow-backwardtraverseARG_TO_PARAMunchanged, so they now span classes correctly (deep-ingest only, like Natural). Overloads still match by name (over-approx); constructor-argument dataflow deferred. New fixturesfixtures/java/flow/*(CheckoutFlow.checkout(amount)→PriceCalculator.applyDiscount(basePrice)) andJavaCrossClassFlowIT(forward + backward); all 78 server ITs + 11 parser tests green. -
J4. Virtual/override (template-method) dispatch (Java) (2026-07-07) — a template method in an abstract base (
doProcessItem→clearTable/writeEntities) dead-ended at the base's abstract method. NewEdgeTypeOVERRIDDEN_BY(base-classFUNCTION→ same-named overridingFUNCTIONin a subclass), materialized by the always-run, idempotent enrichment stepBUILD_OVERRIDDEN_BY(name-based, transitive overEXTENDS; constructors excluded naturally since a super/sub pair never shares a name). Exposed via a newGET /modules/{name}/functions/{fn}/overridesendpoint (GraphRepository.functionOverrides→FUNCTION_OVERRIDES, returningFunctionOverride{module, name, sourceFile, startLine, endLine}) and thefunction_overridesMCP tool — kept function-level rather than threaded into the module-levelcall-tree. Name-based matching is heuristic (overloads over-match); documented. New fixturesfixtures/java/override/*(abstract template base + two concrete steps) andJavaOverrideIT(2 tests: all overrides for an overridden method, empty for a non-overridden one); all 76 server ITs + 11 parser tests green. -
J3. Interface → implementation resolution (Java) (2026-07-07) — new
EdgeTypeIMPLEMENTED_BY(interface → concrete impl), materialized by the always-run, idempotent enrichment stepBUILD_IMPLEMENTED_BY(the derived inverse of a realIMPLEMENTSedge). Java MODULE nodes now carry anisInterfaceproperty (a small slice of the future module-kind item). New?resolveInterfaces=trueoption oncalleesandcall-tree(REST + the MCP tools):calleeshops an interface callee to itsIMPLEMENTED_BYimplementation(s) viacoalesce(impl, callee)— a single impl is a clean deterministic hop, multiple impls expand, the interface is dropped;call-treedrops interface nodes that have a known implementation (their impls are already reached via the CHA syntheticCALLSedges). Flag-gated so the default query is byte-unchanged. New fixturesfixtures/java/iface/*(single-implNotifier/EmailNotifier) reusing the multi-impl gateway fixtures;JavaInterfaceResolutionIT(3 tests); all 74 server ITs + 11 parser tests green. Note: the interface-traversal dead-end J3 originally described was already fixed by the earlier polymorphic CHA step — this item adds the explicit edge and the caller-facing hop/suppress option. -
J2. DI + class-literal wiring edges (Java) (2026-07-07) —
call-tree/calleeson a job class returned empty because steps are wired by CDI injection and class literals, not method calls. Two newEdgeTypes:INJECTS(a class → an injected bean type:@Injectfields, and constructor params of an injection-point constructor —@Inject-annotated, or the sole constructor of a CDI-scoped class) andREFERENCES(a class → a type used asX.classin argument position, e.g.super(AccountKeyInitStep.class, …)/batchlet(refName(X.class))). Emitted byJavaParser.addWiringEdgesas class-level (MODULE→MODULE) placeholder edges; added toRESOLVABLE_EDGE_TYPESso the existingresolve-placeholderfinalize step resolves them cross-file. Surfaced incallers/callees(query match widened toCALLS|EXTENDS|IMPLEMENTS|INJECTS|REFERENCES) asedgeKind = INJECTS/REFERENCES, mirroring howEXTENDS/IMPLEMENTSalready appear; deliberately excluded fromcall-treeby default (kept CALLS-only; see J9 below for the opt-in traversal). New fixturesfixtures/java/wiring/*(a job with a class-literal step +@Injectcollaborator, a constructor-injected bean) andJavaWiringIT(3 tests); all 68 prior ITs green (shared callers/callees query change verified non-regressive), 11 parser tests green. Re-test 2026-07-07 confirms this works: all 3 probed PUR job digests now list their steps viacallees.REFERENCES(RiskImportJob→RiskInitStep/RiskProcessingStep/RiskEndStep; same forMultiTableImportJob,KeyTableExportJob), and the step'scallers.REFERENCESnames the job. The former job-root dead-end is fixed at digest level. Follow-up usability gap fixed by J9 (below). -
J9. Opt-in traversal that follows
REFERENCES/INJECTS(Java) (2026-07-07) — J2 wired theINJECTS/REFERENCESedges intocallees/callersbut kept them out ofcall-tree(CALLS-only), so there was no automatic transitive tree from a job to its steps to their repositories — an agent had to chaindigest/calleescalls by hand. NewfollowWiringboolean, mirroring the J3resolveInterfacespattern end-to-end (RESTcall-treequery param, MCPcall_treetool arg,GraphRepository.callTree,CypherQueries.callTree): whentrue, the transitive-traversal relationship pattern becomesCALLS|INJECTS|REFERENCESinstead ofCALLS(bothINJECTS/REFERENCESareMODULE→MODULE, same as theCALLSedgescall-treealready traverses, so mixing them into one variable-length path is schema-compatible). Composes withresolveInterfaces.callees/callersunchanged (already surfaced these edges at one hop since J2) — onlycall-tree's transitive traversal needed the flag. Off by default. New testcallTreeFollowsWiringOnlyWhenRequestedinJavaWiringIT(default call-tree excludesINJECTS/REFERENCEStargets;followWiring=trueincludes them), reusing the existingfixtures/java/wiring/*fixtures. -
J1a. JPA/Panache repository calls as Java DB accesses (2026-07-07) — Java
db-accesses/sql-statementswere always empty; now repository/EntityManager/ Panache active-record calls resolve toREADS/WRITESon the entity'sDB_TABLE. Parser (JavaParser): tags a repository class with its managed entity (repositoryEntity, from aPanacheRepository<E>/JpaRepository<E,Id>/… generic supertype); treats aPanacheEntity[Base]subclass as an entity; and emits aDB_ACCESScandidate node (contained under the calling function,dataType=READ/WRITE/DELETEby method-name prefix,value= call text, plusjavaReceiverType/javaMethod/javaArgType) for a persistence-shaped call — a write/delete verb on any non-JDK receiver, or a read verb on a repository-named / static-entity /EntityManagerreceiver. A small JDK denylist keeps candidate volume down (J6 will generalize it). Enrichment (RESOLVE_JAVA_DB_ACCESS, a new always-run, project-wide, idempotent finalize step): derives the entity — argument type forem.persist/merge/remove(x), the repository'srepositoryEntity, else the receiver itself — follows the entity'sMAPS_TOto theDB_TABLE, thenMERGEsDB_ACCESS-[:USES_TYPE]->DB_TABLE(forsql-statements) andfn-[:READS|WRITES]->DB_TABLE(fordb-accesses). No newEdgeType: DELETE shows asWRITESindb-accesses(edge mode) butDELETEinsql-statements(node mode), mirroring Natural. Unresolved candidates (entity not in graph) stay unlinked — never pollutingdb-accesses, incremental-ingest-safe (no deletion). New fixturesfixtures/java/jpa/*(Panache repo + entity, EntityManager service, active-record entity) andJavaDbAccessIT(3 tests); all 68 server ITs + 11 parser tests green. J1b deferred (roadmap):@Query/JPQL/native-SQL string parsing, derived-name filters, and no-generic custom repositories (need symbol resolution). Re-test 2026-07-07 fixed by J7 + J8 (below):JavaDbAccessInheritedRepoIT(repo → project base → Panache base, plus a constant-valued@Entity(name=…)) now passes end-to-end —db-accesses/sql-statementsresolve toRISK. A regression check against the actual 9 PUR jobs is still worth doing as a follow-up acceptance pass, but the two traced root causes are fixed. -
J7. Panache-ness inherited through a project base class (2026-07-07) — J1a only recovered
repositoryEntitywhen a repository directly extended a Panache/JPA base type; a repository extending a project-specific abstract base (e.g.RiskRepository extends AbstractPurRepository<Risk, String>, whereAbstractPurRepository<Entity, Id> implements PanacheRepositoryBase<Entity, Id>) got nothing, since the entity generic sits one inheritance hop away from the Panache marker. Parser (JavaParser):repositoryEntityTypenow returnsnull(not a bogus name) when the repository base type's first type argument is one of its own type parameters rather than a concrete class; a newpanacheEntityTypeParamtags such a project base class with the name of that type parameter, alongside its own orderedtypeParams; a newtypeArgsedge property onEXTENDSrecords the concrete arguments a subclass supplies (e.g.Risk,String). Enrichment (RESOLVE_PANACHE_INHERITED_ENTITY, a new project-wide, idempotent finalize step run beforeresolve-java-db-access): finds the base class'spanacheEntityTypeParamposition in itstypeParams, reads the subclass'sEXTENDStypeArgsat that position, andSETsrepositoryEntityon the subclass — after whichRESOLVE_JAVA_DB_ACCESS(J1a, unchanged) resolves its repository calls exactly as for a directly-Panache repository. Resolves one level of indirection (the observed real-world shape). New fixturesfixtures/java/jpa/inherited/*andJavaDbAccessInheritedRepoIT(2 tests); fullac-code-serverunit + integration suite green. -
J8.
@Entity(name=CONST)/@Tablewith a constant table name (verified 2026-07-07, no change needed) — the physical table name is not always a string literal on the annotation;RiskEntity-shaped entities use a same-class constant reference instead (@Entity(name = Risk.TABLE_NAME)withpublic static final String TABLE_NAME = "RISK";). Turns outresolveTableNamealready resolved this via existing same-class constant-value resolution (collectConstants/resolveAnnotationString, predating J1a) — confirmed by parsing theJavaDbAccessInheritedRepoITfixture in isolation before any code change. No fix required; J7 (above) was the actual blocker for that fixture'sdb-accesses. -
12. Java parser — constructors, parameters, field READS/WRITES (2026-06-16) — constructors →
FUNCTIONnodes named after the class; method/constructor parameters →VARIABLEnodes (CONTAINSfrom the function); field READS/WRITES fromthis.fieldand unshadowed bare-name references (AssignExprtarget → WRITES; compound-assign /++/--→ READS+WRITES; else READS). Parameter/local names shadow field references. Intra-classCALLSnow also scans constructor bodies./variables/{name}/reads|writesgainedFIELDto the matched node types. Known limitations (deferred to item 15): same-named params in one file and overloaded constructors converge on one node. New fixturesOrderService.java+BaseService.java. -
16. Java parser — constant value resolution (2026-06-16) —
JavaParser.literalValuedoes type-aware extraction (StringLiteralExpr.asString(),Integer/LongLiteralExpr.asNumber()— stripsL,Boolean/CharLiteralExpr); non-literal initializers fall through to anullvalue.collectConstantsresolves in-class references (bareNAMEandThisClass.NAME) transitively with cycle guard.resolveAnnotationStringresolves an annotation member to a string via the same-class constants (used by item 13 for@Entity(name = TABLE_NAME)). Deferred: cross-class constant resolution /ConstantRefEnricher. -
13. Java parser — Hibernate/JPA entity recognition (2026-06-16) — detects
@Entity/@Table(table name resolved through item 16, fallback to class name),@MappedSuperclass(noMAPS_TO), and per-@Columnfield metadata (columnName,columnDefinition,nullable,converterTypefrom@Convert,hibernateTypefrom@Type,isId,declaredIn).AstNodegained an optionalMap<String,String> properties(persisted viaSET node += n.properties). NewEdgeType.MAPS_TO. Inherited columns are collected at query time by walking(:MODULE)-[:EXTENDS*0..]->(:MODULE)-[:CONTAINS]->(:FIELD). NewEntityColumnrecord +ENTITY_COLUMNSquery +GET /api/projects/{project}/modules/{name}/columns+ CLIentity-columns;DB_TABLE_COLUMNSgained a 3rd UNION branch. FixturesSampleLegacyEntity+AbstractSampleHistorized+AbstractSampleBase. -
14. Java parser — general inheritance enrichment (all classes) (2026-06-16) — implemented query-time (no new edge types). New
GET /api/projects/{project}/modules/{name}/functions?includeInherited=endpoint + CLIfunctions [--include-inherited]: walks(:MODULE)-[:EXTENDS|IMPLEMENTS*0..]->(:MODULE)-[:CONTAINS]->(:FUNCTION)(MODULE_FUNCTIONS_INHERITED) when inherited is requested, else own only (MODULE_FUNCTIONS_OWN); each row tagged withdeclaredIn. NewInheritedFunctionrecord. Deferred: explicitINHERITS/OVERRIDESedge types. FixturesOrderService.java(overridesdescribe()) +BaseService.java. -
17. Java dataflow (extends item 15 to Java) (2026-06-16) —
JavaParsertags each method/constructor parameter VARIABLE node withparamPositionand captures the positional argument list of each intra-classMethodCallExpronto theCALLSedgeargsproperty. NewLINK_ARGS_TO_PARAMS_JAVAenricher mapsargs[i]to the callee function's parameter withparamPosition = i, where the caller-side variable is the caller function's own parameter or a field of its enclosing class. Theflow-forward/flow-backwardendpoints are language-agnostic. Deferred: cross-class Java calls (resolved in item 20). -
20. Cross-class Java call graph (2026-06-17) — Java
CALLSedges were intra-class only. Now resolved at the class (MODULE) level, like NaturalCALLNAT:JavaParserresolves a method-call receiver to a target class — a typed field/parameter/local (svc.method()), a capitalized bare name (staticFoo.bar()), or anew Foo(...)constructor — and emits aMODULE-[:CALLS]->MODULEedge to a placeholder for that class. The placeholder resolves viaresolvePlaceholderTargets(CALLS)and surfaces as aDependencyRef. Scope: class-level resolution only. FixtureOrderController.java. -
22. Polymorphic Java call resolution + implementation ingest (2026-06-18) — (1) Fan-out —
LINK_CALLS_TO_IMPLEMENTATIONS, run last in thesavepost-processing: for every(MODULE)-[:CALLS]->(base)wherebasehas incomingIMPLEMENTS/EXTENDS, itMERGEs aCALLSedge from the caller to each subtype reachable via anIMPLEMENTS|EXTENDS*1..chain (Class-Hierarchy- Analysis over-approximation). Synthetic edges carryresolvedVia:'INHERITANCE'; dataflow is intentionally not routed through them. (2) Implementation ingest —ingestModulebuilds a reverse index (buildImplementorIndex) so when the BFS ingests an interface/base, its implementations are enqueued too. Fixturesfixtures/java/gateway/*; ITjavaInterfaceCallsFanOutToImplementations+IngestModuleJavaIT. -
J1b. Parse
@Query/JPQL/native-SQL strings + derived-name filters (Java) (found 2026-07-06, done 2026-07-07) — motivated by a batch-job analysis pass over the PURpur-batchmodule (2026-07-06, 9 concreteAbstractPurBatchJobsubclasses). J1a resolves calls whose entity is syntactically recoverable (repository generic param,em.persist/merge/remove(arg), Panache active-record receiver). Added:@QueryJPQL/native-SQL string parsing (leading DML verb → mode,FROM/UPDATE/INTOtarget → entity/table, attached at the method declaration since abstract repository methods have no call site to scan); derived query-method name parsing (findAllByClientAndStatus→derivedFilterproperty, tagged on the existing call-site candidate); and a no-generic-entity fallback (repository interfaces with no resolvable generic type anywhere, e.g. a customIRiskRepository, guess the entity from a declared method's return type). Seex-docs/agent-api-usage-ac-implementation.mdfor semantics/limits.
Project model & ingest endpoints
- 21. Per-project root folder + server-side scan ingest (2026-06-17) —
every project now carries a required
root(server-side path its sources live under) and optionalexcludeDirs, stored on theProjectnode. The server walks the root and parses files: newProjectIngestService+SourceFileshelper, withAstIngestServicesplit intoparse/save. Two new endpoints (@Blocking):POST .../ingest-all(scan the whole root) andPOST .../ingest/{name}(ingest one module + its transitive deps via a server-side BFS). Both return anIngestSummary({ingested, unresolved, duplicates, failed}). Strict duplicate detection (Java FQN; Natural(name, kind)from extension).sourceFilestored relative to the root. Old content-upload path removed. New CLI:project create,project update,project clear,ingest <name>,ingest all. - 12.
GET /api/projects/{project}/modules— list modules in a project (2026-06-15) — newCypherQueries.LIST_MODULES/GraphRepository.listModules/AnalysisResource.modules, returnsMODULEnodes (name,sourceFile), excluding unresolved placeholders. Optional?sourceFile=...filters to the module(s) defined in that file. Also addedtype=<NodeType>filter tosearch/identifier(400 INVALID_TYPEfor an unrecognized type). New IT tests.
Module analysis endpoints (P2 reengineering support)
-
72. A dispatch row reports its whole guard chain, not just the innermost
DECIDE(2026-07-17) — found by a manualVMULTMN4audit against the Natural source, the second such audit to pay for itself.DispatchEntry's contract is "guardField = guardValueroutes toassignedField := assignedValue" — a sufficient condition.NaturalParser.guardPropswalked the control-flow stack and returned on the first (innermost) value-DECIDEwith an active branch (its javadoc said so outright), so for a nestedDECIDEthe outer guard was discarded and a conjunction was served as one condition:128| DECIDE ON FIRST VALUE OF #SHORT-VIEW 165| VALUE 'TABL' 170| DECIDE ON FIRST VALUE OF #FIELD-NAME 171| VALUE 'TX-TABLA' 174| MOVE P-DESCRIPTION TO YTABLMA0.TX-TABLA reported: #FIELD-NAME = 'TX-TABLA' truth: #SHORT-VIEW = 'TABL' AND #FIELD-NAME = 'TX-TABLA'12 of that module's 69 rows (17 %) carried an incomplete condition. It over-generalises, which is the dangerous direction: this table exists to port
DECIDEs to Java, and a condition that is too weak ports into a branch firing where the Natural code never would. Corpus scale, measured on disk: 824 of 7,729 value-DECIDEblocks (10.7 %) are nested, across 178 modules (YFRAMN10.nat: 199). Additive, not breaking (correcting the roadmap's own note):whenField/whenValue/whenValueskeep their exact meaning — the innermost guard — so existing consumers see what they always saw; the chain travels in newwhenChainFields/whenChainValuesand surfaces asDispatchEntry.guards([{field, values}], outermost first, AND-joined; each link'svaluesOR-joined). Encoding needs two levels, so RS (U+001E) separates links above item 64's US (U+001F) for alternatives — neither can occur in Natural source. Decoding is in Java, not Cypher: zipping two nested splits back together needs index arithmetic that Cypher makes unreadable. Theac-uimigration dossier rendered the worst possible form —guardField+ the lossyguardValue(item 64 already said "preferguardValues"), i.e. the innermost guard in its ambiguous encoding, in the one screen meant for porting. It now renders the full conjunction. Live proof after refresh:VMULTMN4:174→#SHORT-VIEW = TABL AND #FIELD-NAME = TX-TABLA, legacy fields unchanged. Vacuity-checked: revertingguardPropsfails 3 of the 4 tests (guardscomes back[]) while the legacy-field test stays green — which is the additive claim, proven rather than asserted. The chain's order is asserted explicitly:ArrayDequeiterates innermost-first, so a missing reversal would silently yield an inside-out chain that a "contains both" assertion would accept. Left behind: #73 (aNONEbranch's condition is a negation no chain can express —guardsis strictly better, not total) and #74 (a duplication in the graph that item 72 exposed, not caused). -
67.
call-tree'sdepthcounts module hops, and says when it truncates (2026-07-17) — the last of the raw-hop siblings.callTreebounded on rawCALLSedges and returnedmin(length(p))as thedepthcolumn, so the number measured how deeply the calling statement happened to be nested rather than dependency distance. Measured live onupms:WGEAGB0S's seven direct module dependencies came back at depths 1..3 (raw path lengths 3,4,5,6,7,7,8) —BGEAGFN0reported 3 — although all seven are one module hop away. All seven now report 1. Depth comes from the path, not from ownership:min(size([n IN nodes(p) WHERE n.type='MODULE']) - 1). The decided semantics were "aFUNCTIONinherits the hop count of the module that owns it", but measurement killed the literal reading: a subroutine defined in a copycode is one shared nodeCONTAINSed by every including module —L4N-ENTER(L4NCOPY.cpy) has 137 owners, and 156 upms functions have more than one. "Its module" has no single answer. Counting module nodes on the traversed path yields one answer per route, sidesteps ownership entirely, and gives the intended result (a root's own subroutines are depth 0). The bound was the hard part, and the first design was wrong. A raw*1..Ncannot express a module-hop bound. The initial plan — a raw budget ofdepth × (1 + budget)— was killed in validation: atdepth=1a budget of 7 means 8 raw hops, and 8 raw hops reach 8 module levels when internal chains are short, so adepth=1request would have expanded ~8 module levels and discarded the rest (costing whatdepth=8costs today). Worse, it could report a wrong depth: a target reachable at module-depth 1 via a long internal chain and at module-depth 5 via a short one would be found only on the short route. Fix:moduleTree(item 65) supplies the modules withinmaxDepthmodule hops, and a QPP with a per-hop predicate prunes the traversal to them while expanding, so it can never wander beyond the requested depth. Verified against liveupmsbefore writing any code (QPP had already refuted one of my designs in item 65).truncated(new response field). The raw quantifier survives as a pure safety cap (agenticcode.call-tree.internal-budget, default 20). Because pruning bounds the search space, the quantifier is free — measured onupms, quantifiers 8/20/40 gave identical results at 2.0/2.6/2.1 s — so it is sized far above the observed maximum. Sizing it from a 300-module sample would have been wrong: that sample said "internal chains ≤ 6", butWGEAGB0Salone reachesISINDATEat 8 raw hops for one module hop, and the true longest path is 10. Since any cap can cut, exhausting it now setstruncated: truerather than returning a short list that reads as complete — the exact failure mode that made items 65 and 68 bugs instead of documented limits. Conservative:trueis possible for a complete result. A "known imprecision" I filed here as #71 was retracted 2026-07-17 — it was my measurement error, not a defect. The claim was that Natural external subroutines (module A performs a subroutine owned by module B) do not incrementdepth, "4,808 such calls inupms". That count applied the single-owner filter to the target but not the source, so every copycode-shared source function (L4N-ENTER: 137 owning modules) was counted once per owner. Source and target single-owner returns 0 — and it must, becauseresolvePlaceholderTargetsnever resolves aFUNCTIONplaceholder across modules, so the cross-moduleFUNCTIONedge the claim assumed cannot exist. Path-based depth has no gap here. Two regressions the existing suite caught, both mine.followWiring(item J9) returned an empty list: the bounding BFS followed onlyCALLS, so withCALLS|INJECTS|REFERENCEStraversal every wiring-only neighbour fell outside the module set and the predicate pruned it away — a bound must follow the same edge types as the traversal it bounds (MODULE_HOP_OUT_WIRING). AndAnalysisResourceITencoded the old semantics; checking the fixture instead of just relaxing the assertion showed the old expectation was the bug:INITIALIZATIONSandR-ADDRESS_SPare both subroutines ofYADDRBN0_SAMPLE.natitself, yetdepth=1omittedR-ADDRESS_SP— aDEFINE SUBROUTINEof the very module being asked about, hidden because it sat two raw edges away. Consequence to expect: a givendepthnow returns more, since the internals of modules within the bound are inside it. Both halves vacuity-checked: reverting the formula madeDEPTHLEAFvanish from adepth=1answer entirely (expected: <1> but was: null), and the truncation test is paired with an assertion thattruncatedis false on a normal query, so an always-true flag could not fake it. Dead code removed on the way: four unusedcallTreeoverloads (CypherQueries×2,GraphRepository×2). -
68.
field-flowbounds on module hops, via a derivedCALLS_MODULEedge (2026-07-17) — the last of the three raw-hop siblings item 65 uncovered.fieldFlowgated producer→consumer pairs onEXISTS { (wmod)-[:CALLS*1..%d]->(rmod) }, so itsdepthcounted a module's internalPERFORMjumps: a consumer called from two subroutines deep sat 3 raw edges away and fell outsidedepth=1. The endpoint then reported nothing downstream consumes this field — a confident false negative, the same failure mode as item 65's "no DB access". Unlike item 65 this is a pairwise reachability test between two arbitrary modules, and the producer is optional ($module IS NULL), so item 65's per-rootmoduleTreeBFS does not apply — it would have to run once per candidate producer. Fix: an enricher materialises(MODULE)-[:CALLS_MODULE]->(MODULE)(the persisted module-level projection of the call graph, same shape asEGO_NEIGHBORS_OUT), andfieldFlowbounds on that, which makes a plain variable-length bound correct again. (Correction: this was expected to supply item 67's bound too. It did not — item 67 turned out to need a set of module names to prune a QPP with, which is what item 65'smoduleTreeBFS already returns, soCALLS_MODULEis used byfield-flowalone. It earns its keep there:field-flowis a pairwise test between two arbitrary modules, where a per-root BFS does not apply.)DELETEbefore rebuild is the load-bearing part, notMERGE. The edge is derived, andMERGEis idempotent only for edges that still exist — it never removes ones that shouldn't. The item-58 sweep deletes stale nodes only, so a surviving module's relationships are never reaped. Proven by removing theDELETEstep: after deleting theCALLNATfrom the producer's source and refreshing,field-flowstill answered[STALECONS]— a dependency that no longer exists in the code. Ordering matters too: the projection runs last, after every step that addsCALLSedges (placeholder + dynamic-CALLNATresolution,link-calls-to-implementations) or removes them (the item-62 data-literal reaping), since a projection is only as correct as the graph at the moment it runs. Both fixes vacuity-checked: reverting the bound made the field-flow test fail with an empty producer list. -
69. The origin file is part of an edge's identity (2026-07-16) — data loss, found by item 66's own regression fixture rather than by review. An edge was identified as
MERGE (a)-[r:%s {lineNo: e.lineNo}]->(b)— (source, target, type, lineNo), with no file. Copycode expansion (item 46a) made that key ambiguous: a host statement on line 10 and a copycode statement on line 10 share it entirely, so the secondSET r += e.propertiesoverwrote the first. The endpoint then returned one row where two statements exist, and a real write was gone from the graph. Proven by reverting the fix: withMOVE 'Y' TO #KEYonCOPYHOST.nat:10andMOVE 'Z' TO &1&onMYTABLECOPY.cpy:10,variables/#KEY/writesreturned only[{sourceFile=MYTABLECOPY.cpy, lineNo=11, ...}, {sourceFile=MYTABLECOPY.cpy, lineNo=10, ...}]— the host's own write had vanished, and the surviving edge even claimedviaCopycode=MYTABLECOPY. Fix:MERGE (a)-[r:%s {lineNo: e.lineNo, originFile: coalesce(e.properties.originFile, a.sourceFile)}]->(b). The fallback isa.sourceFile, not''(as first sketched in the roadmap): for a host statement the edge really does originate in its source node's file, sor.originFileholds the true file on every discriminated edge and item 66'scoalesce(r.originFile, <src>.sourceFile)keeps returning exactly what it did before. An''default would have broken all four of those queries, sincecoalescereplaces onlynull. Scope was 20 MERGE sites, not the 1 the roadmap named: the batch merge plus 7 placeholder-resolution re-merges (which re-key the edge) and 6 dynamic-CALLNATmerges. In the dynamic ones the origin had to be threaded explicitly through theWITH DISTINCT caller, dyn, lit, which drops thesrccall-site node. Deliberately excluded:CONTAINS/INCLUDES/USES_TYPE(unique per node pair by construction — a discriminator cannot prevent a collision there and would add a string property to the most numerous edge type in the graph); the Java{lineNo: a.startLine}merges (copycode is Natural-only); andARG_TO_PARAM— its collisions are real but benign, since two colliding edges carry identical(callSite, position)and no per-site payload, so discriminating them would only add parallel edges for dataflow to walk. The type test is aswitch, not a staticSetfield: the query maps are static finals that call it from their initialisers, and a field declared after them is stillnullat that point (this bit — compiled clean, failed at runtime). Corpus incidence remains unmeasurable after the fact: the collision destroys the evidence one would count. Existing graphs are not migrated — the stale sweep deletes only nodes, so old edges would linger beside the new key; delete + re-ingest the project (~179 s forupms). -
70. Placeholder field resolution kept the copycode provenance (2026-07-16) — found while fixing item 69, by reading the queries around it. Four of the six resolution queries copied only two properties when redirecting an edge onto its real target (
SET r2.value = r.value; SET r2.lineNo = r.lineNo), while the other two usedSET r2 += properties(r). The four discardedviaCopycode,includedAtandoriginFile— the very provenance item 66 had just added. This made item 66 only half-effective, which itsupmsverification could not reveal: a bare field reference (#W-OPTIONS) resolves throughresolveBareIncludedFieldTargets, which copies everything — and that is the field item 66 was verified on. A qualified reference (MYLDA.Q-FIELD, a field of aLOCAL USINGdata area) takesresolvePlaceholderFieldTargetsinstead and silently came back withviaCopycode: null, i.e. indistinguishable from a host statement. The split is inNaturalParser.lookupVariable: bare names go toscope.bareFields(), qualified ones toscope.placeholderFields()under a placeholderDATA_STRUCTURE. Fix:SET r2 += properties(r)in all four. Proven by reverting it — the new fixture returnedexpected: <[QUALCOPY]> but was: <[null]>. Note the two fixes are complementary: item 69's key incidentally carriesoriginFilethrough resolution, butviaCopycode/includedAtneed the+=. -
66. Call sites and variable accesses name the file their line is in (2026-07-16) — found by the same manual
WGEAGB0Saudit as item 65. The graph was already right; the API threw the answer away. The parser stamps copycode-origin edges (item 46a) withviaCopycode+includedAtand gives copycode-origin nodes the real.cpyfile — exactly what this doc promised — butviaCopycodeappeared nowhere inac-neo4j-storeorac-code-server, so a line belonging to a.cpywas served as if it were a host line:before: #W-OPTIONS writes → sourceFile WGEAGB0S.nat, lineNo 18/20/22 truth: WGEAGB0S.nat:18/20/22 = the generated comment banner ("* System : Versis", ...) ISICINDI.cpy:18/20/22 = &1& := '0' / '1' / ' ' (&1& = '#W-OPTIONS', INCLUDE at host line 676)Host and copycode lines sat indistinguishably in one list (
lineNo: 670in the same response was a real host line). Severity differed per endpoint:variables/{name}/reads|writesmade an explicit false claim (sourceFile= host +lineNo= copycode line);callers/calleesgave unattributedlineNos(theirsourceFileIndexis the callee's file, so they never claimed a call-site file). Scale inupms: 11,008CALLS, 1,325WRITES, 996READSedges carryviaCopycode. Root cause was one missing line:CopycodePreprocessor.withViaalready heldorigin.sourceFile()where it stampedviaCopycode, butAstEdgehas nosourceFilefield, so the path was dropped — edges now carry it asoriginFile. On top of that:AggregatedCallRef.lineNos: [int]→sites: [CallSite](lineNo,callSiteFileIndex,viaCopycode,includedAt) — per site, because one target can be called from both the host and a copycode, which entry-level provenance cannot express. Deliberately namedcallSiteFileIndex, notsourceFileIndex: at entry level that word means the callee's file, and reusing it would have baked in the next misreading.variables/{name}/reads|writesreturncoalesce(r.originFile, f.sourceFile)plusviaCopycode/includedAt.payloadalready did this correctly (per-entrysourceFile+lineNo) and was the model. Breaking change: REST/MCP/CLI stayed in sync for free (MCP returns the record, the CLI prints the raw JSON);ac-ui'sIdentifierPopoverwas carrying the same confusion — keying PERFORM lines to the caller's definition file — and now reads each site's own file. Live:ADLML02inWGEAGB0S's callees →lineNo 26, callSiteFile src/manual/copycode/ISIYESNO.cpy, viaCopycode ISIYESNO, includedAt 673, whileBGEAGFN0→lineNo 497, WGEAGB0S.nat, no copycode; the#W-OPTIONSwrites above now nameISICINDI.cpywithincludedAt 676, line 670 unchanged. Tests inCopycodeExpansionIT(call site + a variable written from both host and copycode), both vacuity-checked: withoutoriginFilethey fail withexpected: <MYTABLECOPY.cpy> but was: <COPYHOST.nat>— the defect verbatim. Building that fixture also turned up roadmap item 69 (two edges on the same line number from different files collapse in the MERGE). -
65.
?depth=counts module hops, not rawCALLSedges (2026-07-16) — found by a manual endpoint audit ofWGEAGB0S(upms) against the Natural source.db-accesses?depth=1..4reported zero DB accesses for a module that reachesYGEAGBNH's tables through exactly two module calls. An empty list reads as "this module touches no database" — a confident false negative, the worst failure shape for an agent. Cause: aCALLSedge starts at the statement that makes the call, not at the enclosingMODULEnode, so(m)-[:CALLS*1..N]->(hop:MODULE)also steps through internalPERFORMjumps. The real path isMODULE:WGEAGB0S → FUNCTION:GET-DATA → FUNCTION:GET-MAIN-DATA → MODULE:BGEAGFN0 → FUNCTION:READ-FILE → MODULE:YGEAGBNH— 5 raw edges for 2 module hops, so the answer only appeared atdepth=5, and the required value depends unpredictably on the callee's subroutine nesting.EGO_NEIGHBORS_OUT(item 49) already defined a module hop correctly as(a)-[:CONTAINS*0..]->()-[:CALLS]->(b:MODULE)and BFS'd one hop at a time — soego_graph?depth=2anddb-accesses?depth=2disagreed about the same graph. Fixed by giving the transitive endpoints the same definition: newMODULE_HOP_OUT(batched over a whole BFS frontier — one query per hop, not per module) +GraphRepository.moduleTree, consumed byDB_ACCESSES_FOR_MODULES,SQL_STATEMENTS_FOR_MODULESand the?module=scope ofvariables/{name}/reads|writes(whoseEXISTS { (root)-[:CALLS*1..N]->(owner) }had the same flaw). Done in Java because Cypher cannot express the hop: a quantified path pattern rejects both the variable-lengthCONTAINS*0..inside it ("Variable length relationships cannot be part of a quantified path pattern") and nesting. Live:WGEAGB0Sdb-accessesnowdepth=1 → 0(correct — none of its 7 direct callees touches a table),depth=2 → 7incl.VDB2-VERSIS_GENAGREE via YGEAGBNHat line 2617 (matches the source'sFIND (1) VDB2-VERSIS_GENAGREE), consistent withego_graph. Regression testModuleHopDepthIT(fixturesDEPTHROOT/DEPTHLEAF: one module hop, but theCALLNATtwoPERFORMlevels deep); vacuity-checked — dropping theCONTAINS*0..step fails 2 of its 3 tests, while thedepth=0test correctly stays green. -
P2-a. Control-flow parsing (
IF/ELSE/END-IF,FOR/END-FOR,REPEAT/END-REPEAT) — Natural construct #6. NewNodeType.CONTROL_FLOWnode per block (dataType=IF/FOR/REPEAT,value= condition/loop-spec text, line span = full block incl. matchingEND-*), nested viaCONTAINSedges. -
P2-b. Module "context bundle" endpoint (
GET /modules/{name}/context) — aggregates functions, callers/callees, db-accesses/sql-statements, and a variable read/write summary for a module in a single response. NewGraphRepository.moduleContext()+ CLIcontextcommand. -
P2-c.
DATA_STRUCTUREfield schema — newGET /api/projects/{project}/data-structures/{name}/fieldsreturns the flattened field list (name, type, dataType/length, const value, immediate parent) of aDEFINE DATA/DDM structure. New CLIdata-structure-fields. -
P2-d.
DB_TABLEcolumn schema — newGET /api/projects/{project}/db-tables/{name}/columnsderives column info for aDB_TABLEfrom theDATA_STRUCTUREused as theINTO VIEWtarget ofSELECTs against that table. New CLIdb-table-columns. -
P2-e. Transitive call graph — new
GET /modules/{name}/call-tree?depth=Nendpoint returns all modules/functions transitively reachable viaCALLSedges from the module (each with its minimum hop-depth).depthdefaults to 3, clamped to 10. New CLIcall-tree --depth.
Correctness gaps (wrong/empty results)
-
77. Bare-field resolution ignored which module it was resolving for (2026-07-17) — found while root-causing item 74, which it does not fix (see the note there).
resolveBareIncludedFieldTargetsdocumented itself as redirecting a bare reference "only when exactly one field of that name is reachable through the module's resolvedINCLUDES", but resolved project-wide across all owners at once. An unresolved bare field is one node per(name, project)(item 76 —MERGE_NODESkeys onsourceFile,""here), so(m)-[:CONTAINS]->(ph)matched it once per owning module;WITH ph, collect(DISTINCT realv)then droppedm, and the redirect'sMATCH (src)-[r]->(ph)never boundsrctom. Two bugs in opposite directions, both measured inupms:- under-resolution — two owners disagreeing on the target made
size(matches) = 1fail for everyone, including owners for whom the name was unambiguous (38 placeholders); - misattribution — an owner with no matching include had its edges redirected onto another module's
field anyway (28 placeholders, 199 module-field pairs). Live example:
BMTABBP0was recorded writing##MSG-NRofCDPDA-M.pda, a data area it does not include, while the other 63 modules touching that node genuinely include it.
Both surface as a fabricated dataflow:
field-flowpairs producer and consumer only when both touch the same node, so two modules sharing nothing but a field name were reported as passing data between them — a dependency an agent would act on.variables/{name}/reads|writescannot see any of it (it matches every node with the name and reports only the accessing side), which is why the defect survived this long.Fixed by carrying
minto the aggregation and bindingsrcwith(m)-[:CONTAINS*0..1]->(src). The*0..1bound is exact, not a guess: a placeholder edge'ssrcis only ever theMODULE(35,572 edges) or aFUNCTIONdirectly under it (157,616; all 16,231 such functions are direct children) — never aCONTROL_FLOWnode, so the walk also stays clear of item 75'sCONTAINScycles. Also aligned the scoped and unscoped twins, which had silently disagreed (*1..10vs unbounded → the same module could resolve differently depending on which pass ran); the shared bound is nowINCLUDE_FIELD_DEPTH, and it is required, not an optimisation, becauseCONTAINSis not acyclic.DELETE_RESOLVED_BARE_PLACEHOLDER_CONTAINSbecame per-module in the same change — its global "no edges left" test would otherwise have kept a resolved module'sCONTAINSalive on the strength of another module's unresolved edges.Cost: the resolver is ~4.4× slower on
upms(34 s → 2 m 30 s forWRITES), paid on deep ingest only. Known residue (item 76): 132 of 18,539 source nodes are copycodeFUNCTIONs shared by several modules; their edge to the placeholder is a single edge, so per-module binding cannot separate what is one node. Needs per-module placeholder identity, not a better query. Verified byBareFieldModuleScopeIT(2 of 3 tests fail without the fix; the third is a control proving resolution still works, since resolution failing everywhere would also yield no flow). - under-resolution — two owners disagreeing on the target made
-
75. Dynamic-
CALLNATresolvers wedged on the cyclicCONTAINSgraph (code + ITs done 2026-07-17; corpus re-verify pending). A whole-rootdeeprefresh ofupmshung finalize step 17 (resolve-dynamic-callnat-intra-indirect) for ~2 h without completing, blocking every later step (including item 77's bare-field resolution — so item 77 could not be corpus-verified). Cause: that step joins three unbounded(caller:MODULE)-[:CONTAINS*0..]->anchors, and per-caller path enumeration over item 75's cyclic copycode containment is cubic (a read-only probe of the exact query for the single callerDAGNTFN0did not finish in 60 s). Fix: the sixRESOLVE_DYNAMIC_CALLNAT_*queries (INTRA/INTRA_INDIRECT/CROSS+_SCOPEDtwins) find a module's own statements bysourceFileequality (aMODULEis 1:1 with itssourceFile; itsCALLNATsites andWRITESstatements share it) instead of descendingCONTAINS— an index hash-join that cannot cycle. The cross-module resolver's dispatch variable may live in an included PDA, so it is scoped to the caller's own file or a data structure the callerINCLUDES(name-only match would pull 20958 unrelated same-named vars → the scope filter keeps 307). Proven equivalent onupms:resolve-dynamic-callnat-intrareturns the identical 31(caller, target, lineNo)triples project-wide; the rewritten indirect step runs project-wide in ~6 s instead of wedging. Existing dynamic-dispatch ITs (intra / indirect / cross / scoped / unresolved- survival, 8 tests) stay green. Does not remove the underlyingCONTAINScycles (still open inroadmap.md#75); it removes this family's dependence on traversing them. Corpus-verified onupmsv68 (2026-07-18): 251 resolved dynamicCALLS(107 indirect), cross-module 113 callers / 165 edges. -
75b.
link-args-to-params(dataflow step 27) rewritten offCONTAINS*(2026-07-18) — the same item-75 blow-up: three unboundedCONTAINS*anchors (src,cv,pv) over the cyclic copycode graph made it the single slowest finalize step onupms(~25 min). All three now resolve through the(project, sourceFile, …)index instead of descent: the call sitesrcbysourceFileequality (module ↔ file is 1:1), and the argument variablecv/ callee parameterpvover the file-set {module's own file} ∪ {files itINCLUDES}, withcvresolved beforepvis unwound so the file-lists are not cross-producted. Corpus-verified:ARG_TO_PARAMstill built (436 edges on the v68 recreate) and the step no longer registers as a running query across 45 s poll windows (≈ seconds, not 25 min). After this, a full deepupmsfinalize (~28 min) is dominated entirely by step 19 (resolve-field- placeholder WRITES), which is an item-76 cardinality problem (shared placeholder nodes), notCONTAINS*— tracked separately. -
76. Per-module identity for field placeholders (2026-07-18) — an unresolved field (
VARIABLE/CONSTANT,sourceFile="") used to be ONE shared node per(name, project), referenced by up to 916 modules. That sharing maderesolve-field-placeholder WRITES(finalize step 19) join(ph)-[:CONTAINS]->(phv)×(src)-[:WRITES]->(phv)into a ~28.7M-row cartesian → ~55 min onupms. Fix:GraphRepository.placeholderOwnerstamps anownerModule(the referencing file) onto field placeholders only; it joins both the in-memory dedup key (mergeKey) and the DB merge key (MERGE_NODES/MERGE_POSITIONAL_NODES).MODULE/DB_TABLE/DATA_STRUCTUREplaceholders keepownerModule=""and stay shared (aCALLNAT/USINGtarget still resolves once). A companion sweepDELETE_STALE_PLACEHOLDER_NODESreaps a re-ingested file's obsolete placeholders (the file sweep skipssourceFile=""). Corpus (v69): step 19 55 min → 55 s, whole deep finalize ~59 min → ~6 min, item-77 misattribution stays 0, graph grows only +2,207 nodes (resolved placeholders are deleted in step 26),ARG_TO_PARAMwell-formed (30,583 edges, 0 position mismatch). Field/dynamic ITs green. -
P1-h. Transitive DB access resolution (2026-06-16) —
/db-accessesand/sql-statementsreturned[]for modules whose SQL lives behind one or moreCALLNAThops (e.g.WGEAGB0S→BGEAGFN0→YGEAGBNH). Added?depth=Nparam; whendepth>0, newdbAccessesTransitive/sqlStatementsTransitivequeries walk up todepthCALLShops and annotate each result withvia(the intermediate module name). Clamped to 10. -
P1-i.
/db-tables/{name}/columnsis effectively broken (2026-06-16) — returned only one column for tables whose column schema is encoded in.pdafiles.parseDataAreanow detectsVIEW OF <table>and creates aUSES_TYPEedge from theDATA_STRUCTUREto aDB_TABLEplaceholder.DB_TABLE_COLUMNSupgraded to a UNION query covering both the PDA path and the SELECT INTO VIEW path. New unit testpdaViewOfCreatesUsesTypeEdgeToDbTable. -
P1-j. SQL statement text is truncated (2026-06-16) — every
sqlStatementsentry hadstatementcut off at...WHERE. TheDB_READhandler now reads ahead through subsequent lines until an empty line orEND-FIND/END-READ, accumulating all lines intoDB_ACCESS.value. New unit testmultiLineFindStatementTextIsCapturedFully. -
P1-k.
FIND (1)misparsed as a table named(1)(2026-06-16) —FIND (1) <view>produced a phantom DB-access entry. Fixed by adding(?:\\(\\d+\\)\\s+)?to theDB_READpattern. New unit testfindWithRecordLimitIsNotMisparsedAsTableName. -
P1-m. Dynamic
CALLNAT <var>calls are invisible in the call graph (2026-06-21) — the parser/enricher only recorded literalCALLNAT 'NAME'edges, so variable-target calls (CALLNAT #WIF ...) produced noCALLSedge. Real case:W-MNT-N0dispatches ~90Wxxxx*SXML-interface programs via a singleCALLNAT #WIF; the target is resolved byKDWWIFN0. Done — and v1 went further, doing both intra- and cross-module. Parser detectsCALLNAT <var>(newCallKind.CALLNAT_DYNAMIC), emits a variable-named placeholder + marker edge, captures multi-line argument lists;PARAMETER USINGadvancesparamPosition. Enrichment addsRESOLVE_DYNAMIC_CALLNAT_INTRAandRESOLVE_DYNAMIC_CALLNAT_CROSS(followsARG_TO_PARAMinto the callee — recovers theW-MNT-N0 → KDWWIFN0 → WGEAGB0Schain), plus scoped variants and placeholder cleanup. Required fixingLINK_ARGS_TO_PARAMSto source calls via(:MODULE)-[:CONTAINS*0..]->(src)-[:CALLS]so subroutine-nested CALLNATs match. Live-validated onupms: deep-ingestingW-MNT-N0(8.4 s) yieldscallers/WGEAGB0S=W-MNT-N0(edgeKind=CALLNAT_DYNAMIC) andcallees/W-MNT-N0= 124 resolved dynamic targets. -
P1-n. Data-structure fields lack
scopeand often have nulldataType(2026-06-21) — the parser now stamps everyDEFINE DATAfield/variable with ascopeproperty (PARAMETER|LOCAL|GLOBAL|INDEPENDENT), surfaced as ascopefield on bothDataStructureFieldandIdentifierMatch(/search/identifier) — the latter exposes the motivating inline param#P-CALLED-PROG. Covered byNaturalParserTest+AnalysisResourceIT. -
P1-o. No search by literal value (2026-06-21) — a program name like
WGEAGB0Scan exist in the graph only as a string literal assigned to a field. AddedGET /search/value?value={v}(SEARCH_BY_VALUE,ValueMatch): quote-insensitive, returning both nodes carrying the value as a constant (kind=NODE) and literal assignments to a variable (kind=ASSIGNMENT). Missingvalue→400 MISSING_VALUE. -
P1-p. Edge provenance on call results (2026-06-21) — done via the
edgeKind=CALLNAT_DYNAMICprovenance carrier: stamped on every inferred edge and surfaced through the existingedgeKind = coalesce(r.callKind, …)projection, so callers/callees distinguish dynamic (inferred) from literalCALLNAT/PERFORM(static) with zero DTO changes. -
P1-q.
/data-structures/{name}/fieldsover-returns across modules (field pollution) (2026-06-21) — for a structure name shared by many modules (e.g.W-WIF-A1, used by ~80Wxxxx0Ssubprograms), the endpoint returned a polluted list (269 rows spanning ~80 unrelated modules).DATA_STRUCTURE_FIELDSnow selects the canonical definition(s) — those with a non-emptysourceFile— falling back to empty-sourceFilestubs only when no real def was ingested, and constrains each field'sparentto lie within the chosen structure (p = s OR (s)-[:CONTAINS*1..]->(p)). Live:W-WIF-A1266→115 rows with 0 module-parented leak;YGEAGROW430→36. Covered bydataStructureFieldsAreScopedToCanonicalDefinition. -
P1-r. Module→data-structures endpoint (2026-06-21) — added
GET /modules/{name}/data-structures(MODULE_DATA_STRUCTURES,ModuleDataStructure) returning one row per referenced structure:name,relationship(USINGcopybook viaINCLUDES/INLINEgroup viaCONTAINS),area(PDA/LDA/GDA/INLINE/UNKNOWN),fieldCount,sourceFile. LiveWGEAGB0Snow exposes its whole interface (W-WIF-A1PDA/118,W-WIF-A2PDA/10,W-WIF-A4PDA/5, the LDAs) and flags unresolvedCDPDA-M/CDPDA-PasUNKNOWN/0. Note: precisePARAMETER/LOCAL/GLOBALUSING scope is not stored on theINCLUDESedge (areais the proxy). Covered bymoduleDataStructuresListsReferencedAreas. -
P1-s. Referenced PDAs/copybooks not fully ingested (empty field lists) (2026-06-21) —
WGEAGB0SdeclaresPARAMETER USING CDPDA-M, butGET /data-structures/CDPDA-M/fieldsreturned[]. Root cause: the file is ingested, butCDPDA-M.pda's sole top-level group is namedMSG-INFO, while a module references it by file/DDM name. Fix: when a data area has exactly one top-level group whose name differs from the file name,parseDataAreawraps it in a rootDATA_STRUCTUREnamed after the file. Multi-top-group areas and matching-name areas keep their existing shape. Covered byparsesPdaWithRedefineGroup+usingResolvesCopybookWhoseGroupNameDiffersFromFileName(fixturesMSGAREA.pda/MSGUSER.nat). Re-ingest required. -
P1-u. Surface a module
purpose/description(2026-06-21) — parsers stamp adescriptionproperty on theMODULEnode: Natural scans the leading comment banner (priority**SAG TITLE:→* Title :→**SAG DESCS(n):→* Function :); Java uses the class Javadoc's first line. Surfaced asModuleContext.description(/context) viaMODULE_SOURCE_FILE(fetched as aModuleHeader),nullwhen no banner. Covered byNaturalParserTest,JavaParserTest,AnalysisResourceIT. Re-ingest required. -
P1-x.
/callerssilently omits incomingEXTENDS/IMPLEMENTSedges (2026-06-21) — "who extends/implements this class?" returned nothing (live onpur,GET /modules/AbstractUPMFESvc/callerswas[]despite 22 controllers). Root cause:callers(scope)matched only[r:CALLS], whereascallees(scope)already traversed[r:CALLS|EXTENDS|IMPLEMENTS].callers(scope)now mirrorscallees: matchesCALLS|EXTENDS|IMPLEMENTSand tags inheritance callersedgeKind=EXTENDS/IMPLEMENTS.scope=internalstill excludes inheritance;scope=externalincludes it. Covered bycallersSurfaceIncomingInheritanceEdges. -
P1-y. Double-hash (
##) field names truncated to#(2026-06-22) — the data-area field regexDATA_AREA_FIELDaccepted only a single leading#, so a Natural variable named##MSGwas parsed withname = "#". WidenedDATA_AREA_FIELD's name class to[#A-Za-z][#\w-]*so a leading run of#is captured whole.LEVEL_FIELDandIDENTIFIER_TOKENalready handled##. Covered bydoubleHashFieldNamesAreNotTruncated. Re-ingest required for##names. -
P1-z. Dispatch table queryable —
DECIDE/IFvalue→assignment correlation (2026-06-22) — analyzing the routerKDWWIFN0could recover the set of 130 dispatched programs but not the mapping (#P-OBJECT-TYPE = 'genagree'→#P-CALLED-PROG := 'WGEAGB0S'). TheNaturalParsernow tracks the activeDECIDE ON [FIRST] VALUE OF <subject>branch: it records the subject per DECIDE block and the currentVALUE '<literal>'(handlingNONE/ANY VALUEresets and comma-separated value lists), and stamps every literal assignment made inside that branch — including ones nested in an innerIF— withwhenField/whenValueproperties on theWRITESedge. New endpointGET /modules/{name}/dispatch-table(DISPATCH_TABLE,DispatchEntry) returns rows{guardField, guardValue, assignedField, assignedValue, lineNo}. Covered bydecideValueAssignmentsCarryDispatchGuard+dispatchTableRecoversObjectTypeToProgramMapping(fixtureKDDISP.nat). Live-validated onupms(2026-06-22):GET /modules/KDWWIFN0/dispatch-tablereturned ~140 rows reproducing the hand-builtkdwwifn0-prog-routing.csvmapping (genagree → WGEAGB0S/WGEAGX0S, etc.). Re-ingest required. Known limitations: the innerIF #L-LIST(list-vs-detail) sub-guard is not separately captured (both branch assignments share the DECIDE'swhenValue); the field-placeholder resolution does not copy edge properties, so a guarded write to a copybook field would lose its guard on redirect.
Token efficiency (payload shape)
- P1-l. Repeated absolute
sourceFilepaths inflate payloads (2026-06-16) —callers/calleesandcall-treenow return a wrapper with a deduplicatedsourceFiles: [...]index and items that carrysourceFileIndex: int.ModuleContextgains a top-levelsourceFilefield. New DTO records:CallRefResponse,AggregatedCallRef,CallTreeResponse,CallTreeItem. - P1-m. Call-site aggregation (2026-06-16) —
callers/calleesqueries usecollect(r.lineNo) AS lineNoswith aGROUP BY (name, type, sourceFile, edgeKind);db-accessesgroups by(name, mode).lineNossorted in Java. - P1-n. Separate
PERFORM(intra) fromCALLNAT(inter) (2026-06-16) — addededgeKindfield (CALLNAT/PERFORM/EXTENDS/IMPLEMENTS) toAggregatedCallRef, and?scope=external/?scope=internalfilter on/callersand/callees. CLI gains--scope. - P1-o. Field projection + sub-array pagination on
/context(2026-06-16) — added?include=functions,dbAccesses,...projection and?limit=N&offset=Npagination applied per sub-array. CLIcontextgains--include,--limit,--offset. - 7. Pagination (2026-06-17) — added
?limit=N&offset=Ntocallers/callees/db-accesses/search-identifier(and the transitivedb-accessesvariant), defaultlimit=50,offset=0. CLI gains--limit/--offset. New ITcalleesPaginationLimitsAndOffsets. - P1-t.
/contextis heavy by default — make sub-arrays opt-in, return counts (2026-06-22) (found 2026-06-21) — the liveWGEAGB0S/contextwas ~50 KB, ~70% of it thevariableAccessesarray (347 entries) which the analysis never used.?include=projection existed but the default still returned every section fully expanded. Flipped the default to lean: heavy sub-arrays (variableAccesses, largesqlStatements) return a summary by default — e.g.variableAccesses: { count, byFunction: {...}, byMode: { READS, WRITES } }— and the full list is emitted only via?include=. A summary line replaces 347 objects; the biggest single token lever found in theWGEAGB0Sanalysis. - P1-v. Tiny
/modules/{name}/digesttriage endpoint (2026-06-22) (found 2026-06-21) —/contextis the "expand" call (tens of KB); there was no "should I dig deeper" call. AddedGET /modules/{name}/digestwith a deliberately small contract:description(P1-u), function count, callers/callees names only grouped byedgeKind, DB table names, referenced data-structure names + field counts (P1-r) — designed to a token budget, not a full dump. - P1-w. Names-only / field-projection mode on list endpoints (2026-06-22)
(found 2026-06-21) —
call-tree,callers,callees,search/identifierare often used only to enumerate names, yet each row still carriedsourceFile(Index), line ranges,dataType, etc. Added a?fields=namevariant that drops per-row detail to just the name (+ type) when the agent is enumerating, on top of the existingsourceFiles-index dedup (P1-l). Complements P1-t/P1-v. - P1-x. Dynamic
CALLNATvia lookup array + keep unresolved dynamic calls visible (2026-06-22) — aCALLNAT <var>whose dispatch variable is loaded from a lookup array (e.g.ASSIGN #TBL(1) = 'WPARTD2S',#W-ACT-PROG := #TBL(#I),CALLNAT #W-ACT-PROG— WPARTX2S L994) was dropped entirely: (1) the parser missed the subscripted literal write (ASSIGN #TBL (1) = …, space before(), and (2) the unresolved dynamic marker was deleted, erasing the call site. Fix: parser now captures subscripted assignment targets; a new intra-module indirect resolver follows the dispatch var's same-lineREADSto the source array and resolves its literals (taggedindirect); unresolved dynamic markers are now kept (only resolved-site markers are reaped) so an agent can still see/investigate the dynamic call incallees/context.
Agent API / MCP tooling gaps
-
78.
project recreate— re-initialise a project from its own stored config (2026-07-17) —POST /api/projects/{p}/recreate(+?deep=true) andac project recreate <name> [--deep]. There was no way to say "wipeupmsand set it up exactly as it was": recreating meant readingproject list, writing downroot/language/excludeDirs/generatedDir/userExitDir, deleting, and re-supplying them by hand. That is not hypothetical — on 2026-07-17 theacproject was deleted, the re-create failed with400 LANGUAGE_REQUIRED(mandatory since item 47), and the project was gone for a while. A typo ingeneratedDirwould not even error; the item-47 LoC split would just report nonsense. The(:Project)shell is not deleted, only itsAstNodes. The shell holds nothing but config (verified:keys(p)= name, root, language, excludeDirs, generatedDir, userExitDir), so keeping it is observably identical to delete-then-create while removing the window in which the config can be lost. The root is resolved before anything is deleted (withResolvedRoot, which already existed forrefresh): unlike create — where an unreachable root costs nothing — recreate would otherwise wipe a good graph and then find nothing to rebuild from. A moved root now fails400 ROOT_NOT_FOUNDwith the graph untouched. Not atomic and not claimed to be: a failure after the delete leaves the project with an empty graph (re-run to finish), and the root could vanish between check and scan. It turns the common silent failure loud; it does not make the operation transactional. No MCP tool — MCP is deliberately read-only + ingest, and this is destructive. Verified byProjectRecreateIT(3 tests); the missing-root guard was vacuity-checked by moving the delete ahead of the check, which fails exactly that test and no other. -
79. One
deleteverb, with the blast radius spelled out (2026-07-17) —clearused to be bound twice, at two levels, with drastically different reach:ac clearemptied the entire database whileac project clear <name>deleted one project — a forgotten word apart, and only the first prompted for confirmation.ProjectCommand.ClearCommandwas also a byte-for-byte duplicate ofDeleteCommand(same body, same endpoint), so it was dead weight on top of an overloaded name. Both are gone; deleting is now:ac project delete <name> one project ac project delete --all every project (prompts; -y skips) ac project delete usage error — never "delete everything"--alland a name are mutually exclusive and one is required, enforced by picocli'sArgGroup(multiplicity = "1")rather than a hand-rolled check. The confirmation prompt is inherited verbatim from the oldac clear. Covered byProjectDeleteArgsTest(6 tests, parse-level:apiClient()builds its client inside the call, so there is no seam to inject a fake and executing would fire real HTTP; what a valid parse does is covered byProjectRecreateITat the API level). Breaking change:ac clearandac project clearno longer exist. -
Batched deletes (2026-07-17, with item 78) —
DELETE_PROJECT/CLEAR_ALLwere a singleDETACH DELETEin one transaction. At the corpus's real size —upmsalone is 454,300AstNodes joined by ~1.3M relationships, 536,210 nodes across all projects — that is one transaction the heap must hold at once, with noserver.memory.heap.max_sizeconfigured (only a 512M pagecache). Deleting was rare enough to get away with; item 78 makes it routine. Both now useCALL { ... } IN TRANSACTIONS OF $batchSize ROWS(agenticcode.delete.batch-size, default 10000). This forced the delete path offsession.executeWriteonto implicit transactions — Neo4j rejectsCALL { } IN TRANSACTIONSinside an explicit one ("can only be executed in an implicit transaction") — which is why the existence check is now a separate statement. Trade accepted deliberately: an interrupted delete leaves a partially emptied project that re-running finishes, versus an unbatched delete that risks not completing at all. -
deploy.sh→manage-ac.sh(2026-07-17) — renamed (viagit mv, history preserved) and given a command list.manage-ac.sh deployis the oldup, unchanged. No argument now prints help instead of deploying — a full deploy bumpsagenticcode.versionand rebuilds everything, which a bare invocation should not trigger by accident. Noupalias: keeping two names for one action is the exact mistake item 79 removed. Addedrestart(restart the server without a build — also the way to abort a running server-side deep refresh, whose HTTP client can be killed without stopping the job),status(containers, server version and projects; every probe guarded so it reports a down stack instead of dying on it underset -e), anddown(whole stack incl. neo4j — deliberately not-v, which would delete the graph volume). All in-repo references updated, including the two user-facing CLI strings that told people to run./deploy.sh cli;features.md's historical entries below are left as they were written.
Found while dogfooding AgenticCode on its own codebase for the J1b task (2026-07-07): friction points where the agent-facing tools didn't answer a question the graph already had the data for.
-
27. Generic "inspect node" tool (done 2026-07-07) — added
GET /nodes/{id}(+ MCPinspect_node): every property of a node, viaNode.asMap()so new parser properties show up automatically.search/identifierandsearch/annotationnow also returnidso there's a way to get one — without that they'd have been unreachable in practice. Nodeids are regenerated on every re-ingest (merge key is(type, name, sourceFile, project);idis unconditionally overwritten) — documented as a limit, not fixed. -
28. Source-snippet-by-node tool (done 2026-07-07) — added
GET /nodes/{id}/sourceandGET /modules/{name}/source?startLine=&endLine=(+ MCPnode_source/module_source), reusing the ingest-time root resolution and the same UTF-8/ISO-8859-1 fallback decoding the parsers use. -
29. Text/string-literal search (done 2026-07-07) —
search_identifieronly matches identifier names; there was no way to search annotation names or string-literal node values (e.g. "does anything reference@Query", "does any SQL string containorders"). Added:search/valuegained acontainsparameter (case-insensitive substring, vs. the existing exact match) for literal/assigned string values; a newsearch/annotationendpoint (+ MCP toolsearch_annotation) finds Java classes/methods/constructors/fields carrying a matching annotation, backed by a new genericannotationsproperty captured at parse time on every such node (independent of any annotation's own specific interpretation elsewhere, e.g.@Entity/@Query). Seex-docs/agent-api-usage-ac-implementation.mdsection 7-8 for usage. -
31. Resolve inherited
REFERENCES/wiring edges on concrete subclasses (found 2026-07-07/08, PURpur-batchre-evaluation, done 2026-07-09) —call_tree/module_digestonly show aREFERENCESedge on the class that syntactically declares it. Where the wiring (e.g. JBeret step construction) lives in a shared abstract base class rather than the concrete subclass —AbstractKeyTableImportJob.jobSteps()wiresKeyTableImport{Init,Processing, End,PassInit,Logging}Step— the concrete subclasses (KeyTableImportValidationJob/KeyTableImportCommitJob) show noREFERENCESat all andcall_tree(..., followWiring=true)on them returns empty, even thoughcall_tree(AbstractKeyTableImportJob, followWiring=true)reaches the full step tree. For 7 of 9 sibling job classes that declare their own wiring directly,followWiring=truealready worked well — this item was specifically about the case where the wiring is declared on an ancestor. Fixed by also followingREFERENCESedges declared onEXTENDSancestors (transitively) when traversingfollowWiringfrom a concrete class, the same way method calls already resolve inherited members. Materialized as a synthetic edge (resolvedVia: 'INHERITANCE'), same class-level over-approximation as the rest of the CHA-style resolution (doesn't know whether the subclass overrides the specific method that declares the wiring). Covered byJavaInheritedWiringIT.subclassInheritingWiringAlsoResolvesIt(classDeclaringItsOwnWiringResolvesTodayis the control test). -
32. Resolve DB-table access on the Repository/Entity itself, not only on the calling Logic class (found 2026-07-07/08, PUR
pur-batchre-evaluation, done 2026-07-09) —db_accesses/sql_statementscorrectly resolve a table when a Logic class calls a repository method (via: "RiskLogic"etc., per item J1a/J1b), butdb_accesses(RiskRepository)/db_accesses(RiskEntity)anddbTablesin their ownmodule_digeststayed empty — even though the repository interface statically carries its generic entity type (AbstractPurRepository<RiskEntity>/IRiskRepository) and the entity carries its@Entity(name = TABLE_NAME). Fixed by attaching the resolved table directly to the Repository/Entity module's owndbTables/db_accesses, taggedmode: "DECLARES"(not a read/write, just "this module maps to this table") — besides the existingREADS/WRITESfrom called functions. Lets a reference implementation's Repository/Entity pair reveal its table directly instead of requiring a call chain through whichever Logic class happens to use it first. Covered byJavaRepositoryOwnTableIT.repositoryOwnDigestListsItsTableWithoutAnyCallerand.entityOwnDigestListsItsTableWithoutAnyCaller. -
33. Hook-contract query for a base class (found 2026-07-07/08, PUR batch-job family comparison, done 2026-07-09) — when comparing/porting a template-method-style family (e.g.
AbstractSteuertabellenProcessingStepwith hooks likevalidate()/writeEntities()/clearTable()), there was no way to ask the graph which of a base class's methods areabstract(subclass must implement),final(fixed, subclass must not override), or a plain overridable default — this had to be read from the base class source by hand every time. AddedGET /modules/{name}/functions?kind=abstract| final|overridable, sourced from the modifiers already visible to the parser (each function also carries akindfield in the unfiltered response,nullfor a Natural subroutine or a constructor). Turns "what must a new sibling implement" into one call instead of reading the whole abstract class. Covered byJavaFunctionKindIT'skindAbstractReturnsOnlyValidate/kindFinalReturnsOnlyCommit/kindOverridableReturnsOnlyLog(unfilteredListsAllThreeMethodsTodayis the control test). -
34. Bulk
function_overridesfor all abstract methods of a base class (found 2026-07-07/08, PUR batch-job family comparison, done 2026-07-09) —function_overrides(base, function)took one method name per call; profiling how an entire family (5+ hook methods × 5+ sibling classes) implements its contract needed one call per hook method. Added an optionalfunctionparam — when omitted,GET /modules/{name}/functions/overridesreturns overrides for every abstract method ofbaseat once, grouped by method then by declaring subclass. Complements item 33 (which methods exist) with "how does each sibling implement them". Covered byJavaBulkFunctionOverridesIT.bulkEndpointGroupsOverridesByHookMethod(perMethodOverridesAlreadyWorkForEachHookis the control test). -
35.
modules?extends={name}filter (found 2026-07-07/08, PUR batch-job family comparison, done 2026-07-09) — listing every concrete subclass of a base class previously required readingcallers(base, ...)and filtering for theEXTENDSedge kind by hand. Added a direct filter on the existingGET /moduleslisting (?extends=AbstractSteuertabellenProcessingStep), making "show me every existing implementation of this pattern" a one-line query, consistent with the existing?moduleKind=/?sourceFile=filters. Restricts to directEXTENDSsubclasses only (one hop, not transitive). Covered byJavaModulesExtendsFilterIT.extendsFilterListsOnlyDirectSubclasses(unfilteredListingContainsAllFourClassesTodayis the control test).
Integration & ops
-
9.
docker-compose.ymlcommitted + wired into the README (2026-07-12) — the rootdocker-compose.yml(Neo4j 5 + theac-code-servercontainer, built fromsrc/main/docker/Dockerfile.jvm) is tracked in git and driven by thedeploy.shwrapper (up/stop/logs/cli). The README's "Getting Started" now leads with the Compose/deploy.shfull-stack path and documents a dev-mode variant that runs only the Neo4j service from Compose (docker compose up -d neo4j) alongsidemvn quarkus:dev, replacing the previous manualdocker run neo4j:5command. -
5. MCP endpoint (HTTP/SSE) (2026-07-07) — implemented the previously empty
ac-code-server/.../mcppackage as a Quarkus MCP server (io.quarkiverse.mcp:quarkus-mcp-server-sse1.9.1, augments cleanly against Quarkus 3.36.2), exposing the API to agents over HTTP/SSE at/mcp/sse(server nameagenticcode). Two@ApplicationScopedtool beans mirror the REST surface by delegating to the sameGraphRepository/ProjectIngestService— no query logic duplicated.McpQueryTools(22 read tools, all@Blocking, returning the same concise JSON as REST viaMcpSupport):list_projects,list_modules,module_digest,module_context,module_functions,module_data_structures,module_dispatch_table,module_columns,callers,callees,call_tree,db_accesses,sql_statements,data_structure_fields,db_table_columns,search_identifier,search_value,variable_reads,variable_writes,flow_forward,flow_backward,field_flow.McpIngestTools:ingest_all,ingest_module,ingest_call_graph. Deliberately read-only + ingest — destructive project create/update/delete/clearAll are not exposed. Errors reuse the REST{error,code,details}shape as MCP tool errors (PROJECT_NOT_FOUND,ROOT_NOT_SET,INVALID_TYPE, …); the deep-query tools return the same deep-ingest hint (pointing atingest_module) when a module isn't deep-ingested. Extracted the project-root guard shared with REST intoProjectRootResolverso the validation can't drift between transports. NewMcpToolsITdrives the whole pipeline through tools only (connect →ingest_all→list_modules/list_projects- a
PROJECT_NOT_FOUNDerror case); all 58 REST ITs still green after the resolver refactor. Docs:agent-api-usage-ac-implementation.mdgained an "Access via MCP" section.
- a
-
30. Version number in startup log / API / MCP (done 2026-07-08) —
agenticcode.versionis a manually-bumped release counter inapplication.properties(deliberately independent of the Maven project version, which stays a build/packaging concern). Single source of truth (VersionInfo), exposed three ways: an explicitINFOstartup log line (VersionLogger),GET /api/version, and the MCPversiontool — plus the MCP protocol's ownserver-info.versionhandshake field, which was previously hardcoded to1.0.0(already drifted from the real build) and now referencesagenticcode.versioninstead of duplicating it. -
36.
ac-cliparity with the REST/MCP API (done 2026-07-09) — the CLI was missing wrappers for 12 endpoints that already had MCP tools:modules,module-data-structures,dispatch-table,digest,function-overrides(bulk + single),search-value,search-annotation,inspect-node,node-source,module-source, andingest call-graph. Added all as newac-clicommands. Also added a hard rule (CLAUDE.md §0) requiring every new/changed REST endpoint to ship with both an MCP tool and anac-clicommand in the same change, so the three stay in sync going forward. -
37.
ac versioncommand (done 2026-07-09) — closes the last REST/MCP-vs-CLI gap:GET /api/versionalready had an MCPversiontool but no CLI counterpart.deploy.shnow stamps ac-cli's bundledagenticcode.properties(version=) fromagenticcode.version(the same release counterVersionInfouses) at build time (stamp_cli_version, called from bothbuild_allandbuild_cli_only), so a jar built outsidedeploy.shreportsdevinstead of a stale number.ac versionprints both the ac-cli and connected server's version and exits 1 with an error if they differ (skipped when the CLI is adevbuild). The interactive shell also prints the ac-cli version on startup and the same mismatch error (VersionCheck, shared with theversioncommand). -
38. Log every MCP tool call (name + arguments) (done 2026-07-09) — new
@McpLogged/McpLoggingInterceptorCDI interceptor pair (com.agenticcode.codeserver.mcp), applied at class level toMcpQueryToolsandMcpIngestTools. LogsINFO com.agenticcode.mcp: MCP tool call: {toolName}({arg=value, ...})for every@Toolinvocation, reading the tool name off@Tool.name()and real parameter names via reflection (relies on the existing-parametersjavac flag). Verified live viaMcpToolsIT. -
54. Source-text regex search (2026-07-14) —
GET /api/projects/{p}/search/source?regex=&limit=&ignoreCase=greps the source text of all modules from disk (deduped by file, bounded by a matchlimit→truncated), returning{module, sourceFile, lineNo, line}. Case-insensitive by default (legacy Natural).SourceSearchServicereuses the parsers' UTF-8/ISO-8859-1 decoding;400 INVALID_REGEXon a bad pattern. REST + MCP (search_source) + CLI (ac search-source <regex> [--limit --case-sensitive]) in sync; Web-UI explorer gains a names / source mode toggle with a results list (each hit opens its module at the line). Complementssearch_identifier(declared names). Covered bye2e/module-explorer.spec.ts. -
56.
search_identifiersigil-insensitive name match (2026-07-14) — the name search now ignores a single leading Natural sigil (#user,&AIV,+GDA) on both sides, soname=K-OUT-MAXfinds the declared#K-OUT-MAX(and#K-OUT-MAXstill works). Implemented in the shared query path (CypherQueries.SEARCH_IDENTIFIERstripsleft(n.name,1);GraphRepositorystrips the search term) so REST + MCP (search_identifier) + CLI (ac search-identifier) inherit it by construction; exact matches for sigil-less names (Java, tables) are unchanged. Covered byAnalysisResourceIT.searchIdentifierIsLeadingSigilInsensitive. (Prompted by the Web-UI popover, where Ctrl-click always captures the#but the global search box did not.) -
52. Function-level callers ("who PERFORMs this subroutine") (2026-07-15) — new
GET /modules/{name}/functions/{function}/callersreturns theFUNCTIONnodes thatCALLSthe named subroutine/method — the intra-modulePERFORMsites (Natural) or cross-class method callers (Java) — with call-sitelineNos, in the sameCallRefResponseshape as module-level/callers.CypherQueries.FUNCTION_CALLERS(reusestoCallRefRow/buildCallRefResponse) + MCPfunction_callers+ CLIac function-callers <module> <function>. Covered byAnalysisResourceIT.functionCallersListsPerformSites(DYNAIDX: INIT-TBL ← DISPATCH; a subroutine called only from the module main body has 0 function callers). UI wiring: the identifier popover's definition-click picker now attributes each PERFORM site to the calling subroutine viafunction_callers(useFunctionCallers), falling back to "module body" for top-level PERFORMs — the complete site list still comes fromcallees?scope=internalso body-level PERFORMs (e.g.INITIALIZATION) are not lost. Covered bye2e/identifier-popover.spec.ts(ADD-XML-LINE ← GEN-XML-LINE attribution). -
53.
search_identifierscoping (sourceFile/module filter) (2026-07-15) — added optionalsourceFile=<relpath>andmodule=<name>filters tosearch_identifierso a caller can ask "this name, in this module" directly (previously the paginated, project-wide search could push the module-local declaration off the page).module=resolves to the module's source file via anEXISTS { MATCH (:MODULE {name}) WHERE .sourceFile = n.sourceFile }subquery;sourceFile=filters exactly. REST (?sourceFile=&module=) + MCP (search_identifierargs) + CLI (ac search-identifier --module --source-file, plus a previously-missing--type) in sync. Covered byAnalysisResourceIT.searchIdentifierCanBeScopedToAModule.
Tests & tooling
- 8.
ac-clitest coverage (2026-06-17) — module had zero tests. AddedVariableAccessQueryTest(query-string builder) andCommandResolutionTest(resolveProject()precedence,apiClient()URL normalization, picocli option-parsing). 12 tests. - P1-p. Fix pre-existing
AnalysisResourceITfailures (2026-06-16) — six IT tests were red onmainbut unnoticed because*ITis excluded from the default surefire run. (1) Five tests double-encoded the#in the path; switched to RestAssuredpathParamso#is encoded exactly once. (2)transitiveDbAccessesAndSqlStatementsIncludeCalleeResultspassed inline Natural source as theclasspathResourcearg; added aningestInline(...)helper. - 11. Natural/Java parser construct coverage review (2026-06-15) —
checked
NaturalParseragainst the CLAUDE.md priority construct list and theWGEAGB0Sfixture set (26 files): all 7 priority constructs are covered. Found and fixed one gap:DECIDE FOR/DECIDE ON(see P1-f). Other constructs (ESCAPE,RESET,EXAMINE,COMPRESS,SEPARATE,WRITE) have no graph-relevant effects and are intentionally out of scope.
Web UI — code understanding & navigation
A React/TypeScript web UI (ac-ui/) that makes the AgenticCode graph navigable
for humans — primary use case: understanding legacy Software AG Natural in
order to migrate it to Java (who calls a module, what it calls incl. dynamic
CALLNAT, which fields flow where, which DB tables it reads/writes, what breaks
on change). Built on the existing REST API; query-driven (never load the whole
graph — upms is ~1000–6000+ modules), WebGL graph rendering, OpenAPI-first
contract. Full vision/architecture: x-docs/ui-proposal.md. Stack: React + TS +
Vite, React Router (URL-driven, shareable deep-links), TanStack Query over the
generated OpenAPI client, Sigma.js/graphology for the large call-graph, CodeMirror
6 with a custom Natural language mode + built-in Java. Out of scope: parser/ingest
semantic changes (the UI is a consumer) and write access to code (read-only; notes
are UI-side). Item 51; M6 (scale & polish) remains open — see roadmap.md.
Backend prerequisites (Milestone M0):
- 48. OpenAPI spec + generated TS client (backend 2026-07-13) — added
quarkus-smallrye-openapiand annotated all REST endpoints (@APIResponse/@Schema, since they return rawResponse) so the spec is fully typed; served at/q/openapi(+ Swagger UI in dev). TS client generation (openapi-typescript+openapi-fetch) lands with the frontend unit. - 49. Ego-graph subgraph endpoint (2026-07-13) —
GET /api/projects/{p}/modules/{name}/graph?depth=&direction=&limit=returns a bounded module-level neighbourhood (nodes + edges,truncatedflag,unresolvedstyling) via a BFS inGraphRepository.egoGraph. REST + MCP (ego_graph) +ac-cli(ac ego-graph) in sync. - 50. SPA serving + CORS (2026-07-13) — CORS enabled
(
quarkus.http.cors.enabled=true), restricted to the Vite dev origins.ingestStatus/ingestDepthnow joined into theGET /moduleslist rows (was only oninspect_node) for the UI's status badges.
Frontend milestones (item 51):
- M0 Foundation + API contract (2026-07-13) — items 48–50 (backend) +
React/Vite frontend (
ac-ui/): project picker, virtualized module explorer (filter + search), ingest-status badges, project/module refresh, thin module-detail fromcontext, generated OpenAPI TS client (openapi-typescriptopenapi-fetch).
- M1 Navigation (2026-07-13) — whole-file source endpoint (REST + MCP
module_source+ CLIac module-source, omit range = whole file); CodeMirror 6 source viewer with a custom Natural StreamLanguage mode + built-in Java, Ctrl/⌘-click name-based identify (popover viasearch_identifier), caller/callee panels, lazy call-tree (expand-on-demand viacallees, cycle-guarded), URL-driven tabs +linedeep-links,STALE_SOURCErefresh prompt. - M2 Call-graph visualisation (2026-07-14) — interactive ego-graph
(Sigma v3 / graphology, WebGL) as a lazy-loaded "Graph" tab: seeds on the current
module via the item-49
/graphendpoint, ForceAtlas2 auto-layout, click-to-select, expand-on-demand (one hop, merged), direction/depth/limit controls +truncatedbadge,unresolved/dispatch (CALLNAT_DYNAMIC)/inheritance edge styling + legend, open-module (disabled for unresolved). No backend change (pure consumer). - M3 Migration dossier (2026-07-14) — a lazy per-module Dossier tab
(
MigrationDossier.tsx) bundling the "porting profile": payload (I/O contract), data structures (expand-on-demand → fields), DB-access matrix (READ/WRITE badges), SQL statements, and the dynamic-CALLNATdispatch table. Line numbers deep-link into the Source tab (?tab=source&line=). Pure consumer of existing endpoints (payload,data-structures/data-structures/{name}/fields,db-accesses,sql-statements,dispatch-table). Follow-ups: (a) a program-view "ingest +callers/callees" button that deep-ingests the program's whole call-graph neighbourhood (the module + its transitive callers and callees, and via each module's USING/INCLUDE fan-out their data structures) —POST /refresh/{name}?scope=neighborhood(seeds the multi-module BFS ingest with the both-direction ego-graph closure;ProjectIngestService.ingestModules)- CLI
ac refresh <name> --neighborhood(no MCP tool —refreshis deliberately REST+CLI only); nginx/apiproxy timeout raised so the long request survives; (b) source-by-path so aUSINGdata area's field line-links open its own file, not the current module — newGET /api/projects/{p}/source?file=&startLine=&endLine=(reusesSourceSnippetService, rejects root-escaping paths with400 INVALID_SOURCE_FILE) + MCPfile_source+ CLIac file-source; the Source tab gains a?src=<file>mode with an "included file … back to module" banner.
- CLI
- M4 Data-flow & impact (2026-07-14, M4.1–M4.3) — flow-forward/backward /
field-flow visualisation + "what breaks?" impact analysis.
- M4.1a subroutine navigation — Ctrl/⌘-click a subroutine name resolves via
the scoped
module_functions(reliable despite dozens of same-named subroutines across modules — the global search is capped): clicking a reference (PERFORM) jumps to the definition; clicking the definition goes to its caller(s) — the PERFORM sites fromcallees?scope=internallineNos(one → jump, many → a picker). Covered bye2e/identifier-popover.spec.ts(Playwright; 5 cases incl. the upms 60-match reproduction). - M4.1 actionable variable matches — identifier-popover non-
MODULEmatches now expand: showdataType/scope, a jump-to-definition (openssourceFile:startLine, via the M3 source-by-path infra so a field defined in aUSINGPDA opens its own file), and lazy reads/writes lists (variable_reads/variable_writes,VariableAccessLocation) with each access site clickable to itssourceFile:lineNo.MODULEmatches still navigate. - M4.2 flow visualisation — a dedicated Data-flow tab (
DataFlowView.tsx) for a selected variable (URL?tab=flow&flowvar=, seeded on the current module): backward (flow_backward) and forward (flow_forward)DataflowStepchains (indented by depth; variable re-targets the trace, module opens it) and field-flow producer→consumer pairs (field_flow, line numbers deep-link into Source). Reached via a "data-flow →" action in the identifier popover. The flow endpoints auto-deep-ingest the scoped module; a409(warm couldn't complete) shows a DEEP_INGEST_REQUIRED prompt wired to the neighbourhood-ingest button. - M4.3 impact ("what breaks?") — a per-module Impact tab (
ImpactView.tsx): the transitive callers (blast radius if the module changes), via the ego-graphdirection=in(useImpactCallers, depth 10 / limit 500), grouped by call distance (direct callers vs. N-hops-away), each dependent clickable to open; unresolved dynamic-dispatch callers shown but not navigable;truncatedbadge when the node cap is hit.
- M4.1a subroutine navigation — Ctrl/⌘-click a subroutine name resolves via
the scoped
- M5 Understanding boosters (2026-07-14, the consumer-feasible parts).
- M5.1 source + structure side by side — a filterable, collapsible outline
(
SourceOutline.tsx,module_functions) beside the Source editor; clicking a function scrolls the editor to it (SourceView re-scrolls onlinechange without a remount — also fixes repeat dossier line-links). Hidden while viewing an included file. - M5.2 notes + saved views — per-module migration notes (
ModuleNotes.tsx, in Overview) and header saved views (SavedViews.tsx), both persisted tolocalStorageviauseLocalStore(the API is read-only, so annotations stay browser-side). - Deferred (need more than a consumer): diff after refresh (backend must expose
before/after snapshots) and LLM summaries (no LLM in the stack; the existing
module_context.descriptionis shown in Overview as the current summary).
- M5.1 source + structure side by side — a filterable, collapsible outline
(
- M6 — Explorer regex filter (2026-07-14) — the module-list search box accepts a
case-insensitive regex over name/sourceFile (falls back to substring while the pattern is
incomplete). Covered by
e2e/module-explorer.spec.ts. (Rest of M6 — virtualisation/large-graph performance, multi-project, auth, theming, export — remains open inroadmap.md.)
UI-sweep bug fixes (2026-07-14/15)
Found via a UI test sweep + source cross-check.
- 55. Dispatch table: multi-value DECIDE branch guardValue artifact (2026-07-14) — a
DECIDE ON … VALUE 'X', ' 'branch (a real literal + a Natural "also match blank" catch) yielded adispatch-tableguardValueof"X, "(the blank' 'trimmed to empty, leaving a trailing", ") instead ofX. Fix:NaturalParser.quotedLiteralsnow drops whitespace-only alternatives (keeping one blank only if a branch has nothing but blanks, so the guard isn't lost); genuine multi-literal branches ('A', 'B') stay comma-joined. Verified live onWGEAGB0S(3 rows @533/537/545 now clean, 0 trailing-comma artifacts across all 11 rows). Covered byNaturalParserTest.multiValueBranchDropsBlankAlternativeFromGuard. - Dossier data-structure fields shown out of order (2026-07-14) — the
data-structures/{name}/fieldsAPI returns fields in graph order ([12,26,40,55,56,41,…]); the UI now sorts bystartLineso a data area reads top-to-bottom (this out-of-order display was what made a correct line number look "wrong" earlier). - 57. Field format glued to the name (no space) parsed as one token (2026-07-15) — a
DEFINE DATAfield whose format is attached with no space — e.g.01 KEY(A1/1:3,1:V),01 #DELIMITER(A1)(common in NATURAL-CONSTRUCT generated code likeCDRANGE) — was stored with the whole token as the identifier name (KEY(A1/1:3,1:V)) anddataType=null, and thus mis-typed as aDATA_STRUCTUREinstead of aVARIABLE. 234 of 376CDRANGEidentifiers were affected, making them unsearchable (a search forKEYfound nothing) and dataType-less. Root cause: theLEVEL_FIELDname group(\S+)greedily swallowed the attached(…). Fix: name group now stops at(([^\s(]+) — Natural identifiers never contain one — in bothNaturalParser.LEVEL_FIELD(deep) andNaturalCoarseScanner.LEVEL_FIELD(Tier-1 identifier index), which carried the identical bug. Covered byNaturalParserTest.fieldFormatAttachedWithoutSpaceIsSplitFromName+NaturalCoarseScannerTest.fieldFormatAttachedWithoutSpaceIsIndexedByBareName. Verification note (superseded by #58): the old glued nodes once required a project re-create to clear; since #58 a plainrefreshreconciles and purges them. - 58. Stale nodes survive
refresh— diff-based reconciliation (2026-07-15) — nodes MERGE on(type, name, sourceFile, project), so a re-parse that renames/removes a field (e.g. after the #57 fix) produced a sibling node and the old one lingered —refreshMERGEd but never deleted. On the liveac/legacy graphs this left Tier-1 identifier-index nodes created at project-create time coexisting with the deep parse's nodes (CDRANGE: 460 = 234 stale-glued + 226 clean), inflating counts and returning ghost matches; the only remedy was a project re-create. Fix (roadmap option b): after a batch of fresh nodes is merged,GraphRepository.mergeResultscollects each real source file's fresh (canonical) node ids and runsCypherQueries.DELETE_STALE_FILE_NODES—MATCH (n {project, sourceFile}) WHERE NOT n.id IN freshIds DETACH DELETE n. BecauseMERGE_NODESoverwritesnode.idto the fresh UUID, a surviving-key node keeps a fresh id (kept) while a renamed/removed field keeps its old id (deleted); moved positional nodes (DB_ACCESS/CONTROL_FLOW) are swept the same way. Gated by areconcileflag threaded throughpersist/persistBatch: true for every full parse (whole-rootrefresh— bothdeep=trueand the default call-graph pass — and the per-module deeprefresh/{name}), false for the coarse Tier-1 scan (subset node set; runs only on an empty graph at create, so nothing to delete).sourceFile=""placeholders (shared cross-file targets) are never swept; reconciliation runs at persist-time, before enrichment, so enrichment-derived nodes/edges are unaffected and cross-module edges are rebuilt byfinalizeProject. Does not remove nodes for deleted files (still roadmap #43). Covered byRefreshReconciliationIT(rename a declared field on disk →refresh→ old identifier gone, new one present); verified live (acwhole-root refresh, 373 files, clean). No REST/MCP/CLI surface change —refreshalready exists on all three; the behaviour change is internal. - 59. Shared Natural field-declaration tokenizer (2026-07-15) — the
LEVEL_FIELDline pattern and itsNN [REDEFINE] name (typeSpec) restdecode were duplicated verbatim inNaturalParser(deep) andNaturalCoarseScanner(Tier-1), so #57 had to be fixed twice and the two tiers could drift. ExtractedNaturalFieldTokenizer(recordNaturalField{level, redefine, name, typeSpec, rest}+ staticparse(line)), the single owner of the pattern;typeSpecis the raw parenthesised content (what the deep parser stores asdataType, byte-identical to before), with derivedbaseFormat()/arrayDims()splitting it andconstValue()/initValue()pulling the trailing clause. Both tiers now call the tokenizer instead of holding private copies; node-type/emission policy is unchanged per tier (a pure refactor — the #57 regression tests in bothNaturalParserTestandNaturalCoarseScannerTeststay green). NewNaturalFieldTokenizerTest(12 cases) covers the edge matrix: format with/without a leading space,A1/1:3,1:Vdims,REDEFINE,CONST<>/INIT<>,#/##/&/+sigils, dimension-only group arrays, and group (no-format) fields. Kills this class of tokenizing bug. - 61. Comments parsed as CALLNAT targets (2026-07-16) — found dogfooding
WGEAGB0S(upms): the unanchoredCALLNAT/CALLNAT_DYNAMICpatterns matched inside comments, socallees/unresolvedwere polluted by phantom modules. Evidence: tree-wide falseunresolvedWAS/RESULTED/DOES(from prose like* What the callnat does:,/* …last callnat resulted in end-of-data);WGEAGB0S'sISINDATEcall-sites included commented lines 526 & 600 (* CALLNAT 'ISINDATE'); a phantomADLML02at copycode lines 18 & 22 (* CALLNAT 'ADLML02'). Root cause: the deepNaturalParsermain loop matched on the raw line (no full-line*skip, no inline/*strip), and the coarseNaturalCoarseScannerstripped inline/*but not a leading*. (Anchored patterns —^\s*PERFORM,^\s*DECIDE,^\s*INCLUDE— were already immune; only the unanchored CALLNAT ones leaked.) Fix: both parsers now skip a full-line comment (first non-blank char*, incl.**SAG) and strip inline/*before statement matching. Verified live:WAS/RESULTED/DOESgone (WGEAGB0S treeunresolved21→17),ISINDATEnow[597,626,1267],ADLML02now[26](real call only). Covered byNaturalParserTest.commentedAndInlineCommentCallnatsAreNotParsedAsCallsandNaturalCoarseScannerTest.commentedCallnatIsNotIndexedAsACall. - 62. Data literals recorded as false MODULE call targets (2026-07-16) — found in a corpus-wide
sweep of
upms(6311 files): 45 distinct data names (browse keys / codes such asCO-TABLA,COD-EMISOR,AGENT-SP,NAME-DESC-SP) sat in the graph as placeholderMODULEcallees, carrying 282 falseCALLSedges — pollutingcallees/call-tree/unresolvedacross the corpus. Root cause: aCALLNAT <bareword>(a browse key reaching the call site through a copycode/macro argument, e.g.INCLUDE YFRAMGC2 'C-MOD-GET' '"CO-TABLA"'expanding toCALLNAT &3& …) is parsed as a dynamic call and gets a placeholder target. It never resolves — no module of that name exists — andDELETE_DYNAMIC_CALLNAT_PLACEHOLDER_EDGESonly drops markers that did resolve, so the false target was kept forever. Fix: a new project-wide enrichment stepdelete-data-literal-call-placeholders(CypherQueries.DELETE_DATA_LITERAL_CALL_PLACEHOLDERS), running after the dynamic-call resolution and marker cleanup and beforestamp-unresolved-placeholdersso the unresolved flags see the reaped graph. It reaps a placeholder only when all four hold: (1) the name carries no Natural sigil (#/&/+) — a genuine dispatch variable is always a sigil'd user variable, soCALLNAT #WIF-style markers stay; (2) the name is a realVARIABLE/CONSTANTof the project (DATA_STRUCTUREis deliberately excluded — a module and its interface PDA routinely share a stem name); (3) no realMODULEof that name exists (else it would simply resolve); (4) every incomingCALLSedge is inferred (CALLNAT_DYNAMIC/INCLUDE_MACRO) — a staticCALLNAT 'X'is trustworthy, which protects the real external modulesRPC-CNTXandUSIX081Xthat collide with field names. Covered byDataLiteralCallCleanupIT(barewordCALLNAT CO-TABLAreaped; sigil'dCALLNAT #DISPmarker kept; staticCALLNAT 'REALMOD'resolved). Verified live on a cleanupmsre-create (2026-07-16):delete-data-literal-call-placeholders: 139 ms; rels +0/-274, nodes +0/-42, and the only survivors of the name-based suspect query areRPC-CNTXandUSIX081X— both reached exclusively by staticCALLNATedges, i.e. gate 4 protecting them exactly as designed. - 63. CALLNAT matched inside a string literal (2026-07-16) — the last of the unanchored-
CALLNATfamily (#61 = comments, #62 = data literals). TheCALLNAT/CALLNAT_DYNAMICpatterns match the keyword anywhere on a line, so prose inside a quoted string fabricated a call. Filed as "low, 1 occurrence"; that was wrong on both counts — a corpus sweep ofupmsfound 3 distinct call sites (each doubled acrosssrc/andgenerated_src/):PRINT '==> callnat before DREQUFN0'(DREQUDN0) →before;WRITE(#MSG) 'NACH CALLNAT ISINGEAG:'(JE0012N0) →ISINGEAG;#ERR-TYPE := 'Callnat USIA008N'(JA0016N0) →USIA008N. Two of the three are the dangerous class:ISINGEAGandUSIA008Nare real modules, so the phantom was not placeholder noise but a real → realCALLSedge in the call graph (PROCESS-AGENT-CHANGE→ISINGEAG,MAIN-PART→USIA008N) — and one invisible as a problem, since onlybeforecarriedunresolved: TRUEwhile the other two wereNULLprecisely because a real module of that name exists. Item 62's reaper structurally cannot clean these (its gate 3 requires that no real MODULE share the name), so only a parse-time fix removes them. (The hundreds of'Start of Callnat'-style lines are harmless: a quote directly follows the keyword, so no identifier matches.) Fix: new sharedNaturalLines(theNaturalFieldTokenizerprecedent from item 59) holdingisInsideStringLiteral/findOutsideStringLiteralplus the formerly duplicatedstripInlineComment; both tiers now gate the two CALLNAT matches on the keyword start lying outside a quoted literal — testing the start, not the whole match, is what keeps a realCALLNAT 'MOD'working (keyword outside, argument inside). Natural's'/"delimiters and doubled-delimiter escapes ('IT''S') are honoured. Considered and rejected: handling a/*inside a literal (whichstripInlineCommentwould truncate, leaving an unbalanced quote) — measured 0 such lines that also containCALLNAT, so it stays out rather than buying speculative complexity. Covered byNaturalParserTest.callnatInsideAStringLiteralIsNotParsedAsACallandNaturalCoarseScannerTest.callnatInsideAStringLiteralIsNotIndexedAsACall, both verified to fail against pre-fix behaviour. Verified live on a cleanupmsre-create (2026-07-16): thebeforeMODULE node is gone entirely, and the fabricatedCALLNAT_DYNAMICedges intoISINGEAG/USIA008N(2 each) are gone while their genuine staticCALLNATedges (10 and 2) remain — the real → real corruption is removed without touching a single real call. - 64. Ingest summary contradicted the graph; dispatch guards lost their alternatives (2026-07-16)
— found by dogfooding a deep ingest of
WGEAGB0S(upms) and hand-checking every endpoint against the Natural source. Two independent defects: (a)unresolvedreported data fields as missing modules. The deep ingest listedMODULE CO-TABLA,MODULE COD-ENTIDAD,MODULE NAME-DESC-SP,MODULE NAME-VALUE-SP— all(A8)fields, no such module anywhere — although the graph was already clean (the scoped enrichment reaped them:rels +0/-7, nodes +0/-5). Root cause:IngestSummary.unresolvedis built during the BFS walk from a filename index (ProjectIngestService:581) and never consults the graph, so it applied none of item 62's gates. This made the item-62 fix incomplete: it cleaned the graph surface and missed the summary — an asymmetry item 63 doesn't share, since a parse-time fix stops the ref existing at all. Fix: filter the unresolved refs through item 62's gates — no Natural sigil, reached only by inferred (CALLNAT_DYNAMIC/INCLUDE_MACRO) edges (reconstructed from the parse result'sCALLSedges, sinceDependencyRefcarries no provenance), and the name is a realVARIABLE/CONSTANT. That last gate is asked of the graph, project-wide (CypherQueries.DATA_FIELD_NAMES), not of the ingested tree: a first, tree-local attempt fixed only 2 of 4 becauseNAME-DESC-SPis declared inBMTABBN2.nat, which is outsideWGEAGB0S's 164-file tree. The static-CALLNAT gate keeps genuinely-missing modules whose names collide with field names reported (theRPC-CNTX/USIX081Xclass). Verified live: WGEAGB0S's unresolved list went 18 → 13, all four false positives gone, while theUSR*modules, sigil'd#GETSHORT-MODUL,NDBERRand theVDB2-*views all stayed. Covered byDataLiteralUnresolvedSummaryIT(incl. theEXTMODcase: a declared field that is also a statically called missing module must stay reported), verified to fail pre-fix. (b) dispatch guards dropped theirVALUEalternatives.guardValueis a single comma-joined string, soVALUE 'GENAGREE-WOUT-SP', ' '(a Natural "also catch blank") silently lost the blank —quotedLiteralsdrops whitespace-only alternatives to avoid a trailing", "— and a multi-alternative branch yields a synthetic"A1, A2"the guarded field never equals. Fix: the parser now also recordswhenValues, every alternative in source order including blanks, andDISPATCH_TABLEsplits it into a realguardValueslist.guardValueis untouched, so REST/MCP/CLI/UI keep working (all three surfaces return the record directly, so the field propagates automatically). Encoded with U+001F rather than modelled as a list becauseAstEdge.propertiesisMap<String, String>; US cannot occur in Natural source, unlike the comma that madeguardValueambiguous. Verified live: WGEAGB0S now reports 3 branches that also catch blank (0 before), and correctly distinguishes theDECIDEat 531 (blank alternatives) from the one at 602 (none). Covered byNaturalParserTest.multiValueDecideBranchKeepsEveryAlternativeIncludingBlank. Scale, measured honestly: 121 blank-bearingVALUEclauses corpus-wide, 3 real in WGEAGB0S; 0 joined multi-value guards observed in the deep-ingested subset (238 guarded WRITES), so the joined-string flaw is real in the code path but unobserved in practice — a full deep ingest would be needed to settle its true rate.
Qualified-field access & re-ingest reconcile (items 78, 79, 74) — 2026-07-18
Three defects found dogfooding a deep WGEAGB0S audit of upms, all in the Natural field
READS/WRITES path, fixed together. All three verified with Testcontainers ITs (independent of the
live server); full ac-code-server IT suite green (190/0/0) with no regressions.
-
78 — group-name-qualified access dropped. A field written qualified by its data area's inner group name rather than the
USINGname (MSG-INFO.##MSG-PGM := *PROGRAM, whereMSG-INFOis the top group of aUSING CDPDA-M) produced no edge:NaturalParser.lookupVariablerecognised onlyUSING-name qualifiers and fell through to a local-only lookup. Fix: when the qualifier is not a known include and the field is not local, mint a bare included-field placeholder (gatedallowIncludePlaceholder && hasIncludes && identifier-shaped) soresolveBareIncludedFieldTargetsredirects it to the real PDA field. TestQualifiedFieldResolveIT.groupQualifiedFieldAccessIsCaptured. -
79 — qualified read inside an expression dropped. The read side of an assignment/expression/
IFtokenised onIDENTIFIER_TOKEN(no dot), so#C := QCPDA.QC-DESCandIF MSG-INFO.##MSG-NR EQ …split into two non-resolving tokens, andIFconditions were not read-scanned at all. Fix: newOPERAND_TOKENkeeps a dotted operand whole;addExpressionReadsresolves a qualified token withallowIncludePlaceholder(legitimate explicit field access) but a bare token without (no placeholder for the "expression soup", so item 74's by-name path is not fed); the assign-RHS loop and theIFhandler both call it. TestQualifiedFieldResolveIT.qualifiedFieldReadInsideAnExpressionIsCaptured. (The originally-filed "qualifier does not disambiguate a shared name" was disproved by a controlled fixture — the qualifier resolves correctly; the live "unresolved" observation was item-74 stale state.) -
74 —
USINGfield misattributed to a same-named local group in another subprogram.WGEAGB0S'sBXFRABA4.#L-FORWARD-VIA-PRIwas linked not only to the realBXFRABA4.pdabut also into unrelatedJX0124N6/JX0129N0. (First mis-diagnosed as a stale-edge coexistence surviving re-ingest — a cleanrecreatedisproved that: the misattribution is created fresh in a single pass.) True cause: a NaturalUSING Xreferences a data area (PDA/LDA/GDA) — always a top-level member of a data-area file — but the placeholder resolver (buildResolvePlaceholderTargetQueries+ the two by-name field resolvers) matchedUSING Xto anyDATA_STRUCTUREnamedX, including a1 BXFRABA4group those subprograms declare inline, and linked every field under it. Fix: aDATA_STRUCTUREplaceholder now resolves only to a real area not owned by aMODULE(NOT EXISTS { (:MODULE)-[:CONTAINS]->(real) }) — a file-level data area, never a program-internal group. Data-area files produce noMODULEnode, so real PDAs are kept and inline.natgroups dropped. TestUsingResolvesToDataAreaNotLocalGroupIT.- Secondary safeguard (separate scenario, kept):
mergeResultsalso runsDELETE_STALE_RESOLVED_FIELD_EDGESin thereconcile(deep-re-ingest-only) branch — deletes a re-parsed file's prior cross-fileREADS/WRITESto aVARIABLE/CONSTANTso finalize rebuilds them, so a module that changes which area it includes does not keep the old resolved edge. Scoped toVARIABLE/CONSTANTand cross-file targets; gated onreconcile. Two-phase testQualifiedWriteReconcileIT.reIngestDeletesTheStaleResolvedFieldEdge.
- Secondary safeguard (separate scenario, kept):
Parallel parse phase (item 24) — 2026-07-18
The ingest parse phase (read file + parse + shell/user-exit metric enrichment) ran as a sequential loop
over all candidate files. It is per-file independent — the JavaParser/NaturalParser instances hold no
mutable state, and copycodes/userExit are read-only — so ProjectIngestService.ingestRoot now submits
one task per file to Executors.newVirtualThreadPerTaskExecutor() (extracted helper parseCandidate).
Results are collected in candidate order (the futures list is parallel to candidates), so cross-file
duplicate detection and persist order stay deterministic; a per-file parse/read failure is still recorded
as an IngestSummary.Failure for that file only. Full IT suite green (192/0/0).
Qualified group-target resolution (item 80) — 2026-07-18
A Natural qualified reference can name a group (a DATA_STRUCTURE), not only a leaf field — e.g.
WGEAGB0S writes BGEAGBA0.#P-DESC-NAME, a level-1 group of BGEAGBA0.pda. The qualifier-scoped field
resolvers matched a real target of type VARIABLE/CONSTANT only, so such writes/reads stayed
unresolved (sourceFile=""). Fix: the two qualifier-scoped (INCLUDES-gated) resolvers
(buildResolvePlaceholderFieldTargetQueries + …ScopedQueries) now accept a DATA_STRUCTURE target as
well (realv.type IN ['VARIABLE','CONSTANT','DATA_STRUCTURE']). Left the bare-field resolvers (which
carry a size(matches)=1 guard) leaf-only, so adding groups cannot turn a previously-unique bare match
ambiguous. Test QualifiedGroupTargetResolveIT.qualifiedWriteToAGroupResolvesIntoTheDataArea. (The
#MAP-T.* reads that were unresolved in the same WGEAGB0S snapshot are leaves, a separate concern, not
covered here.)
Group-qualifier field resolution (item 81) — 2026-07-18
A qualified reference can name a leaf via a group the parser cannot see as a USING member — e.g.
WGEAGB0S reads #MAP-T.V-ID, where #MAP-T is a group inside the included BGEAGA01.pda and the leaf
V-ID also occurs in dozens of other PDAs. After items 78/79 the read was captured but fell to the
bare-field resolver, whose size(matches)=1 guard rightly refused the project-wide-ambiguous leaf, so it
stayed sourceFile="". Fix: NaturalParser.lookupVariable now keeps the group qualifier as a
qualifierGroup property on the bare placeholder (cached under the compound STRUCT.FIELD key so a
group-qualified and a truly-bare reference to the same leaf are distinct); the two bare-included resolvers
(buildResolveBareIncludedFieldQueries + …ScopedQueries) then keep only a candidate nested under a group
of that name, so the qualifier pins the leaf to the one right PDA. The placeholder is a bare VARIABLE
under the module (not a DATA_STRUCTURE area), so the by-name resolvers never see it — no
#74-style misattribution risk. Test GroupQualifiedLeafResolveIT.groupQualifierDisambiguatesAnAmbiguousLeaf;
full IT suite green (193/0/0).
Read-path CONTAINS bounding (item 75 read path) — 2026-07-19
The runtime read queries traversed a module's statements with an unbounded (m)-[:CONTAINS*0..]->(src).
CONTAINS is not acyclic (shared copycode CONTROL_FLOW nodes, parser line-range artefacts), so on a
heavy module (ACCNPE01) that expansion blew up and callees/digest/context hung (callees >2 min).
Every edge source is a MODULE (depth 0) or a FUNCTION that is a direct CONTAINS child (depth 1) —
verified corpus-wide (0 sources deeper; DB_ACCESS parents only FUNCTION/MODULE) — so the traversal is
provably equivalent when bounded to *0..1, which cannot walk the cycles. 20 read-side traversals updated
(callees, MODULE_HOP_OUT, DISPATCH_TABLE, EGO_NEIGHBORS_*, VARIABLE_ACCESSES, DB_ACCESSES,
SQL_STATEMENTS, FUNCTION_CALLERS, SEARCH_BY_VALUE, fieldFlow, BUILD_CALLS_MODULE, and the
*_FOR_MODULES batch variants). Query-only change (no recreate). Guard ReadPathBoundedTraversalIT
(EXPLAIN asserts no unbounded CONTAINS expand); full IT suite 193/0/0. Live (v76): ACCNPE01 callees
2 min → 0.16 s,
digest>10 s → 3.3 s,context>10 s → 3.1 s.
Version bump moved from deploy into the build — 2026-07-19
agenticcode.version (the integer build counter reported by /api/version and used by ac version for
stale-CLI detection) used to be incremented by manage-ac.sh deploy (shell bump_version), so a plain
mvn clean install never bumped and only a deploy did. It now lives in the Maven build: a new
ac-mvn-plugins mojo bump-version, bound to ac-code-server's generate-resources phase, increments
the counter and stamps the same number into ac-cli's agenticcode.properties. Binding to
generate-resources (before process-resources) means the freshly built jar already reports the bumped
number, keeping the jar's baked version and the CLI stamp in lock-step. The mojo only acts when a
requested goal is package/install/deploy (read from MavenSession.getGoals()), so mvn test,
mvn compile and quarkus:dev do not bump; deploy bumps because it runs a full install. Every such
build bumps unconditionally (no source-change gating — high numbers are harmless). manage-ac.sh's
bump_version was removed; stamp_cli_version is kept only for the CLI-only rebuild path (cli), which
syncs the CLI to the current server version without bumping. Logic unit-tested (BumpVersionMojoTest,
7/0/0); live-verified: mvn generate-resources left the counter unchanged, mvn package bumped 77→78 and
stamped ac-cli 78, and target/classes/application.properties carried 78 (timing correct).
Manual override for unresolvable dynamic CALLNAT targets (item 82) — 2026-07-19
Surfaced by the WGEAGB0S deep-API audit: the dynamic-CALLNAT resolvers cannot recover every target —
e.g. YGEAGGNH's CALLNAT #GETSHORT-MODUL, whose name is assembled by MOVE 'YGEAGKEY' TO #GETSHORT-MODUL + MOVE 'GN0' TO SUBSTR(#GETSHORT-MODUL,6,3) → YGEAGGN0 (a real, ingested module).
Such a site is left an unresolved CALLNAT_DYNAMIC placeholder. A human or agent can now pin it via a
REST/MCP/CLI override, keyed by (originFile, lineNo) and fanning out to one or more target modules
(deliberate branches). Stored as a :DynamicCallOverride node deliberately not labelled :AstNode,
so DELETE_PROJECT_NODES (which matches :AstNode {project}) never wipes it on a refresh; an enrichment
step apply-manual-dynamic-callnat (after the auto resolvers, before the placeholder cleanup) re-applies
every override automatically — MERGEing a CALLS edge (callKind=CALLNAT_DYNAMIC, resolvedBy='manual')
to each target and flagging the placeholder marker manualHidden (kept, not deleted) so the read queries
suppress it. Precedence: an override applies only while the site is still unresolved; once an auto-resolver
resolves it, the override is skipped and listed obsolete. Reset (DELETE, one site or all) clears the
manual edges and un-hides the placeholder inline — no refresh needed, no fragile src-node tracking.
Target validation rejects a non-module name (400 UNKNOWN_TARGET) so a typo can't reintroduce a phantom.
Bug B (same audit) fixed alongside: callees/digest now surface the unresolved flag the graph
endpoint already exposed, so an unresolved dynamic target is machine-distinguishable without inspecting
sourceFile. API surface (REST + MCP + CLI, in sync): GET /dynamic-calls/unresolved,
GET/POST/DELETE /dynamic-calls/overrides; MCP list_unresolved_dynamic_calls,
list_dynamic_call_overrides, set_dynamic_call_override, reset_dynamic_call_override; CLI
ac dynamic-calls unresolved|overrides|set|reset. The agent system prompt (agent-api-system-prompt.md)
now requires an agent to investigate and pin an unresolved dynamic call rather than report a dead end.
Covered by DynamicCallOverrideIT (6/0/0): resolve, multi-target, unknown-target rejection, inline reset
restore, obsolete precedence, and override survival across a refresh?deep=true.
MCP surface removed (item 26) — 2026-08-04
The MCP server was removed from the product. It existed as a second transport in front of the same
services the REST API already exposes: McpQueryTools (770 LoC, 38 @Tool methods), McpIngestTools
(2 tools), plus McpSupport/McpLogged/McpLoggingInterceptor — 973 LoC of main code and the
145-LoC McpToolsIT, all deleted, together with the quarkus-mcp-server-sse/-test dependencies
(root POM dependencyManagement + mcp-server.version property + ac-code-server POM), the
quarkus.mcp.server.server-info.* properties, and the repo-root .mcp.json client registration.
Rationale. Item 26 (MCP session reliability) never became fixable from this codebase: calls failed
with "the first message from the client must be initialize: tools/call", at first intermittently and
by 2026-08-02 from the first call of a session onwards, while the equivalent REST endpoints answered
normally. Rather than carry a second, unreliable transport plus the CLAUDE.md rule that every REST
change be mirrored into an MCP tool, the surface was dropped.
No capability lost. All 40 tools were verified to have a REST twin before deletion — including the
non-obvious ones: version → GET /api/version, ego_graph → GET /modules/{name}/graph,
file_source → GET /source?file=, project_loc → GET /loc, and the three dynamic-call-override
tools → GET/POST/DELETE /api/projects/{p}/dynamic-calls/overrides. The MCP layer held no logic of
its own; McpSupport only serialized service results into the same JSON the REST resources return.
Docs & rules updated. CLAUDE.md (sync rule now REST + ac-cli; tool priority now REST → CLI →
grep/Explore; architecture diagram, project structure, ADR table), README.md (architecture diagram,
tech stack, endpoint list, "Agent usage" section replacing "MCP"), x-docs/agent-api-system-prompt.md
("Access via MCP" section and its REST↔tool mapping table dropped), x-docs/agent-api-usage-ac-implementation.md
(tool names replaced by their endpoint paths throughout; filename kept to avoid breaking ~6 inbound
references), x-docs/agenticcode-ueberblick.md, x-docs/presentation.md, prompts/CLAUDE.md, both
prompts/*-deep-api-audit.md, and .claude/settings.local.json. Historical MCP mentions in this file
and in earlier roadmap.md entries are left intact as record.
Closed roadmap items (moved out of roadmap.md 2026-08-28)
roadmap.md states in its header that it tracks only open work, but 59 completed items had been left
standing in it. They are moved here verbatim — no summarising, no rewriting — grouped by the section
they came from and in their original order. Two sections whose content was entirely narrative about
finished work moved with them.
One correction was made in transit: a second item had also been numbered 91 (search/identifier
?priorityModule=) and is now 151; the original 91 (copycode provenance on
db-accesses/workfile-accesses/sql-statements) keeps its number.
High priority — Agent API gaps found in the UPMS→PUR reengineering (2026-08-18)
Six gaps reported after a week of daily agent use on upms / pur / app, ordered by the time they
cost. All probes below were run against a freshly refreshed graph and are reproducible as written.
Items 125 and 127 were the two said to make the API return a wrong answer rather than a missing
one, which is why they led the list. 125 is fixed (2026-08-18); 127 was retracted the same day —
its probe was not ambiguous, so the answer had been correct all along (see the retraction below).
All six are closed: 125 and 127 (2026-08-18, 127 retracted), 126 and 129 (2026-08-18), 128 and 130 (2026-08-19).
-
125.
search/identifierdid not match a type declaration's short name — and silently ignoredcontains( fixed 2026-08-18)Symptom. The class is in the graph, but the search endpoint cannot find it.
GET /pur/search/identifier?name=PartnerOtherUpdateLogic → [] GET /pur/modules/com.uniqagroup.pur.logic.upms.partner.PartnerOtherUpdateLogic/context → the module, 8 functions GET /pur/search/identifier?name=PartnerOtherUpdate&contains=true → []Only methods, fields and variables are indexed. This breaks the single most common question an agent asks before writing anything: does this name already exist? The only workaround is to guess the FQN and try
/modules/{FQN}/context— and then a404means "guessed wrong" or "does not exist", which are not distinguishable.&contains=trueworks on/search/valuebut is accepted and ignored here — it returns[]rather than a match or an error. "Every identifier containingupd" was the central question of a rename pass and could not be answered at all.Diagnosis (2026-08-18). The reported cause was wrong in a way worth recording: type declarations are indexed —
?name=com.agenticcode.neo4jstore.graph.GraphRepositoryreturns theMODULEnode.SEARCH_IDENTIFIERmatchedn.namealone, and a Java module'snameis its FQN (item 117) while the short form lives inn.simpleName. So the index was complete and only the short key was missing — which produced exactly the reported symptom.containswas not a parameter of the endpoint at all, so JAX-RS dropped it silently.Delivered.
SEARCH_IDENTIFIERmatches the sigil-strippedn.nameorn.simpleName;?contains=true(CLIac search-identifier --contains) switches both to a case-insensitive substring match as on/search/value, andcontainswithout anameis400 MISSING_NAMErather than a full node dump. EachIdentifierMatchnow carriessimpleNameandmoduleKind(CLASS/INTERFACE/ENUM/RECORD,PROGRAM/SUBPROGRAMfor Natural,nullfor non-MODULEhits) — the discriminator the item asked for. The unindexed substring scan is bounded tooffset+limitrows via a$scanCapon the query'sLIMIT, chosen sopaginate()'s existing "limit <= 0means unlimited" rule is untouched. Covered byIdentifierTypeDeclarationIT(8 tests);x-docs/agent-api-usage-ac-implementation.mdupdated.Measured after deploy (2026-08-18).
pur?name=PartnerOtherUpdateLogic→ the class (was[]);pur?name=Builder&type=MODULE→ 1 match,AuthorizationInterceptor→ 2 (which is what item 127 could not check). The substring scan onupms(454,300 nodes) is 1.2–2.3 s warm — usable, but not free: it is a label scan, so keep alimit. -
126.
/projectscarried no ingest metadata, so "not found" was never evidence (fixed 2026-08-18)Symptom.
GET /api/projects {"name":"upms","root":"…","excludeDirs":["target"],"language":"java", …}No
ingestedAt, no file count, no failure list. Theupmsproject is known to be incompletely ingested, but nothing in the API says so — so every negative answer has to be cross-checked against the file system, and the consuming project has had to write that rule into its own agent instructions. The same blind spot hides staleness: after a code change there is no way to tell whether an answer predates the edit short of running a refresh.Delivered. The
(:Project)shell records what the last whole-root ingest did, returned as a nestedingestobject onGET /projectsand on the newGET /projects/{p}(CLIac project show):ingestedAt,mode(tier1/call_graph/full),filesExamined,filesPersisted,filesFailedfailures(capped at 200 with an explicitfailuresTruncated, so a short list is never read as the whole story),durationSeconds,serverVersion. Three deliberate semantics, each pinned by a test inProjectIngestMetadataIT(6 tests):
ingest: nullmeans never recorded, not "ingested nothing" — a zero-filled record would state a fact nobody measured, which is the same class of error as the one this item reports.- Only whole-root passes write it. A by-name
refresh/{name}, a deep ingest or a fan-out warm ingests real files but walks a fraction of the tree; letting one moveingestedAtwould advertise the whole project as freshly walked because one module was deepened. ingestedAtis not a freshness guarantee — it dates the walk, not the match against disk. That is item 129's hash work;GET /{p}/source?file=(409 STALE_SOURCE) is today's real check.
Kept apart from
UPDATE_PROJECT(user config,COALESCEd) so neither can clobber the other, andProjectInfocarries the state in a nested component rather than flat — the same record is the input to an ingest, and a run must not be handed the state it is about to replace.Not delivered (deferred, not forgotten). "Ideally every response carries the project's
ingestedAt" is a cross-cutting envelope change — responses are bare arrays today. It belongs with item 130's scope-echo work, which rewrites the same envelopes. -
128. No endpoint for "every reference site of this symbol" (fixed 2026-08-19)
Symptom.
callersgives module-level call edges with line sites. There is no way to ask for all occurrences of a name — import, type position, field type, annotation argument, test reference. A rename of four classes had to be scoped by unioning thesourceFilesof fourcallersresponses plus severalsearch/identifiercalls, and two affected files (PartnerControllerTest,AbstractUPMSServiceAcceptanceTest) were only found because the agent already knew they existed.Delivered.
GET /search/references?name=&kind=(CLIac references) returns{sourceFile, lineNo, kind, inModule, target}per mention, unioningCALL,IMPORT,TYPE(declared field/parameter/return),ANNOTATION,EXTENDS,IMPLEMENTS,INJECTS,CLASS_LITERALand Natural'sINCLUDE. The name may be the identity or the short form (item 125); an unknownkindis400 INVALID_KIND, never an empty list.SearchReferencesIT, 9 tests.The Java parser now emits the mention edges it never had — imports, declared type positions and annotation usages — measured on
puras ~26.6k imports and ~13.4k declared-type positions against ~194k existing edges, so roughly +20%.Measured ingest cost (2026-08-19 deploy): deep refresh went
ac19s → 106s,app37s → 106s,pur29s → 125s,upms1230s → 1496s. A 3-5x slowdown on the Java projects is the price of this index, and worth re-reading before assuming it is free.Verified live on
pur:?name=PartnerControllerreturns 65 sites (41 CALL, 20 INJECTS, 2 IMPORT, 2 TYPE) — includingAbstractUPMSServiceAcceptanceTest, one of the two files this item reported as "only found because the agent already knew they existed".Two decisions worth keeping:
- A new
MENTIONSedge type, notREFERENCES. ReusingREFERENCESregressed the call graph —callers/callees/call-tree/ego-graphfollow it as wiring, so animportsurfaced as a caller (javaCrossClassCallGraphSpansFilescaught it:edgeKind: REFERENCESwhere a method call belonged). A separate type keeps "mentions" out of "calls" by construction rather than by remembering to filter it in every existing query. - Imports are filtered to project-internal ones (sharing the importing file's first two package
segments), and the JDK/framework name list moved to
ac-parser-core(ExternalTypeNames) so the parser and the ingest apply the same exclusion. Otherwise everyjava.util/framework import mints a placeholder node that is created, persisted and swept again on every ingest.
Known gaps, stated rather than hidden: local-variable types and generic type arguments are not indexed (
List<Target>recordsList); same-package references have no import, so within one package the index rests on declared-type positions; Natural has no import or type-position concept and contributes call/include/inheritance kinds only. And the index is only as complete as the last re-parse — existing graphs need arefreshbefore it is populated. - A new
-
129. Refresh was all-or-nothing (fixed 2026-08-18)
Symptom. Twelve changed files;
POST /pur/refreshexamined 2 981. The feedback loop after a code change is minutes, which discourages the verify-after-change step the workflow depends on. Worse, an aborted deep refresh leaves the graph half-updated with no marker, so every later query silently answers from that state — a hazard the consuming project documents in its own instructions because the API does not surface it.Delivered, in three independent pieces (CLI in step:
ac refresh --paths/--changed-only):POST /refresh?paths=a/B.java,c/D.java— targeted deep re-ingest of the named files plus their dependencies, on the existingingestFilesBFS. Paths matching no file come back inunresolved(a typo'd path must not read as a successful refresh), and paths escaping the root are refused by the same guard thesourceendpoints use. It deliberately skips the deleted-file sweep and does not moveingestedAt— both are only meaningful for a whole-root walk.ingest.incompleteon/projects— set before a whole-root pass, cleared on success, so a run that never finished stays flagged. It cannot self-heal (a killed process clears nothing), which is the correct direction to fail: a false "incomplete" costs one refresh, a false "clean" costs trust in every answer.?changedOnly=true— hash-based skipping, opt-in rather than the default the item asked for.
Why
changedOnlyis not the default. Implementing it that way would have silently corrupted the graph in three ways, all found while validating rather than after shipping:- Natural copycodes are inlined at parse time. A module whose
.cpychanged parses differently while its own hash is unchanged — so it would be skipped and keep a stale expansion, with nothing in the graph to indicate it. A changed copycode now disables skipping for the whole run (tested). This alone rules out "incremental by default" forupms, where 137 modules share one copycode. - Duplicate detection groups the files it parsed — a subset can only confirm duplicates among changed files. Markers are never cleared (the query only MERGEs), so this loses discovery, not recorded facts.
- User-exit LoC annotation is re-stamped only on re-parsed files.
Two bugs found verifying against the live server (2026-08-19), both fixed with regression tests:
?paths=pom.xmlwas accepted rather than reported — a path that exists but is not an ingestible source file passed the guard, was listed asexamined, and leftunresolvedempty. That is exactly the "I ingested 2 of your 3 files" invisibility the list exists to prevent.- The copycode stand-down fired on a Java project:
accarries.cpyfiles as Natural test fixtures that its Java walk never ingests, so they had no stored hash, counted as changed, and disabled skipping entirely — 487 of 509 unchanged files were re-parsed, i.e.changedOnlydid nothing at all. The guard is now Natural-only (an unknown language keeps the conservative path).
Honest about the win: enrichment is project-wide and still runs in full, so this cuts parse+persist only. The premise in the symptom above has also moved — a full deep
purrefresh measured 29 s on 2026-08-18 (upms: 20 m 30 s, the one project where this really pays). Covered byTargetedRefreshIT(8 tests).Not delivered: surfacing
incompleteon every query response, as opposed to on/projects. That is the same envelope change item 126 deferred, and it belongs with item 130's scope echo. -
130. Two smaller ones: REST-path routing, and project scope invisible in the answer (fixed 2026-08-19)
Delivered.
GET /rest-endpoints(CLIac rest-endpoints) returns the composed path, verb, declaring class and handler, with?module=to narrow. The parser now persistsrestPath(class and method) andhttpMethod— annotations were stored by name only, so the path string was not in the graph at all. A@Pathwritten as a constant reference resolves the same way a JPA@Columnname does; a method with no verb annotation is not an endpoint. Narrow by choice: this is two properties, not a general annotation-argument store, which would be a JSON blob per node against item 111d's grain.- Scope and freshness are echoed as response headers —
X-AC-Exclude-Dirs,X-AC-Ingested-At,X-AC-Ingest-Incomplete— on every project-scoped response. This also closes the "every response carriesingestedAt" half that items 126 and 129 deferred.
Why headers rather than body fields. Most endpoints answer with a bare JSON array (
db-accesses,functions,search/identifier, …). Adding a scope field there means restructuring array → object, which breaks the web UI's generated client, the CLI printers and any agent that indexes[0]— too high a price for a hint. The cost of the choice is stated in the docs rather than hidden: an agent reading only the body will not see them. The project shell is cached ~10 s so the headers add no query per request, and the ingest path invalidates that cache explicitly — a staleincomplete=falseduring a running refresh would point exactly the wrong way.RestEndpointsIT, 11 tests, including the headers riding on a bare-array response.Three bugs that only real data exposed (found verifying against
purafter the 2026-08-19 deploy, fixed the same day)://file— the path halves were joined with a singlereplace('//', '/'), and replacement is non-overlapping, so'///upload'collapsed to'//upload'. Each half is now stripped of its own leading/trailing slash before joining.- 183 duplicate rows of 436 — the graph holds more than one
CONTAINSedge between the same module and function (see item 75), so every such endpoint was emitted twice.RETURN DISTINCT. POST /— JAX-RS inherits@Pathfrom a base class or interface, which this codebase uses heavily (AbstractFileTransferUiSvc). Those endpoints reported the bare/: a wrong answer, not a missing one. The class path now falls back to the nearest ancestor's@Path.
Worth recording that none of the three was visible in the fixture-based tests that passed first time — they were found by looking at the real output on
purand disbelieving it. A fourth followed from the same habit: three@RegisterRestClientinterfaces were listed as served endpoints, which states the traffic's direction backwards. They now carryoutbound: truerather than being dropped, since "what does this application call out to" is a real question.Not changed: the projects still have different
excludeDirs(app:["test","target"],pur/ac:["target"],upms:[]). The header makes the asymmetry visible; silently normalising someone's ingest scope to make answers look consistent would be the wrong fix.
High priority — the silent-truncation defect of item 103 is still open on the three
search/* endpoints (2026-08-19)
-
131.
search/identifier,search/valueandsearch/annotationsilently truncated at the defaultlimit=50— a bare array with no total and notruncatedflag (fixed 2026-08-19)This is item 103's own follow-up, in 103's words: "If a
truncatedflag is wanted anyway, it should be a separate item covering all paginated endpoints, not just these two." The fix for 103 was deliberately narrowed todb-accesses/workfile-accesses; the three search endpoints match 103's criterion exactly and were left as they were.Symptom. Measured on
pur, freshly refreshed:GET /pur/search/annotation?name=Immutable → 50 rows, ends at FolderEntity GET /pur/search/annotation?name=Immutable&limit=500 → 95 rowsThe response is a bare JSON array, and there is no header either — checked for
X-Total-Count,Content-RangeandLink: none is sent. The caller cannot tell the two answers apart from the outside.This produced a wrong finding in a real audit, not a hypothetical one.
TASKS.md58 of the UPMS→PUR project recorded that 17 of 35 read-only tables had lost their@Immutableannotation, and concluded that a generator run would silently break optimistic locking on them. The entry was derived from the truncated page. Re-checked on 2026-08-19 against all 114 rows ofTables_meta.csvin both directions: 94 withWritable=falsecarry@Immutable, 20 withWritable=truedo not, zero deviations. 16 of the 17 named entities are precisely the rows the cut dropped; the 17th,FolderEntity, is the hit on the boundary. The finding cost a day and was pure artefact.Cause.
AnalysisResource.searchIdentifier(:872),searchByValue(:789) andsearchAnnotation(:812) all passeffectiveLimit(limit), i.e.DEFAULT_PAGE_LIMIT = 50(:176-180).uncappedLimit(:198) — added for item 103 with a Javadoc that states the criterion as "endpoints where a silently-capped default is actively harmful… a bare JSON array with no total and notruncatedflag" — is not applied to any of them.Why these three and not every paginated endpoint. The cap only bites where the query means enumerate, and only
search/annotationdoes at scale. Measured spans:| Endpoint | Sample | Range | | :--- | :--- | ---: | |
search/identifier| 10 queries | 1 – 41 | |search/value| 5 queries | 0 – 4 | |search/annotation| 12 queries | 14 – 3 630 |identifierandvaluenever approach 50 in practice — uncapping them is free.annotationis where the damage is, and also where an uncapped default is expensive: a row costs ~272 bytes (~68 tokens), so@Column(3 630 rows) would be ~1 MB / ~250 k tokens in one response. Removing the cap outright, as 103 did, is therefore not obviously right here.Proposed fix, in order of preference.
- Keep the default, make the cut detectable. Send
X-Total-CountandX-Truncatedheaders on the paginated bare-array endpoints. Zero bytes in the body, no contract change for any existing REST/MCP/CLI/UI consumer — exactly the objection that stopped 103 from adding an envelope. This also makes paging usable for the first time: the caller learns after call one whether paging is worth it. - If only the number may change:
uncappedLimitforidentifierandvalue(their result sets are small), and raiseannotationto ~300 — above the sets anyone asks a completeness question about (@Path276,@RunInTransaction233,@Entity160,@Immutable95) and far below@Column. ?countOnly=true. A completeness question is a counting question. It would answer "how many classes carry@Immutable" for a few bytes instead of 20 k tokens.
Paging already works and is not the fix. Verified on the 95-row set:
offsetis honoured, the order is stable across repeated calls (five identical checksums), five pages of 20 reassemble the single fetch exactly — 95 rows, 95 distinct, no gap, no duplicate, same order — and reading past the end returns200with[]rather than an error. So "short page = done" is a sound stop signal. But paging is opt-in, and the caller who does not know the answer was cut is exactly the caller who will not page. Two caveats worth recording: the order is by internal node id, not alphabetical (KeyTableEntryEntitysorts last), so it is stable only within one graph state — arefreshbetween pages can shift the set; and a total that divides evenly by the limit costs one extra empty call to detect.Documentation gap, independent of the code.
limit/offsetare documented inagent-api-system-prompt.mdonly forcontext(§2). For the three search endpoints (§8) the parameters are not mentioned at all, so an agent reading the guide has no reason to suspect a cap. Fixed in the same pass — see §8 and the pitfalls list.Delivered (2026-08-19), option 1 + option 3 of the three proposed above.
X-AC-Total-CountandX-AC-Truncatedon all three endpoints. Bodies stay bare arrays — no contract change for UI/CLI/MCP/agents, which is the objection that stopped item 103 from adding an envelope, and consistent with item 130'sX-AC-*scope headers.?countOnly=true(CLI--count-only) returning{"count": n}— the completeness question answered for a few bytes rather than ~250 k tokens on@Column.- Default limits unchanged (option 2 rejected): raising them moves the cliff and makes the
expensive case worse. With the cut visible and
countOnlyavailable, 50 is the right default. - CLI parity, which headers alone would have broken.
ApiResponsewas(statusCode, body)and dropped headers, so a truncated CLI answer would have looked complete — the same silent cut, one layer out. It now carries headers and prints a note to stderr (stdout stays pipeable intojq).
Two implementation decisions worth keeping:
- Each query is split into a shared predicate core plus row/count projections, composed from one constant. Two hand-maintained copies would drift, and a total that disagrees with its rows is worse than no total — it turns a visible truncation into a confident wrong number.
- The count query runs only when the page comes back full. A short page is provably the end of
the set, so its total is arithmetic. That matters most for
search/identifier?contains=, an unindexed scan measured at 1.2-2.3 s onupms; counting unconditionally would have doubled it. The critical trap, pinned by a test: the count must not inherit item 125's$scanCap, or the total equals the row count every time and the whole item is silently undone.
Covered by
SearchTruncationIT(8 tests). - Keep the default, make the cut detectable. Send
Follow-ups found while closing 125-130 (2026-08-19)
-
134. A post-deploy API smoke test (2026-08-20)
Why. Items 129 and 130 both shipped with bugs that no test caught and that only manual live probing exposed:
//filein a REST path, 183 duplicate rows of 436, an inherited JAX-RS@Pathcollapsing toPOST /,?paths=pom.xmlbeing accepted as a source file. The integration tests pin semantics against fixtures; nothing checked the deployed server against the real graph.Delivered.
x-scripts/verify-api.sh—curl+python3only, no Maven, no Docker, ~10s againstac/app/pur, ~25s againstupms../x-scripts/verify-api.sh [-p <project>], exit 0/1, onePASS/FAIL/SKIPline per check. Five groups: reachability and version; the item-130 scope and freshness headers; the item-131 paging contract (X-AC-Total-Countnumeric,countOnlyagrees with the full total,limit=1setsX-AC-Truncated, consecutive offsets are disjoint); data plausibility per endpoint family; and the negative cases (structuredPROJECT_NOT_FOUND/MODULE_NOT_FOUND, and the item-129?paths=pom.xmlregression).Deliberate limits. Every assertion is an invariant (
> 0, no duplicates, header present, required field set) — never a fixed row count, because counts move with each refresh and differ per project. Endpoint families that a project legitimately lacks (rest-endpointsin a pure Natural project, an annotation search with no hits, a leaf module with no callees) reportSKIP, notFAIL. The run is read-only apart fromrefresh?paths=pom.xml, which resolves no file and starts no ingest. The duplicate-row check onrest-endpointscannot surface item 75 — that query already appliesDISTINCT; it only guards against theDISTINCTbeing dropped again. This is not a test replacement and must not be treated as a quality gate.It paid for itself on the first run — three real defects, now items 135 and 136 below.
-
135.
search/referencesandrest-endpointswere left out of item 131 — they still truncate silently ( 2026-08-20)Symptom.
verify-api.shreports on all four projects:search/references: X-AC-Total-Count is numeric — got '' (status=200), same forrest-endpoints. Neither response carriesX-AC-Total-CountorX-AC-Truncated.Cause. Both handlers in
AnalysisResourcereturnok(...)on a plainListinstead ofpaged(...)on aPage<>.searchReferencesuseseffectiveLimit(limit), so it caps at the default 50 with no signal at all — exactly the failure item 131 set out to remove.restEndpointsusesuncappedLimit(limit), so it does not lose rows today, but it is equally silent about how many there are.Delivered. Both queries split into
*_CORE+ a shared row projection + a*_COUNTthat isWITH DISTINCT <the same columns> RETURN count(*)— the same column set as the row projection, on purpose: a narrower distinct key would count rows the page never delivers. Neither count carries$scanCap, for the reason already pinned onSEARCH_IDENTIFIER_COUNT.restEndpointsPage/searchReferencesPagego through the existingwithTotal(...), so the count query runs only when the page comes back full — and withrest-endpoints' default uncapped limit it never runs at all. Both endpoints gained?countOnly=true, both CLI commands--count-only;warnIfTruncatedalready sat inprintResponse, so the stderr note started working the moment the headers appeared.SearchTruncationITgrew from 8 to 15 tests.One thing the fixture taught us: a reference to a type in the same package produces no edge at all — there is no import, and
TypeResolvercannot turn the bare simple name into an identity. The test fixture had to put the referenced class in its own package to have anything to page over. That limit is documented onJavaParser#addReferenceEdgesand is worth knowing before scoping a rename inside a single package. -
136.
search/identifierfaults with an unstructured 500 on a deep page (2026-08-20)Symptom.
GET /api/projects/ac/search/identifier?name=e&contains=true&limit=500answers 500 with a plain-text Quarkus error page.limit=300is fine, so it is not the limit itself but which rows land in the page. Server log:org.neo4j.driver.exceptions.value.Uncoercible: Cannot coerce NULL to Java int.Cause. The identifier row mapper coerces
startLine/endLineunconditionally, and at least one node inaccarries NULL there. Only reachable past ~300 rows, which is why no test and no manual probe hit it.Two defects, not one. Beyond the mapping bug, the response violated the API principle that errors are structured JSON — an agent got an HTML-ish body with no
codeto branch on.Delivered. Three changes, because one alone would have been a patch over a symptom:
- Cause.
MARK_DUPLICATE_IDENTITIES(item 114) was the only node-creating query that set neitherstartLinenorendLine. It now sets both viacoalesce(n.startLine, 0)—coalescebecause itsMERGEkey is deliberatelyMERGE_NODES', so marker and reference placeholder can be one node whose real lines must survive.0is a convention, not a truth: a marker has no line. The alternative — letting two properties be null on a handful of nodes — pushes null handling into every row mapper in the project, and that is exactly how this bug happened. - Defence. The identifier row mapper uses the existing
intOrZero(record, key)helper. A row mapper must never be the thing that faults an endpoint. The otherasInt()call sites were reviewed and left alone: they readFUNCTION/FIELD/SQL rows that always carry lines, and churning 50 call sites would have been a different, larger change. - Contract. New
ApiExceptionMapper(@Provider,ExceptionMapper<Throwable>) returns500 INTERNAL_ERRORwith anerrorIdindetailsmatching the logged stack trace, which is not echoed to the client. TheWebApplicationExceptionpass-through is load-bearing —Throwableis the least specific mapper there is, and without it every deliberate 404/400/409 would have become a 500.ApiExceptionMapperTestpins both directions.
Not retroactive. The write-side fix takes effect on the next ingest; the 11 broken nodes in
acstay until then. The mapper hardening makes the endpoint correct immediately, refresh or not. - Cause.
Lazy / deferred ingest (three-tier model)
Reworks ingest from eager whole-project parsing into a lazy, on-demand model.
Three tiers: Tier 1 = cheap eager reference index (per file: nodes,
identifiers, coarse call/DB references — no deep bodies); Tier 2 = lazy deep
ingest (control flow, statement-level dataflow, precise reads/writes) triggered
on demand; Tier 3 = source served from the filesystem, no longer stored on
nodes. Reverse queries (callers, search_identifier, flow_backward) stay
answerable because Tier 1 pre-indexes coarse references globally.
Items 36–43 are done (Tier-1 reference index + tri-state status, Tier-2 lazy
deep-ingest, depth/node caps, unresolved-reference nodes, Tier-3 source-from-disk
with stale check, the refresh surface, and hash-based auto-invalidation) — see
x-docs/features.md. No open items remain in this track.
Ingest performance
(Items 24 — the stale-file sweep's missing (project, sourceFile) index, the actual persist
bottleneck — and 25 — batch persist, found already implemented — completed 2026-07-16 and moved to
x-docs/features.md. A full upms call-graph refresh (6311 files) now takes 179 s end-to-end
(~63 s parse, 104.5 s persist across 32 batches, ~12 s finalize), against ~2.5–2.8 min per batch
before. The parked "parallel parse phase" idea was implemented 2026-07-18 (item 24) — see
x-docs/features.md. No open items remain in this track.)
-
111a.
NodeTypeas a Neo4j label, stage 1: written on every persist (done 2026-08-05)Measured problem. The node type lived only in the
typeproperty, so every expand-and-filter read it out of the property store. Profiled onupms(570,739 nodes, 2,002,935 relationships):| query | dbHits | rows | |---|---|---| |
(:AstNode{type:'MODULE',project})-[:CALLS]->(:AstNode{type:'MODULE'})| 1,598,264 | 5,384 | | same, without the target'stypefilter | 663,156 | 25,843 |~935k dbHits — 58% of the query — spent reading one string property, ~36 hits per candidate node because
typesits in an 11-property chain. A label lives in the node record and costs no property-store access. Of the 200type: '…'occurrences inCypherQueries, 106 filter ontypewithout bindingname, and none bind both — so the composite index(project, type, name)is never used with all three columns, and the win is in traversal filtering, not in index seeks.Delivered (stage 1, deliberately additive).
SET node:$(n.type)in both persist paths (MERGE_NODES,MERGE_POSITIONAL_NODES— dynamic labels verified available on Neo4j 5.26.27), plusBACKFILL_TYPE_LABELSrun fromensureSchema(), batchedIN TRANSACTIONSso 570k nodes do not build one transaction state larger than the 2 GB Neo4j heap. Labels take no part in the MERGE identity, so merge semantics are unchanged. Thetypeproperty stays, so none of the ~200 queries change behaviour and nothing has to migrate at once. Covered byTypeLabelIT.Not yet done — stage 1 alone buys nothing. It only creates the precondition:
- 111c: drop the
typeproperty and the fourAstNodecollection indexes once no query reads it (−9 chars/node).
- 111c: drop the
-
111b. Hot-path queries anchored on the label instead of the
typeproperty (done 2026-08-05, builds on 111a)Migrated in
CypherQueries:callers(both scopes),callees(both branches and itsscopefilter, which runs on every expanded candidate),MODULE_HOP_OUT/MODULE_HOP_OUT_WIRING(the call-tree traversal),FUNCTION_CALLERS, and the two module anchors inSEARCH_IDENTIFIER. Its$typefilter stays a property comparison on purpose — it is a bound parameter, not a literal, so it cannot become a static label.Two new indexes were required, not optional. A Neo4j index serves exactly one label, so
(m:MODULE {project, name})cannot use the:AstNodeindexes. Withoutmodule_project_name/function_project_namethe anchor lookup would degrade from an index seek to a label scan — turning the migration into a regression at the very point it is meant to help.~158
type: '…'literals remain in the enrichment queries; they migrate incrementally, each with its own measurement.Baseline captured before deploying (upms, best of 3 per endpoint, version 160):
| module | endpoint | before | |---|---|---| | WGEAGB0S |
callers/callees| 24 ms / 14 ms | | WGEAGB0S |call-tree?depth=3/depth=5| 266 ms / 312 ms | | ZINERR01 (1689 callers) |callers| 329 ms | | YFRAMN04 (916 callers) |callers/call-tree?depth=5| 335 ms / 262 ms |Measured after deploying (version 162) — the performance case did not hold up.
| module | endpoint | before | after | |---|---|---|---| | WGEAGB0S |
callers/callees| 24 / 14 ms | 21 / 16 ms | | WGEAGB0S |call-tree?depth=3/depth=5| 266 / 312 ms | 255 / 292 ms | | ZINERR01 |callers| 329 ms | 323 ms | | YFRAMN04 |callers/call-tree?depth=5| 335 / 262 ms | 277 / 141 ms |At the Cypher level, where dbHits are deterministic and HTTP noise is absent:
| query | before | after | change | |---|---|---|---| |
callersYFRAMN04 | 16,202 dbHits | 14,369 | −11% | | module hop (call-tree core), 4 seeds | 1,864 dbHits | 1,850 | −0.8% |Why the 15.8× microbenchmark did not transfer. It scanned every module in the project and filtered the expansion; the real queries seek one module by name and expand over hundreds of edges. There are simply almost no property reads left to eliminate. The
YFRAMN04 call-tree1.9× is not attributable to the migration — the module-hop core it is built from improved by 0.8%. A 12-run repeat of the one apparently slower endpoint (WGEAGB0S/callers) gave min 21 ms / median 43 / max 66: no regression, just a measurement dominated by noise.Kept despite this, on a different justification than the one it was proposed under: it is a small consistent improvement and no regression, and — the actual reason — 111c cannot happen without it. Dropping the
typeproperty requires that no query reads it. The remaining value of 111a/b is the storage reduction 111c unlocks, not query speed.Lesson recorded deliberately: the pre-implementation benchmark was chosen for convenience, not for resemblance to the production query shape, and overstated the benefit by more than two orders of magnitude. Benchmark the query the code actually runs. 295/295 ITs pass.
Write cost — measured 2026-08-05, no regression. A
project recreate upms --deepwith the label writes in place took 541 s for the parse/persist phase (6311 files), against 853 s for the same phase without them. Not a like-for-like comparison — a recreate creates nodes rather than merging onto existing ones and skips per-file reconciliation, so it is expected to be faster — but there is no sign of the slowdown that would have forced a revert.Payoff confirmed on the real graph — bigger than estimated. Same graph, same query, same result (5381 rows), only the filter style differs:
| | dbHits | time | |---|---|---| |
(:AstNode{type:'MODULE',project})-[:CALLS]->(:AstNode{type:'MODULE'})| 1,601,071 | 512 ms | |(:MODULE{project})-[:CALLS]->(:MODULE)| 101,120 | 25 ms |15.8× fewer dbHits, 20× faster — but do not read this as the payoff of the migration. It is not. This query starts from every module in the project and expands, so the property filter is applied to hundreds of thousands of candidates. The real API queries anchor on a single module by name via an index seek and expand from there, over hundreds of edges rather than hundreds of thousands. See 111b for what the migration actually delivered on those (11% and 0.8%, not 15×). The number above is a property of this microbenchmark, not of the codebase.
-
111d-1. The 36-char node UUID is gone — replaced by two inline longs (done 2026-08-05)
Measured share. Of 100,000 sampled nodes, 100% carried an
idover Neo4j's ~15-char inline threshold (vssourceFile96%,name33%,value13%,type0%). At ~2.45 long strings per node against a 385 MB string store, the UUID is roughly 41% of it (~157 MB).What it actually did — and the mistake that nearly shipped. The first proposal asserted the id had "exactly one load-bearing use" (the item-58 reconciliation sweep). That was wrong: it is also the join key wiring a batch's edges to the nodes merged in the same transaction (
mergeEdgesBatch). The assertion came from greppingn.idpatterns and classifying the hits, without following the dataflow — the edge query was in the same file and was missed. Removing the property on that basis would have produced a graph with nodes and no edges, failing silently. The change was reverted and re-proposed. Enumerate by dataflow, not by grep sample.Design. Per persist transaction: one
ingestGenstamp (wall-clock seeded, monotonic, so a fresh JVM can never reuse an earlier run's generation) plus a per-transactionnidcounter. Both inline longs. Verified precondition:persistandpersistBatchboth wrapmergeResultsin a singleexecuteWriteWithoutResult, so nodes and their edges are always in one transaction — which is why a counter suffices where a UUID was needed.- edges:
MATCH (a:AstNode {nid: e.sourceNid, ingestGen: $ingestGen}), (b …) - reconciliation (items 58/76):
WHERE n.ingestGen IS NULL OR n.ingestGen <> $ingestGen. TheIS NULLhalf is required — Cypher's three-valued logic makesNULL <> xyieldNULL, so legacy nodes would otherwise be undeletable. - new index
(ingestGen, nid)— required, not an optimisation: without it every one of upms's 2M edge endpoints degrades from a point seek to a label scan. ast_node_id_uniquedropped explicitly (deleting the CREATE alone would leave it in place on every existing database), andDROP_LEGACY_NODE_IDSstrips the retired property from pre-existing nodes —SET node += n.propertiesdoes not remove properties, so without it every surviving node would keep its old id forever.
The API contract got stronger, not weaker.
/nodes/{id}now returns Neo4j'selementId(4:<db-uuid>:<internal-id>), which survives a re-ingest — where the UUID was overwritten on every single merge, which is what forced the old "same ingest generation only" caveat. VALIDATE had called this a deliberate weakening; that was backwards. Affectsac inspect-node/ac node-source; the web UI never used node ids.Also removes the ~1.1M UUID strings shipped as query parameters per full refresh (two id lists). Covered by
NodeIdentityIT; 300/300 ITs pass.Not yet verified: the net store saving after subtracting the new
(ingestGen, nid)index. Only measurable on the store files after a full re-ingest. - edges:
-
112. A whole-project refresh logs no completion line (done 2026-08-06)
A finished refresh was only visible as
POST /api/projects/upms/refresh?deep=true -> 200, an access-log line naming neither the project's outcome nor how long it took. Anything watching the log for completion — a monitoring script, an agent,rebuild-and-refresh.sh— had to grep HTTP status lines, which is why a watcher written during the 111d-1 verification missed the end of a 1191 s run outright.ProjectIngestService.refreshProjectnow emits one INFO line covering both modes:Refresh finished: project='upms', mode=deep, files=6311, modules=6311, failed=0, 1191 sPlaced after
sweepDeletedFileOrphansso the timing covers the deleted-file sweep too — the line means "the refresh is done", not "the ingest is done". Per-modulerefreshModuleis deliberately left alone: it is short, called often, and would only add noise. -
113. Every server operation logs when it finishes (done 2026-08-06, extends 112)
Rule applied: an operation that logs a start must log an end. Four violated it.
The access log was already there — its duration was just never recorded. Every request has been logged as
POST /api/projects/upms/refresh?deep=true -> 200 (-ms); the(-ms)is not a missing feature butquarkus.http.record-request-start-timedefaulting tofalse, so%{RESPONSE_TIME}has nothing to render. Setting it gives every endpoint a completion line with a real duration for one property and oneSystem.nanoTime()per request. It adds no new log lines — the access log already covered every call. The property lives inVertxHttpConfig, notVertxHttpBuildTimeConfig, i.e. it is a runtime property, so it does not have to be repeated in the test-sideapplication.properties.Two domain completion lines, both at the shared worker rather than the entry points:
ProjectIngestService.ingestRoot→Project ingest finished: project=…, mode=tier1|call_graph|full, files=…, persisted=…, failed=…, duplicates=…, N s— covers project create, the call-graph pass and the deep whole-root pass.ProjectIngestService.bfsIngest→Deep ingest finished: project=…, seed=…, files=…, failed=…, unresolved=…, truncated=…, N s.
bfsIngestis the single worker behind bothingestModules(by name) andingestFiles(the fan-out warm). Lines first written intoDeepIngestCoordinatorat the two call sites were removed again once that was traced — they would have double-logged every deep ingest. Thefilescount is per operation, deliberately: the walk emits oneIngesting <module> …start line per file, so the single end line is what those N start lines add up to.No
try/finallywith anoutcomeflag, though the proposal called for one. All three sites already log their failure path (Fan-out warm … failed,Auto deep-ingest … failed, and aningestRootexception surfacing as a 500 with duration once the access log works). Buying symmetry would have meant restructuring the definite-assignment flow of a 90-line method on the hot ingest path for a log line — a bad trade.A refresh now logs two lines by design:
Project ingest finished(parse + persist) andRefresh finished(item 112, additionally covering the deleted-file sweep). The difference between them is the sweep's cost, which was not visible anywhere before.refreshModulestays silent — short, frequent, pure noise.
Known bugs
-
119. Enums, records and annotation types are modules (found 2026-08-06 in the
pursource-vs-API cross-check; done 2026-08-06)The Java parser iterated
ClassOrInterfaceDeclaration. Everything else was invisible — inpur76 annotation types, 67 enums and 25 records, 168 types with noMODULEnode at all (plus 384package-info.java, correctly ignored).GET /modules/BatchParam/digestanswered404; a record referenced from another file stayed an unresolved placeholder and answered409 NOT_INGESTED. The API did not lie — item 107 saw to that — but a whole category of type was unanalysable.The loop now iterates
TypeDeclaration. Exactly three things differ between the four kinds, and they are resolved once in aTypeFactsadapter instead of scatteringinstanceofthrough a 250-line body: only a class/interface canextends, only an annotation type can do neither, andisInterfaceis a class/interface notion.moduleKindgainsENUM | RECORD | ANNOTATION.What the probe corrected in the plan. Rather than reason about the JavaParser API I ran a parser over a fixture with all four kinds, and two assumptions were wrong:
- A record's compact canonical constructor (
Point { … }) is not returned bygetConstructors()—CompactConstructorDeclarationdoes not extendCallableDeclarationand so does not fit the callable machinery. Left uncovered deliberately (1 occurrence inpur), and named here because the calls in its body stay invisible. - An annotation type's
String value();is not a method —getMethods()returns none. Building it as planned would have given 76 annotation types a module node with an empty body: "analysed, nothing found" again. Members are modelled asFIELDs withdefaultValue, because the question asked of an annotation is which attributes it carries.
Record components and enum constants are captured for the same reason — a record DTO reporting zero fields is worse than no answer. The
TypeResolverand the enclosing-type walk had to follow, or a reference to an enum declared in the same file would have been unqualifiable — item 117 running backwards for precisely the types it just gained.Kept class/interface-only on purpose: the JPA/Panache repository heuristics. A record is not a Panache repository, and generalizing guesswork to types it was never written for produces confident wrong answers rather than silence.
Watched, not asserted: the CHA fan-out grows, since an enum implementing a project interface is now a real implementation target. And 168 additional types can make a previously unique
simpleNameambiguous, so a short name that used to work may now answer409— the honest answer under 115, but a visible change. - A record's compact canonical constructor (
-
118b. A control character in one constant disabled two guards, and the tests that should have caught it asserted the same wrong literal (found 2026-08-06 in the
purcross-check; done 2026-08-06; both regressions from 116b)A — the marker leak.
UNRESOLVED_FIELD_RECEIVER_PREFIXcontained a strayU+0001:JavaParser.java:157 = ".field:"; hex: 3d 20 22 01 66 69 65 6c 64 3a 22 3bSo the parser wrote
field:xwhile both Cypher guards matchedSTARTS WITH 'field:'— the cleanup inDELETE_UNRESOLVED_FIELD_RECEIVERSand the exclusion inRESOLVE_SIMPLE_NAME_REFERENCES. Neither ever matched. Inpur, 202 marker nodes survived, 180 of them still wired with 850 edges, and the scaffolding was served from the public API:GET /modules/…KundeService/callees → { "name": "field:e", … }Why it survived review and tests. The two ITs written to guard exactly this checked
m.name STARTS WITH 'field:'andnot hasItem("field:repo")— the same wrong literal. They were true because they matched nothing. Code and test were wrong in the same way, so the test could not see the bug it existed for.Fixed at the source, and both predicates now match
CONTAINS 'field:': a ':' cannot occur in a Java or Natural module name (verified across all four projects), so it is equally sharp — and it also clears markers left by an earlier build, which a refresh would never reach otherwise (placeholders have nosourceFile, so the per-file sweeps do not touch them).B — lambda parameters taken for inherited fields.
isProbableFieldReceiverasked "lower-case and not indeclaredTypes?". Lambda andcatchparameters are in neither the callable's parameter list nor itsVariableDeclarationExprs, so.map(e -> e.getX())looked like an inherited field namede— 24 markers from one class alone, none of them ever resolvable. The bound names now go into the existingshadowedset. Deliberately not intodeclaredTypes: an implicit lambda parameter has typeUnknownType, and feeding that to the receiver resolver would turn a silent omission into a confident edge to a module named after a non-type.Both fixes proven by reverting them. With A restored to its broken form, three assertions fail — including
field:somethingUndeclaredreaching/callees. The first B test I wrote was itself vacuous: with A fixed the cleanup deletes every marker, so nothing is observable end-to-end. It moved toJavaParserTest, where reverting B turns it red with[field:entry, field:ex, field:inherited]. The new assertions match the marker anywhere in the name and, separately, assert the invariant that was actually violated — no module name contains a control character. -
118. The Java DB resolvers scanned the project once per candidate row (found 2026-08-06 while watching a
purrefresh; done 2026-08-06; regression from [117])Symptom. A deep refresh of
pur(2988 files) spent 391 s of 447 s in two of 48 finalize steps — and both produced zero edges:[10/48] resolve-java-db-access : 305 382 ms rels +0/-0, props 0 [11/48] resolve-java-query-jpql : 85 537 ms rels +0/-0, props 0 the other 46 steps : ~10 s totalCause, and it was mine. 117 added
OR m.simpleName = …so endpoints accept the short form, and an index to keep that a seek — but declared it asFOR (n:MODULE), while these enrichment queries anchor on(:AstNode {type: 'MODULE'}). The label indexes are therefore invisible to them, and the plan confirms what that costs:+Union | +Filter repo.simpleName = a.javaReceiverType … | +NodeIndexSeek RANGE INDEX repo:AstNode(project, sourceFile) | 216674 rows | +NodeIndexSeek RANGE INDEX repo:AstNode(project, type, name) | 198 rowsThe
simpleNamebranch scans the whole project — under anApply, i.e. once per DB_ACCESS row. 3563 candidates × 55 030 nodes ≈ 196M property reads. The same 117 note that (correctly) refused to putOR simpleNameinside the 18 API queries for exactly this reason missed that four enrichment queries already had it.Fix. Anchor on the dynamic labels:
(a:DB_ACCESS …)<-[:CONTAINS]-(fn)instead of a label scan over all 573kAstNode+ 639kCONTAINS, and(repo:MODULE …)/(entity:MODULE …)so all fourORbranches become seeks (216 674 estimated rows → 1 and 14). Applied toRESOLVE_JAVA_DB_ACCESS,RESOLVE_JAVA_QUERY_JPQL,RESOLVE_JAVA_QUERY_NATIVE_SQLand the twodb-accessesAPI queries that share the pattern.This is not a pure plan change, and it was verified as such. It depends on every node carrying its type label; if one lacked it, the module would silently drop out of resolution and report "no DB access" — the [114] failure class. Checked both directions on live data (all four projects, including the freshly refreshed
purwith its placeholders):AstNode{type:'MODULE'}without:MODULE= 0,:MODULEwithout the property = 0, same forDB_ACCESS. Labels come fromSET node:$(n.type)in both node-merge paths, i.e. at ingest time, and no other query createsAstNodes.Equivalence, measured rather than argued. Same 200 candidate rows through both formulations, result triples
(function, access, table)sorted and hashed:6c829e01…for old and new alike — 12.3 s vs 1.7 s. Full query over all 3563 rows: 1.7 s (was 305 s).Not claimed: the JPQL step cannot be verified on
pur— it has zero JPQL rows there, so its 86 s were pure anchor cost and only the anchor fix applies. And every timing above is warm-cache; the 305 s arose cold, right after 2988 file writes. What the finalize actually costs after this gets measured on the next refresh, not predicted here. -
117. A Java module's identity is its fully-qualified name (done 2026-08-06; the root fix behind [115])
115 stopped the wrong answers; this removes the cause. A Java module node's
nameis now the FQN (com.example.OrderService, anda.b.Outer.Innerfor a nested class — verified against JavaParser, which also returns empty for local/anonymous classes, so those keep the simple name).simpleNamecarries the short form for display. Endpoints accept either: the FQN resolves exactly, the short form is the convenience and falls back to 115's409 AMBIGUOUS_NAME.The rename is the easy half; references are the hard one. All eight reference sites in the parser created their placeholders from
simpleType(...). Renaming only the definitions would have left every reference unable to find its definition — the whole Java call graph, silently. ATypeResolverper compilation unit now qualifies from the file's own declarations and its explicit imports, and refuses to guess: "not imported, therefore same package" would have producedcom.example.String. What stays unqualified is picked up by the newresolve-simple-name-referencesenrichment step, which matches onsimpleNameand only when exactly one module matches. Measured justification: 12 of 2988 files in one real codebase use a wildcard import (0.4%).Resolution happens once, in the API guard.
MODULE_INGEST_STATEmatchesname OR simpleNameand returns the resolved identity, which the guard hands to the endpoint; every downstream query then runs on the identity alone. The alternative —OR simpleNameinside all 18 module queries — would have dropped them from an index seek to a label scan on 500k nodes, undoing items 111a/b. A(project, simpleName)index was added to keep the short form a seek as well.What the tests caught that review did not. Two regressions, both silent:
- Cross-class dataflow went empty. The
?module=scope filters are not behind the guard, so they got a short name and matched nothing. Fixed by resolving the scope centrally (resolveModuleName), also used by the module-hop BFS seed and the?extends=filter. - The
em.persist(...)write path disappeared.RESOLVE_JAVA_DB_ACCESStestsjavaReceiverType IN ['EntityManager', 'Session', 'StatelessSession']— hardcoded simple names. QualifyingEntityManager(it is imported, so the resolver could prove it) made that branch dead. The fix was to stop qualifying the DB type properties rather than qualify three hardcoded lists: the module lookups accept the simple form viasimpleName, so nothing is lost where it matters.
Isolating the first one is worth recording: disabling the new enrichment step and seeing the test still fail is what proved the cause lay in the query, not in the resolution.
Two of my own errors, for the record. The reference-resolution query was first written with APOC (not available here) and then registered as enrichment step 24 although its own comment said it must run first — leaving every later join reading a placeholder about to vanish.
Also unified on the way:
declaredInis a display label in every producer (column metadata and both function queries), i.e. the short name. Two of the three had drifted to the identity.UI (done 2026-08-06, after the deploy that unblocked the codegen). The client is generated from the running server's OpenAPI, so this had to wait for a build that knows
simpleName. Regenerating pulled in exactly the 115/117 contract and nothing else: 18sourceFilequery params, 11 reworded 409 descriptions, 1simpleName.The find was that this is not a cosmetic task.
Explorer.tsxfilters the module list by regex overname, andnameis now the FQN — so an anchored pattern on the class itself returned nothing:^AbstractLogic$ → 0 Treffer (name = com.uniqagroup.common.base.AbstractLogic)The existing explorer spec runs against
upms(Natural, no dots in any name) and stayed green throughout. The filter now also matchessimpleName; the package path and the FQN remain searchable.Display: eight sites render the short name with the identity in the
title— module list, module header, graph node labels and the selected-node chip, callers/callees, call tree (including its cycle markers), impact list, dataflow steps. Identity is untouched everywhere it matters: routing, query keys,key=, requests, and the graphology node key. Where the server sendssimpleNameit wins over the client-side rule, because it knows the cases a string rule cannot — a class in the default package, or a local class for which JavaParser reports no qualified name.The derivation ("text after the last dot") was checked against the server rather than assumed: for all 4734 qualified modules of
pur, it equalssimpleName— 0 divergences; and 0 of 3587upmsmodule names contain a dot, so Natural is provably unaffected. Worth noting the graph is currently mixed-generation (onlypuris re-ingested post-117), which is what makes the prefer-then-derive fallback necessary rather than nice.Accepted loss: a nested class shows as
Inner, so twoOuter.Innerin different outers look alike in a list. The FQN is one hover away and the list also shows the source file.New e2e spec
java-fqn.spec.ts(the only one that runs against a Java project) — and it was verified to actually catch the regression by reverting the filter fix and watching it go red.tsc --noEmit+vite buildclean; Playwright 21 passed, 1 pre-existing failure inidentifier-popover.spec.ts(expectsline=323, gets324— reproduced with all UI changes stashed, so it predates this work and belongs to the upms data, not the UI). - Cross-class dataflow went empty. The
-
115. Module endpoints address Java classes by simple name and silently merge distinct classes that share one (found 2026-08-06,
pursource-vs-API cross-check; done 2026-08-06)Symptom.
GET /modules/BrokerHistoryTests/functionsreturns 627 functions. The class it names has 107. The other 520 belong to four unrelated classes:MATCH (m:MODULE {project:'pur', name:'BrokerHistoryTests'}) RETURN m.sourceFile → HistoryCalculationServiceTest.java 84 HistoryComplexServiceTest.java 203 HistoryRecordPageServiceTest.java 128 HistoryRecordServiceTest.java 107 HistoryVariantServiceTest.java 105 = 627Five distinct
@Nestedclasses, one per test file, each legitimately its own module — correctly ingested, correctly distinguished in the graph bysourceFile. The API collapses them, because/modules/{name}/…matches onnamealone and the path offers no way to disambiguate (verified against the OpenAPI spec: the only parameters areprojectandname).Not a test-only problem.
WorkingStorageis six nested classes in six production files (LastAgentLoopLogic,AuthorizationCheckStandardProcessingLogic,InitializationService, …). Itsdigestreports all six enclosing classes asCONSTRUCTORcallers, reading as one shared class constructed in six places when in truth each file has its own.Scale in
pur: 163 ambiguous names covering 385 modules — ~8% of all 4738 modules. Worst cases collide 7-fold (ProcessTests,VermittlerstammTests). Java makes this normal: nested test classes,Builder,Config,Handler,WorkingStorage.Why it matters. This is worse than a wrong count. Every derived answer — callers, callees, db-accesses, call-tree — is a union across unrelated classes, presented with no indication that a merge happened. An agent asking "who uses
WorkingStorage?" gets six answers for six different classes as though they were one, and cannot tell.Solution sketch.
- Make ambiguity visible before making it addressable. When a name resolves to more than one
module, answer
409 AMBIGUOUS_NAMEwith the candidatesourceFiles indetailsrather than silently unioning. This is the same shape as the 107 fix and stops wrong answers immediately, at the cost of breaking queries that currently "work". - Then make it addressable. Accept an optional
sourceFile(or fully-qualified name) query parameter on the module endpoints to pick one candidate;search/identifieralready returns thesourceFileper hit, so an agent has what it needs to disambiguate in one extra call. - Natural is unaffected in practice — its module names are file-stem-unique by language convention, and genuine collisions are already handled (skipped) by the duplicate detection; see [114], which is the other half of this problem.
Not yet decided: whether Java modules should carry the FQN as
nameinstead. That would fix addressing at the root but changes every response body and the whole UI, so it needs its own proposal rather than being folded in here.Delivered. Both halves together — the guard alone would have made 385 modules in
purand 326 inappunreachable by name, trading one wrong answer for another.ModuleIngestStatecarriescandidates(the real, non-placeholdersourceFiles for the name) andambiguous().MODULE_INGEST_STATEreturns them and honours a$sourceFilefilter.withModule/withIngestedModule(the item-107 guards) answer409 AMBIGUOUS_NAMEwith the candidates indetails. Guard order matters: absent → ambiguous → placeholder. A placeholder is never a candidate, so one real module beside a placeholder is not ambiguous and#SUBPROGRAM-style references keep answeringNOT_INGESTED.?sourceFile=on the module endpoints, threaded through 18 Cypher queries and the repository.ac --source-fileon all 17 module commands, via a picocli@Mixinwhoseappend()picks?or&— several commands add their own parameters after the base path and a fixed?emitted two.- UI:
409no longer maps blindly toNOT_INGESTED; the body'scodedecides, and the banner lists the candidates. Listing them is the point — a bare "ambiguous" is barely better than the merge.
Two corrections worth keeping. First, a line-based inventory found 27 affected Cypher sites; a block-based one found 33 — the same grep-instead-of-dataflow error as item 111d-1, 20% low. Second,
EGO_NEIGHBORS_IN/OUTwere patched and then unpatched: they are keyed on the BFS frontier's name, not the root's, so filtering them by the root'ssourceFilewould have broken traversal. Ambiguous neighbour names still merge inside a call tree — the same scope boundary item 107 has, and it is not closed here.The runtime hazard predicted in review actually fired. Cypher rejects an unbound parameter at runtime, not compile time;
moduleFunctionsbuilt its parameters as a hand-rolledHashMap(becausekindis nullable andMap.ofrejects nulls), so it was the one site the centralmoduleParamshelper did not cover, and it returned500. Caught byAmbiguousModuleIT, which asserts a real selection (1 method vs 2) rather than just a200. - Make ambiguity visible before making it addressable. When a name resolves to more than one
module, answer
-
116. A method call on a field the calling class does not declare produces no edge (found 2026-08-06,
pursource-vs-API cross-check; done 2026-08-06)Symptom.
HistoryRecordService(53 methods, exercised by 185@Testmethods) reportscallers: {}— nobody calls it. The tests call it constantly:class HistoryRecordServiceTest { private HistoryRecordService service; // line 33, outer class @Nested class BrokerHistoryTests { … service.applyControlCardFilters(…) // line 50, inner class → no edgeNeither the class nor the invoked method names appear among
BrokerHistoryTests' callees; only receivers that are type names (static calls, constructors) resolve.Cause, narrowed by counter-example. Instance-field calls resolve correctly within one class:
AbstractPartnerLogic.partnerRepository.findFirstByHistorySpOptional(…)yieldsIPartnerRepository, confirmed in both directions (33 callers). The failure is specific to the field being declared in the lexically enclosing class — field-type resolution does not cross that boundary.Why it matters.
@Nestedis the standard JUnit 5 layout, so an entire codebase's test → production call graph can be missing while the API reports it as empty rather than unknown. "Which tests cover this service?" and "is this method still used?" both answer wrongly, in the direction that looks like a clean result.The original write-up was too narrow, and the counter-example that produced it did not hold. It read as an enclosing-class problem because
AbstractPartnerLogic.partnerRepository.findFirst(…)resolved correctly — but that class declares the field.PartnerCommonLogic, which inherits it, produces no edge either: theIPartnerRepositoryedge there isINJECTS(propagated byLINK_INJECTS_TO_SUBCLASSES), notMETHOD_CALL. The real rule is any field not declared in this very type, and it splits by what the parser can see:116a — lexically enclosing types (done).
JavaParserbuildsfieldTypesfrom thefindAncestor(ClassOrInterfaceDeclaration)chain, applied outermost-first so an inner class's own field shadows an enclosing one. Purely in-file, so no enrichment involved. This is the@Nestedcase — the standard JUnit 5 layout, which had been costing the entire test→production call graph.116b — inherited fields (done). The parser reads one file and cannot know a supertype's fields, so it no longer drops the call: it records the receiver's identifier on a
field:<name>placeholder edge. The newresolve-inherited-field-receiversenrichment step walks theEXTENDS/IMPLEMENTSchain the graph does know, finds the declaring field, normalises its declared type (generics and package stripped in Cypher, since modules are keyed on the simple name) and re-points the edge.delete-unresolved-field-receiversthen removes the scaffolding, sofield:*never reaches the module namespace. Ordered beforelink-calls-to-implementations, so a recovered edge is still eligible for the polymorphic fan-out.A bug the test caught that review did not. The resolver first ended with a global
... ORDER BY target.sourceFile DESC LIMIT 1, which collapses the whole query to one row — exactly one marker edge resolved per project. On a one-level hierarchy that looks like success. The two-hop assertion inInheritedFieldCallITfailed and exposed it; the fix collapses per marker edge viacollect().Still open: receivers that are neither locals, own fields, enclosing fields nor inherited fields — static imports and chained calls. Those markers are cleaned up rather than resolved.
-
114. A module skipped as a duplicate identity is indistinguishable from one that does not exist, and caller/callee lists drop it without a word (found 2026-08-06,
upmssource-vs-API cross-check; same failure class as 107, one level deeper; step 1 done 2026-08-06)Done: the state is now visible. Each skipped identity gets a marker carrying the conflicting files, so module endpoints answer
409 DUPLICATE_IDENTITYwithdetails.paths, andGET /projects/{p}/duplicates(+ac duplicates) makes the set queryable long after the ingest response that used to hold it. The marker shares theMERGEkey with the reference placeholder, so a referenced duplicate is one node carrying both facts and the more specific one wins — which also fixes the instability described above, where one cause produced404and409 NOT_INGESTEDdepending only on whether anyone happened to call it.What the tests caught that the plan did not. A third sweep:
delete-resolved-placeholderserases every edgeless placeholder, which is exactly what an unreferenced marker is — soDUPEkept answering404while the referencedCALLEDworked. VALIDATE had checked the two file-level sweeps and missed this one. Markers are now excluded from it explicitly.Also fixed on the way: the duplicate paths were absolute server paths (
/tmp/junit-…/DUPE.nat), unlike every other path the API returns. They are relative to the project root now — in the ingest response too, which had the same flaw since the duplicate detection was written.Deliberately not done — step 2. The missing edges stay missing:
JE999 → USIX045Nexists only ifJE999's body is parsed, and it was not. Caller lists can still be short, they just no longer pretend otherwise. Choosing which of the conflicting files wins is a design decision (a configurable precedence such assrc/manualovergenerated_src), not a bug fix, and this entry must not be read as closing that.Symptom.
GET /modules/USIX045N/callersreturns two callers. The source has three: the call ingenerated_src/subprogram/JE999.nat:321is uncommented and real.JE999is not in the graph at all —digest→404 MODULE_NOT_FOUND,search/identifier?name=JE999→[]. Thecallersresponse carries no hint that anything was left out.Cause — a deliberate decision with an undocumented consequence.
JE999has two colliding identities, andingestRootfilters every conflicting path out oftoPersistrather than silently picking one file:MODULE JE999 generated_src/subprogram/JE999.nat vs src/manual/program/JE999.nat DATA_STRUCTURE JE999 …/local_data_area/new/JE999.lda vs …/parameter_data_area/new/JE999.pdaSkipping is right — picking arbitrarily would be worse. Verified as the only colliding stem in the whole walked tree, in both categories, which matches the item-112/113 completion line exactly:
files=6315, persisted=6311, duplicates=2— two identities, four skipped files.Why it matters. It creates a third module state that 107 does not model, and reports it as the first:
| graph state | answer today | truth | |---|---|---| | no node at all |
404 MODULE_NOT_FOUND| correct | | placeholder (sourceFile = "") |409 NOT_INGESTED| correct | | skipped as duplicate |404 MODULE_NOT_FOUND| exists twice, deliberately not ingested |Worse, the answer is not even stable: a duplicate-skipped module that something references gets a placeholder and answers
409, while an unreferenced one answers404. Nothing callsJE999, which is why it vanished entirely. The same cause produces two different HTTP answers.The duplicate list is computed — it rides in the
IngestSummaryof the ingest/refresh response — but it is unreachable afterwards:Duplicateappears in the OpenAPI schema only inside that response, and no endpoint exposes it. After a 17-minute deep refresh nobody has that body.Solution sketch — two genuinely separate problems; the second is the hard one.
- Make the state visible. Persist a marker node for each skipped identity (a placeholder carrying
the conflicting paths), so the module endpoints answer
409 DUPLICATE_IDENTITYwith the paths indetailsinstead of404. AddGET /api/projects/{p}/duplicatesso the set is queryable after the fact rather than only in an ingest response. This alone fixes the misleading status. - The missing edges stay missing. A marker node does not restore the
JE999 → USIX045NCALLNATedge: that edge only exists ifJE999's body is parsed, and it was not. SoUSIX045N/callerswould still be short one entry, just no longer silently. Closing that needs a decision the graph cannot make alone — e.g. a configurable precedence (src/manualovergenerated_src) that ingests one side and flags the module ambiguous. That is a design question, not a bug fix, and must not be smuggled in with step 1.
Scope note: fixing only step 1 leaves caller lists incomplete. That is still a strict improvement — an agent can see that a duplicate exists and go read the sources — but the roadmap entry must not imply the class of bug is closed, exactly as 107's scope note does.
- Make the state visible. Persist a marker node for each skipped identity (a placeholder carrying
the conflicting paths), so the module endpoints answer
-
107. Every module endpoint answers
200with an empty shell for a module that does not exist — indistinguishable from a real but empty module (found 2026-08-02,upmswebservice-layer audit; contradicts the documented404 MODULE_NOT_FOUND) — done 2026-08-05Symptom. Measured against the live server,
upms:GET /modules/WXSPOD0S/digest → 200 {"name":"WXSPOD0S","description":null,"functionCount":0, "callers":{},"callees":{},"dbTables":[],"dataStructures":[]} GET /modules/NOSUCHMOD123/digest → 200 {"name":"NOSUCHMOD123", … identical shell … }Byte-identical answers apart from the echoed name — for a module that exists nowhere in the graph and for a name typed at random.
callees,call-treeandcontextbehave the same (all200).search/identifier?name=WXSPOD0S&type=MODULEcorrectly returns[], so the graph knows; only the module endpoints invent the row.agent-api-system-prompt.mdpromises404 NODE_NOT_FOUND / MODULE_NOT_FOUND — unknown id / module name.Why it matters — this produced a wrong analytical result, not just an ugly response. The question was "can any
W*webservice module reach the commission calculation?". Five dispatchers (WPOLIX0S,WGARCX0S,WACOMX0S,WCLAIX0S,WOBJPX0S) dispatch dynamically to 13 targets (WXSPOD0S,WXSDAD0S,WXSCMD0S,WCARLD0S,WXSCLD0S,WXSCLD2S,WXSFUD0S,WXSGAD0S,WACOMD0S,WCLAID0S,WOBJPD0S,WCATEX0S,WCATED0R).call-tree?depth=4on each returned200with 0 modules and 0 provenance hits, which reads as "analysed, nothing found". The truth is "not analysable" — all 13 source files are absent from the checkout (findfinds none). An agent that trusts the200concludes "these paths trigger no commission processing"; the honest answer is "unknown". Same failure mode as item 103: an incomplete answer that looks complete.Fix (as implemented 2026-08-05).
ModuleIngestStatenow carriessourceFile, so the three graph states are separable:present()(any node),placeholder()(node exists,sourceFile = ""),ingested()(real parsed node; the oldexists(), renamed). Two guards inAnalysisResourcereplace the project-onlywithProjecton every/modules/{name}/…endpoint:withModule→404 MODULE_NOT_FOUNDwhen no node exists at all.withIngestedModule→ the same404, plus409 {status:"NOT_INGESTED", module, detail, nextAction}for a placeholder, reusing the existingDeepIngestRequiredrecord.
Applied to all 15 endpoints whose answer comes from the module's own source (
digest,context,call-tree,callees,db-accesses,workfile-accesses,sql-statements,functions,functions/overrides,functions/{fn}/overrides,functions/{fn}/callers,data-structures,dispatch-table,payload,columns).callersandgraphget the404but stay200for a placeholder — their data comes from the calling modules and is genuine, so the409copy points callers there./modules/{name}/sourcewas already correct. Covered byModuleNotFoundIT(54 cases, including an explicit assertion that the fixture really produces all three states).ac-uimaps both statuses to one banner instead of ~8 per-panel failures; the CLI needed no change (printResponsealready prints any body and exits 1).Deviation from the fix sketched above. The placeholder case returns
409rather than aplaceholder: true/sourceFile: nullpayload marker. A marker cannot be attached to the array-shaped responses (db-accesses,functions,dispatch-table, …), so the marker approach would have fixed the object-shaped endpoints only and left the rest indistinguishable — exactly the gap this item is about. The409is uniform, needs no DTO changes, and carries an actionablenextAction.Scope limit — the class of bug is NOT closed. This guards only the root module of a request. A
call-treethat traverses into placeholder targets still reports that subtree as empty without flagging it, so the originalWPOLIX0S-style dispatcher whose 13 targets are all placeholders still returns200with a silently truncated tree. Callers must still cross-check individual targets (each now answers409). Flagging unanalysable nodes inside a traversal result is item 103's territory. -
75.
CONTAINSis not acyclic — 22 self-loops and 162 two-cycles inupms(fixed 2026-08-23) (found 2026-07-17 while root-causing item 74; cause NOT established — do not treat the notes below as settled). The containment hierarchy that dozens of queries traverse withCONTAINS*contains cycles:(IF @ JX0031N0.nat:781-966) -[:CONTAINS]-> (FOR @ YFRAMBC0.cpy:65-15) (FOR @ YFRAMBC0.cpy:65-15) -[:CONTAINS]-> (IF @ JX0031N0.nat:781-966)Neo4j's variable-length patterns use trail semantics (no relationship repeats in a path), so queries terminate rather than hang — but the blow-up is real: three probes using an unbounded
CONTAINS*over this region were killed at a 2-minute timeout during the item-74 investigation. The shared-node identity below has a second consequence — see item 124. Because a copycode node is MERGEd per(type, name, sourceFile)and thus shared by every includer, item 124's call-edge reap cannot touch the 605 call edges whose source subroutine lives in a.cpy(14 of them stale): reaping them during one module's refresh would delete edges other modules contributed. Items 86 and 106 have the same hole. Whoever gives copycode-resident nodes a per-module identity closes all three at once — worth knowing before designing a fix for either symptom alone.Candidate, unconfirmed: copycode
CONTROL_FLOWnodes are shared by every including module (one node per.cpyline), and a.cpythat opens a block it does not close (YFRAMBC0.cpyopensFORat line 65; theEND-FORlives in the includer) gets containment edges from every includer's nesting context accumulated onto that one shared node. Related: 625 nodes haveendLine < startLine(531DB_ACCESS, 94CONTROL_FLOW) — theFORabove is65 -> 15. But this explains only 26 of 162 cycles and 4 of 22 self-loops, so it is not the main cause. Left open on purpose rather than guessed at. The cycles are not confined to the statement tree: theDATA_STRUCTUREnode named","inBSUPLFN0.nat(itself a parser artifact worth its own look) carries self-loops, so theINCLUDES -> fieldwalk the bare-field resolvers do runs through cyclic ground too. That is why item 77'sINCLUDE_FIELD_DEPTHbound is a correctness requirement, not a tuning knob — an unboundedCONTAINS*there is what killed the probes above. (Item 77's redirect does not traverse this region:srcfor a placeholder edge is only everFUNCTION(157,616 edges) orMODULE(35,572) — neverCONTROL_FLOW— and all 16,231 such functions are directCONTAINSchildren of a module, so its*0..1bound avoids the cycles entirely.)2026-07-17 — concrete impact established and fixed for the dynamic-
CALLNATfamily (pending corpus re-verify). The blow-up is not merely theoretical: a whole-rootdeeprefresh ofupmswedged finalize step 17 (resolve-dynamic-callnat-intra-indirect) for ~2 h without completing, which blocks every later step — including item 77's bare-field resolution, so item 77 could not be corpus-verified. Measured cause: that step joins three unbounded(caller:MODULE)-[:CONTAINS*0..]->anchors, and the resulting per-caller path enumeration over the cyclic copycode region is cubic — a read-only probe of the exact query for a single caller (DAGNTFN0) did not finish in 60 s. Fix: the six dynamic-CALLNATresolvers (RESOLVE_DYNAMIC_CALLNAT_INTRA/…_INDIRECT/…_CROSS- their
…_SCOPEDvariants inCypherQueries) no longer descendCONTAINSto find a module's own statements. AMODULEis 1:1 with itssourceFile(verified: 3587 files, max one module each), and a module'sCALLNATsites andWRITESstatements all carry that samesourceFile, so the anchors becomesourceFile-equality hash-joins that cannot cycle. The dispatch variable in the cross-module resolver can live in an included PDA, so it is scoped to the caller's own file or a data structure the callerINCLUDES(matching by name alone would pull in 20958 unrelated same-named vars; the scope filter keeps the 307 in-scope ones). Proven equivalent onupms:resolve-dynamic-callnat-intrayields the identical 31 resolved(caller, target, lineNo)triples project-wide, and the rewritten indirect step completes project-wide in ~6 s. Existing dynamic-dispatch ITs (intra / indirect / cross / scoped / unresolved- survival) stay green. Still to do: deep-recreateupmsand confirm finalize reaches 36/36; then finish item-77 corpus verification. This does not remove the underlyingCONTAINScycles — otherCONTAINS*traversals remain exposed if a future step joins several of them; the cycles themselves (parser line-range/shared-copycode artifacts) are still open above.
2026-07-18 — same blow-up confirmed on the READ path (frontend-facing, NOT yet fixed). Only the finalize write-queries were rewritten above; the runtime read-queries the UI (
ac-ui) renders still use unbounded(m:MODULE)-[:CONTAINS*0..]->(src). Measured live against server v71 on a heavy module (ACCNPE01):callees>2 min / hangs,digest>10 s timeout,context>10 s timeout,callers~3.6 s; light modules (BMTABBP0) anddb-accesses/call-tree/graph/functionsstay <1.2 s. So the UI's Callees / module-overview / context panels spin on large modules. Same cure as the dynamic-CALLNATfix (src.sourceFile = m.sourceFilehash-join, INCLUDES-scoped for variables). Affected read constants:CALLEES,CALLERS,DIGEST/CONTEXTquery,FUNCTION_CALLERS,DB_ACCESSES, variableREADS/WRITES,*_FOR_MODULES. 2026-07-19 — fixed. Rather than thesourceFilehash-join, the read queries use a tighter, provably equivalent bound: every {@code CALLS}/{@code READS}/{@code WRITES}/{@code DB_ACCESS}-parent edge source is a {@code MODULE} (depth 0) or a {@code FUNCTION} that is a direct {@code CONTAINS} child of the module (depth 1) — verified corpus-wide (0 sources deeper, 0 non-direct-child edge-source functions, and {@code DB_ACCESS} parents are only {@code FUNCTION}/{@code MODULE}, never {@code CONTROL_FLOW}). So(m)-[:CONTAINS*0..]->(src)becomes(m)-[:CONTAINS*0..1]->(src), which returns the identical set but cannot walk the cyclic copycode region. 20 read-side traversals updated (callees,MODULE_HOP_OUT(+ wiring),DISPATCH_TABLE,EGO_NEIGHBORS_*,VARIABLE_ACCESSES,DB_ACCESSES(+FOR_MODULES),SQL_STATEMENTS(+FOR_MODULES),FUNCTION_CALLERS,SEARCH_BY_VALUE(+_CONTAINS),fieldFlow,BUILD_CALLS_MODULE,FLOW_FRONTIER_SOURCE_FILES). Finalize/resolve queries left as-is (they completed). Guarded byReadPathBoundedTraversalIT(EXPLAIN plan asserts no unbounded CONTAINS expand incallees); full IT suite green (193/0/0). Norecreateneeded — a query-only change against the existing graph. Verified live (v76):ACCNPE01 callees>2 min → 0.16 s,digest>10 s → 3.3 s,context**>10 s → 3.1 s;WGEAGB0S calleesunchanged (7). The underlyingCONTAINS` cycles (parser artefacts) still exist, but both the finalize and the read consumers are now bounded — item 75 no longer has a practical impact.2026-08-21 — cause established and half of it fixed (parser artefacts, scope A). The entry above said "cause NOT established"; it is now. There are exactly two families, on unrelated code paths:
- A — parser artefacts (fixed). A Natural filler is declared
<level><byteCount>Xwith no name (489X= level 4, 89 bytes).DATA_AREA_FIELD_EXPORT's occurrence/length column swallows part of the digits, so each filler line produced aDATA_STRUCTUREliterally namedXat a different, bogus level; the(type, name, sourceFile)merge key collapsed all of a file's fillers onto one node, which then contained itself. 27 such nodes inupms— 20 of the 24 self-loops. The","node (BSUPLFN0.nat11x,JB0067N0.nat5x) is the same bug on the.natsource path: the continuation lines of a multi-lineINIT<...>(166 , /*CREDIT NOTE) tokenize as level 166, name",". Not cosmetic — a swallowed level pops the whole group stack, so the fields after a filler were silently re-parented under it: inSNA27R01.pda,C63P0027-GESCHL(line 28) hung underXinstead of its real group. Fix:NaturalParser.DATA_AREA_FILLERskips filler lines on both data-area call sites (it requires at least one count digit, so a field genuinely namedXstill parses), andNaturalFieldTokenizer.opensMultiLineValue()/closesMultiLineValue()let both ingest tiers skipINIT</CONST<continuation lines — the grammar decision lives in the shared tokenizer so the two tiers cannot drift (item 59). Guarded by three tests that fail on the pre-fix parser —NaturalParserTest'sfillerDeclarationProducesNoNodeAndDoesNotReparentFollowingFieldsandmultiLineInitValuesAreNotParsedAsFields, plusNaturalCoarseScannerTest.multiLineInitValuesDoNotEnterTheIdentifierIndex— andNaturalParserTest.singleLineInitIsUnaffected, which pins the common single-lineINIT<...>form against the new skip. The stale nodes leave the graph on the next deep refresh ofupms. - B — shared copycode nodes (open, own round). The candidate above is confirmed:
YFRAMBC0.cpyopens aFORat line 65 and ends at 68 (theEND-FORis in the includer), so the one shared node accumulates every includer's nesting context — hence the 2-cycles withJX0031N0.nat:781/828/874/…and the65 -> 15range. Same inYFRAMBC4/CH/CI/CM.cpyandJX9901C6.cpy;JX0030C2.cpy:37 <-> 41is the intra-file variant. Sized for the first time: 1185 shared.cpynodes, avg 13.7 includers, max 773 -> a per-includer identity (ownerModule, as item 76 does for field placeholders) yields ~16,199 nodes, +3 % onupms's 459,093 — affordable. This also closes items 124/86/106. Drift measured on the current graph (vs. the 2026-07-17 numbers in the heading): self-loops 22 -> 24, two-cycles 162 -> 21, inverted ranges 625 -> 383. Scope A removes 20 self-loops and 2 two-cycles; the item stays open for B, so a corpus-wideCONTAINS-acyclicity assertion is not yet possible — the guards above are deliberately fixture-level.
2026-08-22 — scope B shipped (per-module identity), and it does NOT fix the cycles. Copycode-resident nodes now carry the including module's file as
ownerModule(GraphRepository.nodeOwner, generalising item 76'splaceholderOwner): a node whosesourceFilediffers from the parse's own module file came from a copycode — no extension test needed, sinceCopycodePreprocessoris the only thing that can produce one.MODULE/DB_TABLEare excluded: a module can be declared inside a copycode (ZDTSTBP6inZDTSTBC6.cpy) and every module lookup binds(project, name, sourceFile)but neverownerModule, so an owned MODULE node would be invisible to them. All five sweeps/reaps (DELETE_STALE_FILE_NODES,..._RESOLVED_FIELD_EDGES,..._NATURAL_TABLE_ACCESS_EDGES,..._NATURAL_USING_EDGES,..._NATURAL_CALL_EDGES) now key on the(sourceFile, ownerModule)pair instead of the file alone — mandatory, not cosmetic: the copycode file is in every includer's fresh-file set, so a file-only sweep would delete the other includers' nodes (one transaction older) together with their edges. That pair key also closes the documented scope limit of items 124/86/106: the 605 call edges whose source subroutine lives in a.cpyare now reapable, scoped to the re-parsed owner. Guarded byCopycodeNodeOwnershipIT(4 tests, 3 of which fail on the pre-fix store).Migration caveat, learned the hard way: a deep refresh does not clean up after an identity change. Legacy shared nodes key as
(cpy, "")and no fresh parse produces that pair any more, so nothing sweeps them — after the refreshupmsheld 2217 legacy nodes beside 27,551 new per-owner ones (486,543 nodes). Arecreate(item 78) is required whenever the node identity changes. Afterrecreate?deep=trueof upms/pur/app: 484,327 nodes, legacy shared down to 1 — exactly the excludedMODULE ZDTSTBP6.The cycles survive the recreate, so the current parser produces them — they are not stale-edge residue (a hypothesis that looked plausible because nothing ever reaps
CONTAINSedges, and was wrong). Post-recreate: self-loops 15, two-cycles 16, inverted ranges 1007 (up from 383 — the65 -> 15range defect is no longer folded onto one shared node but exists once per includer: not worse, just no longer hidden). Every remaining cycle is intra-module — the copycode node already belongs to one host and its cycle partners areIFnodes of that same host. Cause:JX0031N0.nat INCLUDE YFRAMBC0 16x JE0018N0.nat INCLUDE YFRAMBCH 4x JB0025N0.nat INCLUDE YFRAMBCI 5xA module that includes the same copycode at many differently-nested sites gets one node for all of them (same file, same line), so the copycode's unterminated
FORaccumulates the nesting context of every site. Per-module identity cannot separate those; per-include-site identity can.Scope C (implemented 2026-08-22, awaiting a corpus recreate) — identity per expansion site.
ownerModulefor a copycode-resident node is now<hostFile>#<includePath>. Not#includedAt, as first planned: item 104 makesincludedAtthe host's INCLUDE line at every nesting level, so all expansions of a member that a copycode includes repeatedly share it —VPARTC02.cpyincludesL4NLOGIC136 times,JB0028C7.cpyincludesISICDAYS4 times.includePath(the fullfile:line>file:linechain) is the only unique site identity; verified present and non-empty on all 27,551 copycode nodes ofupms(max 111 chars, avg 39.7). The key is a parser-set property, so the keys live inac-parser-core'sCopycodePropertiesrather than as magic strings on both sides, and a missing/empty chain falls back to per-module identity (the item 75-B behaviour — a re-collapse, never an invented identity).Proven on a fixture, not yet on the corpus.
CopycodeNodeOwnershipIT'sHOSTBincludes the same copycode at two differently-nested sites — the shapeJX0031N0.nathas 16 times over. On the 75-B store that fixture produces 2 two-cycles; with per-site identity it produces 0, and 3 of the IT's 5 tests fail without the change. The corpus numbers (15 self-loops, 16 two-cycles) can only be confirmed by arecreate?deep=true.Projected cost (recursive include-chain expansion, cycle-guarded, depth-capped at
CopycodePreprocessor.MAX_DEPTH): copycode nodes 27,551 -> ~77,349, i.e.upms484,327 -> ~534,125 (+10.3%). Extremely skewed:JX0030N0.natincludesJX0030C291 times and accounts for +16,650 nodes on its own;XUPD009P.natreachesYFRAMEC2742 times transitively (but that member has 2 nodes). Two things will look worse afterwards and are expected: inverted ranges (endLine < startLine) multiply again from 1007, because each site's node carries the same65 -> 15defect, andJX0030N0will own ~16,800 copycode nodes — itsdigest/context/graphendpoints need a latency check after the recreate.ownerModulealso grows to ~80 chars on ~77k nodes, which is item 111d-2's territory (a hash would kill graph-side diagnosability, so plain text was kept deliberately).2026-08-23 — closed. Corpus-verified after
recreate?deep=trueofupms(server v252, 6311 files,full,incomplete=false): self-loops 15 -> 0, two-cycles 16 -> 0, and a bounded search for longer cycles (CONTAINS*3..6from every multi-parentDATA_STRUCTURE) finds 0 as well.CONTAINSis acyclic on the corpus. Copycode nodes 27,551 -> 51,895 across 19,565 distinct owners, all but one site-keyed (hostFile#includePath); the one exception is the excludedMODULE ZDTSTBP6, by design. Project total 484,327 -> 508,670 (+5.0%), i.e. half the projected +10.3% — the projection expanded every include chain independently, but sites that resolve to the sameincludePathlegitimately share a node. No legacy residue: the only 2 nodes without aningestGenare the unresolved-target placeholdersMODULE/DATA_STRUCTURE JE999(emptysourceFile), which never carry one. Latency on the worst-case moduleJX0030N0.nat(91INCLUDEs ofJX0030C2) is unaffected:digest0.83 s,context0.25 s,graph0.19 s. The corpus-wide acyclicity assertion mentioned above was not added as a test — the IT suite runs against Testcontainers fixtures, not the corpus, so it would have nowhere to live;CopycodeNodeOwnershipIT(5 tests, 3 failing pre-fix) stays the guard, and the corpus number is a measurement recorded here. Inverted ranges rose 1007 -> 1437 exactly as predicted (each site now carries the65 -> 15defect once) — that defect is tracked separately, see item 137.purandappstill hold 75-B-keyed nodes and need the samerecreate?deep=true. - their
-
140. In a project without JPA entities every Java
DB_ACCESSis a false positive (found 2026-08-27 while investigating whyapphas noUSES_TYPEedges; fixed 2026-08-27). ADB_TABLEnode is created only from a JPA@Entityor a Panache active-record class (JavaParser.java:1297). With none in the project, not oneDB_ACCESScandidate can resolve — whileaddDbAccessCandidate(JavaParser.java:1063) over-approximates on purpose: its read gate ismode != READ || isRepositoryReceiverName(...) || staticReceiver || ENTITY_MANAGER_TYPES..., and|| staticReceiveradmits every static call whose method starts withget/find/read/list/count/... In a legacy codebase built on static utility classes that gate filters nothing.How much of the corpus this was:
| project |
DB_TABLE|DB_ACCESS| resolved | unresolved | |---|---:|---:|---:|---:| |upms(natural) | yes | 14,303 | 14,302 | 1 | |pur(java) | 157 | 3,951 | 1,937 | 2,014 (51%) | |ac(java) | 71 | 335 | 114 | 221 (66%) | |app(java) | 0 | 2,219 | 0 | 2,219 (100%) |appcontains no@Entity,@Table,@Repository,JpaRepository,PanacheEntity,@QueryorEntityManageranywhere, and its fourjava.sqlimports are allTimestamp— it has no database access at all. Its 2219 candidates were led byUserContext.getCurrent()(561x),UpmsSessionUtils.getSupportDaten(140x),Config.getInstance()(84x). Natural, by contrast, is exact, because aREAD/FIND/STOREis an access regardless of view resolution.The damage was agent-visible, contrary to the first version of this entry. That version said both endpoints join through the table so unresolved candidates never surface. True for
db-accesses(plainMATCH); false forsql-statements, which usesOPTIONAL MATCHand returned all 76 candidates of...vermittler.VermittlerGeschaeftsregelnas"table": null, "mode": "READ", "statement": "UserContext.getCurrent()".Fixed by a new enrichment step
reap-java-db-access-without-tables(CypherQueries.REAP_JAVA_DB_ACCESS_WITHOUT_TABLES), after the three Java resolvers: if the project holds noDB_TABLE, delete its JavaDB_ACCESSnodes. Java only — the same reaping in Natural would destroy real accesses.Deliberate limits of the chosen rule, all three pinned by
JavaDbAccessNoEntityIT:- Project-level, not per node. It fixes
appand nothing else:pur's 2014 andac's 221 unresolved candidates keep leaking throughsql-statements, because those projects have tables. 2219 of 4454 corpus-wide false positives are removed, roughly half. Whether the rest are false positives or real accesses whose entity lies outside the ingested root is unmeasured — that question, and any tightening of the heuristic itself, is untouched here. - A project whose DB access is exclusively native SQL has no entity, hence no table, so its
accesses are reaped too. They were already unresolvable (
RESOLVE_JAVA_QUERY_NATIVE_SQLmatches an existingDB_TABLE, which only an entity creates), so this loses no working behaviour — but it turns a silent false positive into a silent false negative. Not present in any of the four projects; constructible. - Recovery needs a full refresh. Once such a project gains its first entity the reaper stops
firing, but the deleted nodes only return for files that are actually re-parsed;
changedOnlyskips unchanged ones.
Rejected alternative (the first proposal in this entry): marking the nodes
unresolved: trueinstead of deleting them. Deleting loses the only available measure of how well the Java heuristic aims — forappthat number is now recorded above, but future projects of this shape will not report it. - Project-level, not per node. It fixes
-
99.
DATA_AREA_FIELDmis-split data-area lines that carry a marker column — the field name was lost and the level invented (found and fixed 2026-07-27 while implementing item 98)Symptom. Natural data-area exports (
.lda/.pda/.gda) contain lines with a marker character between the type/length columns and the level. For those lines the parser emits an invented level and uses the marker (or type) letter as the field name; the real field name never reaches the graph at all.Cause.
DATA_AREA_FIELDis^\s*(.*?)(\d)([#A-Za-z][#\w-]*)\s*(.*)$— the prefix is non-greedy, so the regex takes the first digit followed by an identifier as the level. Normally that is right, because the prefix ends in<TYPE><spaces><LENGTH>and<LEVEL><NAME>follows directly:A 60 2##COMMAND -> prefix 'A 60', level 2, name '##COMMAND' (correct)But when a marker is glued to the length, the regex stops too early — inside the length:
A 4C 2#C-PADRE_START-PATTERN parsed : level 4, name 'C' <- '4' is the length, 'C' the constant marker correct : level 2, name '#C-PADRE_START-PATTERN' S0001A 50M 2CRITERIA parsed : level 1, name 'A' <- 'S0001' is the marker column, 'A' the type correct : level 2, name 'CRITERIA'Measured impact (whole
upmscorpus, 2726 data-area files):- 60 files affected, 931 lines mis-split.
- Of those, 333 lines in 31 files produce a phantom group at level 1.
- The invented names are nearly always marker/type letters:
C(739×, the constant marker),A(117×),I(44×),N(20×),M(9×),B,P. - The invented levels come from the length digits and range from
0to6.
Consequences.
- The real field name does not exist.
search_identifiercannot find#C-PADRE_START-PATTERN;data_structure_fieldsshows a field calledCinstead. - Group nesting collapses. The level is arbitrary, and at level
0thegroupStackis emptied completely (while peek().level() >= level) including the root — every following field in the file loses its parent. - The wrapper root disappears.
parseDataAreaonly creates the file-named root whentopLevelNamesreports exactly one top-level group. A phantom level-1 group makes it two, and a module'sUSING <area>then references a node that does not exist — exactly what happens atYLORDVL1.lda:30. - Identifier-index pollution. 739 nodes named
Cacross the corpus, which can collapse together at thesourceFile=""placeholder level.
Relation to item 98. Item 98 was not blocked by this: its enricher deliberately joins on
area.sourceFileinstead of walkingCONTAINSfrom the root, precisely because that root was not guaranteed to exist. That join stays — it is the more robust one regardless.Done. The export is column-oriented —
[<occ>] <TYPE> <LENGTH>[<marker>] <LEVEL><NAME>— where<occ>is an occurrence/superdescriptor column (S0001,0013) glued to the type and<marker>(Cconstant,Mmultiple-value,*) is glued to the length. New anchoredDATA_AREA_FIELD_EXPORTis tried first, with the permissiveDATA_AREA_FIELDkept as a fallback, so anything the anchored form does not recognise keeps its previous behaviour exactly — regression is impossible by construction.topLevelNamesuses the same split, otherwise a phantom level-1 group would still suppress the wrapper root. Deliberately a targeted grammar, not a complete one: the zero-padded two-digit level inA 8*03COD-GENAGREEis left to the fallback, which already resolves it correctly.Measured over all 2726 data areas: 59 282 lines parse identically, 931 are corrected (in 60 files), 1 338 fall back to the previous behaviour verbatim. Tests:
NaturalParserTest#dataAreaMarkerColumnsDoNotStealTheLevelAndName(one case per marker shape) and#dataAreaLinesWithoutAMarkerColumnAreUnchanged(regression guard — the source form, the export view marker and the plain field all depend on regex backtracking past the occurrence column, so they are pinned explicitly rather than assumed).Knock-on:
YLORDVL1.ldaregains a single top-level group, so its wrapper root reappears and itsUSINGreference resolves — item 98's enricher now reachesVDB2-VERSIS_LISTORDERtoo.Open. What the
*marker means is still unknown (the adjacent comment onA 1002* 2V25C6961-RECORDreads/* #01 - alte Länge, hinting at a superseded field). Both the old and the new grammar treat it as a live field, so including it changes no outcome — but if it marks a removed field, those declarations are wrong in the graph either way, which would be its own item. -
98. View aliases declared in a
USINGdata area were reported as tables (2026-07-27, follow-up to item 95). Item 95's alias pre-scan is per-module over the copycode-expanded lines, but aLOCAL USINGdata area is a separate module, so a view declared there stayed unresolved:YGEAGBNH.nat:2617doesFIND (1) VDB2-VERSIS_GENAGREEwith the view declared inYGEAGVL1.lda, sodb-accessesreported the alias instead ofVERSVW_GENAGREE. Root cause was deeper than cross-file scoping:parseDataAreadid not recognise the data-area export view marker at all. In an export a view isV 1VDB2-VERSIS_GENAGREE VERSVW_GENAGREE DA:00,00…— there is noVIEW OFtext, and theVprefix was read as a data type, so the view became aVARIABLEwith noUSES_TYPEto its DDM and no group push, letting its columns escape to the file root. 45 data areas use this form; 32 modules do DML on an alias only declared there. Done: (a)parseDataAreatreats prefixVas a view —DATA_STRUCTURE+USES_TYPEto the DDM (first token of the rest), fields now nesting under it; (b) new enrichment stepsresolve-view-alias-tables READS|WRITES+resolve-view-alias-access-nodesredirect the module'sREADS/WRITESand itsDB_ACCESSnode onto the real table, scoped by the module's ownUSINGset — a name-based redirect would be arbitrary, sinceNEXT-VIEWalone is declared over 100 different tables corpus-wide.size(reals) = 1leaves a contradictoryUSINGset unresolved rather than guessed; the join is onarea.sourceFile, notCONTAINSfrom the area root, because that root only exists when the file has one top-level group (see item 99). Tests:NaturalParserTest#dataAreaExportViewMarkerLinksToTheTableAndNestsItsFields, ITNaturalCrossFileViewAliasIT(incl. two modules resolving the same alias name to different tables); the enricher was verified load-bearing by disabling it and watching the IT go red. -
97.
call-treeleaked the dynamic-call placeholder a manual override only hides (2026-07-27, third WGEAGB0S deep API audit).call-treeforWGEAGB0Slisted#GETSHORT-MODUL— a variable (YGEAGGNH.nat:443,CALLNAT #GETSHORT-MODUL) — as aMODULEin the closure, whilecalleesfor the same module correctly reported only the resolved targetYGEAGGN0. Cause: a manual override does not delete the marker edge to the variable-named placeholder, it setsmanualHidden = trueand relies on the read queries to suppress it (DELETE_DYNAMIC_CALLNAT_PLACEHOLDER_EDGES).callees/callersfilter it; the BFS behindcall-treedid not —MODULE_HOP_OUT/MODULE_HOP_OUT_WIRINGdid not even bind the relationship. Everything driven by that BFS inherited the pollution (graph,db-accesses?depth=N,sql-statements?depth=N). Done: both hop queries bindrand applycoalesce(r.manualHidden, false) = false, matchingcallees/callers. Characterization ITDynamicCallOverrideIT#callTreeHonoursTheOverrideLikeCallees(placeholder present → override → absent → reset → present again); verified red against the pre-fix query. -
96. Natural
UPDATE(ref.)/DELETE(ref.)were dropped, hiding every access layer's write path (2026-07-27, third WGEAGB0S deep API audit). 34 statement sites across 13 of the 65 modules in theWGEAGB0Sclosure — everyY****MN0CRUD module — produced noWRITESedge, sodb-accessesshowed them as read-only plus a singleSTORE. Cause:DB_WRITE's(?!\()guard (added by item 90 to stop a phantom(OLD.)table) suppressed the phantom but never recovered the real table, andDELETEwas only handled in its SQLDELETE FROMform. Done: newDB_WRITE_BY_REFplus a pre-scan mapping eachFIND/READstatement label to its (alias-resolved) table; an unresolvable reference still records nothing, so item 90's no-phantom guarantee holds — its two tests stay green unchanged and now serve as the negative cases. Shared withNaturalCoarseScannerso tier-1 and deep agree. Tests:NaturalParserTest#updateAndDeleteByReferenceResolveToTheEnclosingLoopTable,#byReferenceWriteWithoutAResolvableLoopNamesNoTable,NaturalCoarseScannerTest, ITNaturalViewAliasDbAccessIT. A label may also introduce a SQLSELECTloop rather than aFIND(YELEMMN0,YMULTMN0hold their record that way) — those resolve through theFROMclause;#byReferenceWriteResolvesThroughALabelledSelectLoop. Verified on liveupms: all 34 by-reference sites in theWGEAGB0Sclosure now recorded, 0 missing. -
95. Natural view aliases were reported as DB tables (2026-07-27, third WGEAGB0S deep API audit).
db-accessesnamed the Natural view variable of a DML statement, not the DDM it is declared over: 58 rows across 13 of the 65 modules in theWGEAGB0Sclosure, 32 alias names standing in for 20 real tables. Worst effects — the generator's boilerplate aliasNEXT-VIEWbecame oneDB_TABLEnode shared by 11 modules meaning 11 different tables (and reporting no columns), and 11VDB2-*-VLOGaliases hid every write toVERSVW_LOGFILE, so "who writes the audit log?" answered nothing. Cause:VIEW OFwas only recognised inparseDataArea(.pdafiles);parseModule— which parses every.nat— never built an alias map, and the DML branches passed the operand verbatim todbTable(...). Done:VIEW_DECLpre-scan over the copycode-expanded lines feedsresolveViewAliasinto theDB_WRITE/DB_READbranches;dbTable()now upper-cases (Natural is case-insensitive andDB_TABLEmerges on the name). Shared withNaturalCoarseScannerso a shallow and a FULL module cannot report different names for the same statement. Tests:NaturalParserTest#viewAliasResolvesToTheUnderlyingTable,NaturalCoarseScannerTest, ITNaturalViewAliasDbAccessIT(incl. the same alias in two modules resolving to two tables). Verified on liveupms: alias rows in theWGEAGB0Sclosure 58 → 1, the phantomNEXT-VIEWnode gone,VERSVW_LOGFILEreachable for the first time. The remaining row is the cross-file case, item 98. -
94.
call-treeno longer enumerates paths;followWiringusable again (2026-07-20, JX0034N0 ↔ MultiTableImportJob functional comparison).call-tree?followWiring=truetimed out onpuratdepth ≥ 2(>120s; depth 1 already took 5.3s), which made the Java wiring closure unobtainable. Measured cause — a single quantified path pattern overCALLS|INJECTS|REFERENCES, bounded bymaxDepth × (1 + internalBudget)(= 42 at depth 2), recovering each target's depth asmin(#MODULE nodes on path) - 1, i.e. by enumerating every path. With the CHA-materialized wiring edges (items 31/92) that is combinatorial:| rawBound | Java,
followWiring, depth 2 | NaturalJX0034N0, depth 5 | |---|---|---| | 3 | 1.6s, 71 targets | — | | 4 | 1.7s, 71 targets | 2.6s, truncated (42) | | 6 | 18.9s, 71 targets | 3.2s, truncated (79) | | 8 / 12 | >120s | 2.2s / 2.4s, truncated (120/172) | | 21 | >120s | 3.9s, converged (182) | | 42 (production) | >120s | 3.2s, 182 |So the budget is necessary for Natural (whose result converges only near 21) and useless for Java (converged at 3) — lowering it globally would silently truncate Natural, the exact failure its javadoc warns about. The blow-up comes from the wiring edges, which
JavaParser.addWiringEdgesand the CHA steps only ever emit class-to-class, so they can never reach aFUNCTION. Fix: split the query.MODULErows now come straight from themoduleDepthsBFS (its hop index is the module-hop depth — verified equal to the old query's module set: 71/71 for Java, 46/46 for Natural), and onlyFUNCTIONrows still traverse,CALLS-only and bounded within one module. Behaviour change: the budget can no longer hide a module whose call site sits behind a long internalPERFORMchain —DEPTHLEAFis now reported at depth 1, which also removes a standing contradiction withdb-accesses/sql-statements, whose module set always came from the same BFS.truncatedaccordingly now means "some module's internal subroutine chain may be cut off". Covered byCallTreeTruncationIT(both tests).Measured live after deploy —
call-tree?followWiring=trueonMultiTableImportJob: depth 2 1.9s (was >120s) with the same 71 modules, depth 6 3.0s, depth 10 2.3s converging at 852 modules. Onupms/JX0034N0at depth 5 the module set grew 46 → 53 with nothing lost; the seven that had been hidden areNDBERR,NDBNOERR,USIX009N,USIX052N,USIX053N,YELEMGN0,YLITEMN0. They are real:USIX052NisCALLNATed byISI173N0(line 474), itself a direct callee ofJX0034N0, so it sits at module depth 2; andYLITEMN0was already reported bydb-accesses?depth=5as avia, which is the contradiction this item removes. -
93. Transitive
db-accesseslost theDECLARESrows (2026-07-20, JX0034N0 ↔ MultiTableImportJob functional comparison).DB_ACCESSESresolves a table from three sources —READS/WRITES, an entity's ownMAPS_TO, and a repository'srepositoryEntity(item 32) — but its transitive counterpartDB_ACCESSES_FOR_MODULES(item 65) only ever had the first. The transitive view was therefore not a superset of the direct one: asking the same module withdepthsilently dropped its table. Minimal repro:db-accessesonMultiTableEntryEntityreturnsmulti_table_entry/DECLARES,db-accesses?depth=1on that same module returns[]. Consequence: a Java caller's transitivedb-accessescame back empty even though the entity it persists through maps to a real table, which made the Java side of a Natural↔Java DB comparison impossible to obtain from the API. Fix:DB_ACCESSES_FOR_MODULESnow carries the same three UNION branches, withvianaming the module that declares the table.SQL_STATEMENTShas no such branches, soSQL_STATEMENTS_FOR_MODULESneeded no change (verified). Covered byJavaRepositoryOwnTableIT.entityTableAlsoResolvesInTheTransitiveView/repositoryTableAlsoResolvesInTheTransitiveView(both red before the fix). -
92. Java inheritance/CHA wiring: three defect classes fixed (2026-07-19, MultiTableImportJob deep API audit — manual Java source pass, project
pur). Three systemic errors in the callees/wiring materialization, all found by comparingcalleesagainst source:- A — inherited
INJECTS/REFERENCESlost their origin file (635× in the MTIJ closure, 255 with alineNopast the caller file's end).LINK_REFERENCES_TO_SUBCLASSES/LINK_INJECTS_TO_SUBCLASSES(item 31) copied the base class'slineNoonto the subclass edge but never setoriginFile, so thecalleessites(coalesce(r.originFile, source.sourceFile)) fell back to the subclass file. Worked example:MultiTableImportJobreportsPurBatchJobListener REFERENCES lineNo=133, but that file has 93 lines — line 133 is inAbstractPurBatchJob.java. Same as the Natural copycode bug (items 66/91), for Java inheritance. Fix: materialized edges now carryoriginFile = coalesce(r.originFile, base.sourceFile)+inheritedFrom = base.name. - B — CHA fanned constructor calls out to subtypes (48× phantom
CONSTRUCTORcallees).LINK_CALLS_TO_IMPLEMENTATIONSapplied class-hierarchy analysis tocallKind='CONSTRUCTOR'CALLS, sonew ArrayList<>()produced a phantom→ InputConstraintHolder [CONSTRUCTOR](itextends ArrayList), andnew BaseException()fanned out to every exception subtype. A constructor is statically bound. Fix:AND coalesce(r.callKind,'') <> 'CONSTRUCTOR'. - C — a qualified same-name supertype resolved to self (3× self-
EXTENDS).DateUtils extends org.apache.commons.lang3.time.DateUtils(andNumberUtils/StringUtils) were resolved by simple name to the project's own same-named class → aDateUtils EXTENDS DateUtilsself-loop that also poisoned the inheritance materialization. Fix:JavaParser.supertypeNamekeeps the FQN when a qualified supertype's simple name equals the declaring class's own name; the materializers additionally guardsub <> base. The parser fix stops new self-edges, but a non-wipingrefreshleaves the old self-EXTENDSbehind (both endpoints are the surviving class node, so node reconciliation never sweeps it — the edge gap item 86 closed for Natural), so adelete-self-inheritance-edgesenrichment step reaps any self-EXTENDS/IMPLEMENTSedge project-wide before the inheritance graph is traversed. - Infra: a
delete-synthetic-inheritance-edgesenrichment step reaps allresolvedVia:'INHERITANCE'edges before the three materializers rebuild them, so a non-wipingrefreshpicks up the new properties/gates (otherwiseMERGE ... ON CREATEnever updates a pre-existing edge). Query-only fixes for A/B + reap; parser fix for C. ITJavaInheritanceWiringIT(3 tests). No response-shape change —originFileflows through the existingcalleessites.callSiteFile.
- A — inherited
-
91.
db-accesses/workfile-accesses/sql-statementscarry copycode provenance (2026-07-19, third WGEAGB0S deep API audit — manual source pass). A DB or work-file access whose statement lives in anINCLUDEd copycode was reported with a copycode-locallineNoand no file context, so the number read as a line of the host module. Concretely: the DB2 sequence readSELECT … FROM SYSIBM-SYSDUMMY1lives inUSIX043C.cpyat lines 31/39/45/51/57;db-accessesfor the 9 including modules (YAPRFMN0, YCUACMN0, YLITEMN0, YMODAMN0, YMTABMN0, YMULTMN0, YPRODMN0, YRAMOMN0, YUGRPMN0) reported those as barelineNosthat land on each host's own comment/DEFINE DATAlines. Root cause: the provenance was already on theREADS/WRITESedge (item 66 stampsoriginFile/viaCopycode/includedAton every edge inCopycodePreprocessor.remap, persisted viaSET r += e.properties), but theDB_ACCESSES/WORKFILE_ACCESSES/SQL_STATEMENTSqueries never returned it — exactly the gap item 85 closed forfunctionsand item 66 forcallees/variables. Fix (pure query + DTO, no re-parse of data):db-accessesandworkfile-accessesnow returnsites: [{lineNo, sourceFile, viaCopycode, includedAt}](newAccessSiterecord) alongside the keptlineNos;sql-statementsgainssourceFile+viaCopycodefrom the DB_ACCESS node's (remapped) file.includedAtis stored as a string, so the site queries wrap it intoInteger(...). Direct and transitive (?depth>0,*_FOR_MODULES) variants. REST auto-serializes the records; MCP returns the same DTOs; the CLI is a JSON passthrough — all in sync. ITWorkfileAndCopycodeFunctionIT#dbAccessSiteNamesTheCopycodeFileForCopycodeSourcedAccess(hostFINDvs copycodeFIND→sites[0].sourceFile/viaCopycodedistinguish the two). -
151.
search/identifier?priorityModule=pins the caller's module into the page; deterministic order (renumbered 2026-08-28 — this item had mistakenly also been given the number 91.) (2026-07-26, UI click-to-identify test). Click-to-identify sentsearch/identifier?name=&limit=25, but the query had noORDER BYand paginated in incidental index order, so for a name declared in >25 modules (e.g.#I-LINE-LEV, 192 declarations) the open module's own declaration was truncated away and the popover falsely reported "0 in this module". Fix:SEARCH_IDENTIFIERgains$priorityModule— it does not filter (unlikemodule=) but computes apinRank(0 for that module's file, else 1) andORDER BY pinRank, sourceFile, startLine, so the local match survives thelimitwhile the global list is preserved; ordering is now deterministic (it was undefined before). Delivered across REST (priorityModule), MCPsearch_identifier, CLI--priority-module, and the UI hook. ITIdentifierPriorityModuleIT(four modules sharing one LOCAL field,PRIO_ZZZ_TARGETsorts last: excluded atlimit=2without the pin, first in the page with it, and the full set still returned at a large limit — i.e. no filtering). -
90.
DELETEno longer mis-parsed as a table write (2026-07-19, second WGEAGB0S deep API audit, Finding 5). Natural DMLDELETE [(label)]deletes the current record of the enclosing READ/FIND loop and names no view, and theEXAMINE … DELETE [FIRST]clause is not a DELETE statement at all — butDB_WRITEcaptured the token afterDELETEas a table, producing phantomFROM(108×, from SQLDELETE FROM <table>),(OLD.)/(*)/label refs (19×), andFIRST(7×) — and, for SQL, lost the real table (it sat afterFROM). Fix:DELETEremoved fromDB_WRITE(STORE/UPDATE keep their view operand); a newDB_DELETE_FROMcaptures the SQLDELETE FROM <table>real table (modeDELETE);DELETE (label)/DELETE FIRSTname no table. The same label-reference shape also affectsUPDATE (label)(UPDATE (OLD.)/(HOLD-PRIME.)) and a(*)read operand — a view never starts with(, soDB_WRITE/DB_READnow carry a(?!\\()guard that rejects a parenthesized label reference (no phantom(OLD.)/(*)table) whileUPDATE <view>still records the real view. Both parsers. Tests inNaturalParserTest(DELETE FROM,DELETE (OLD.),EXAMINE … DELETE FIRST,UPDATE (OLD.)). Follow-up (2026-07-19, final verification): one last(*)phantom survived, from the SQL SELECT parser, notDB_READ. A SELECT column list can contain a hyphenated Natural field whose last segment is literallyFROM(YCOMIROW.DAT-CALC-FROM (*));FROM_VIEW = \bFROM\s+(\S+)treated the hyphen as a word boundary, captured the trailing(*)as a phantom table, and — since the FROM view binds on the first match only — swallowed the realFROM VERSVW_COMISIONclause below. Fix:FROM_VIEWnow uses a negative lookbehind(?<![-\\w.])FROM(FROM must be a standalone SQL keyword, not an identifier tail) plus the(?!\\()operand guard. Both parsers. TestNaturalParserTest#sqlSelectColumnEndingInFromDoesNotShadowTheRealFromClause. -
89.
READ WORK <n>(FILE keyword omitted) recognized as work-file I/O (2026-07-19, second WGEAGB0S deep API audit, Finding 4). The item-84 guard only matchedREAD WORK FILE; the corpus also writesREAD WORK 1 ONCE RECORD …(289× project-wide) withoutFILE, which still fell through toDB_READand produced a phantomDB_TABLE 'WORK'(154 accesses). Fix: theWORK [FILE] nguard and theWORKFILE_ACCESS/WORKFILE_DEFINEpatterns now treatFILEas optional (guarded on a following digit, so a view whose name merely starts withWORKis unaffected). Both parsers;NaturalParserTest(READ WORK 1 ONCE RECORD). -
88. Finalize sweep deletes edgeless
DB_TABLE/WORKFILEplaceholder nodes (2026-07-19, second WGEAGB0S deep API audit). Companion to item 86: reaping a stale access edge left the placeholder node (sourceFile="", never node-swept) behind with degree 0 — invisible todb-accesses(edge-driven) but still surfacing insearch_identifier?type=DB_TABLEand the DB-table inventory (observed: orphanedNUMBER/WORK/FIRSTafter items 84/87). New finalize stepdelete-orphaned-placeholder-tables(DELETE_ORPHANED_PLACEHOLDER_TABLES) removes anyDB_TABLE/WORKFILEwith no relationships; degree-0 only, so a table any file still accesses (or a Java@Entity'sMAPS_TOtarget) is kept. Covered byStaleTableEdgeReapIT.orphanedPlaceholderTableNodeIsDeleted. -
87.
FIND NUMBER <view>no longer mis-parsed as a phantomDB_TABLE 'NUMBER'(2026-07-19, second WGEAGB0S deep API audit).FIND NUMBER <view>is a count-only FIND (natural-grammar.md §7.1);NUMBERis a statement keyword, not the accessed view — but theDB_READregex (in bothNaturalParserandNaturalCoarseScanner) captured it as the table name, sodb-accessesreported a bogusNUMBERtable and lost the real view (e.g.CON-DB2-AUTHPROF-USED-IN-AUTHSPC,NEXT-VIEW). Seen on 11 modules of the WGEAGB0S call tree (YAPRFMN0,YCUACMN0,YENTIMN0,YGARAMN0,YLITEMN0,YMODAMN0,YMTABMN0,YPRODMN0,YRAMOMN0,YTABLMN0,YUGRPMN0; 13 accesses; 182FIND NUMBERoccurrences project-wide). Same class as item 84'sREAD WORK FILE. Fix:DB_READnow skips the FIND optionsALL/FIRST/NUMBER/UNIQUEand theRECORDS/IN/FILEnoise words before the view (and allows a variable record-limit(operand), not just a literal). Covered byNaturalParserTest(FIND NUMBER MY_VIEW/FIND NUMBER IN FILE OTHER_VIEW). Item 86 reaps the existingNUMBERedges on the next refresh. -
86. Re-ingest reaps stale Natural
DB_TABLE/WORKFILEaccess edges (self-healing) (2026-07-19, follow-up to items 84/85). A statement whose access target changed between parses orphaned its oldREADS/WRITESedge forever: the target is a placeholder (sourceFile="", never node-swept) and the source node survives, so neither the item-58 node sweep nor the target-keyed edge MERGE reaped it. Seen as theWORKdb-access that lingered onUSIX052Nafter the item-84 parser fix (a pre-fixREAD WORK FILE→DB_TABLE 'WORK'edge), and it applies to any edited view name too. Fix: before re-merging a re-parsed Natural file's edges,DELETE_STALE_NATURAL_TABLE_ACCESS_EDGESdrops itsREADS/WRITESedges toDB_TABLE/WORKFILEplaceholders; the fresh parse (which always re-emits them) re-creates the current ones, unchanged ones round-trip identically. Scoped tolanguage:'natural'source nodes — Java DB edges are resolver-built (RESOLVE_JAVA_DB_ACCESS) and untouched. Covered byStaleTableEdgeReapIT(editREAD VERSVW_OLD→READ VERSVW_NEW+READ WORK FILE, assert old view +WORKgone). (A clean re-ingest ofupmsalready cleared the existing staleWORK; item 86 prevents recurrence on incremental refreshes.) -
84. Natural work-file access tracking (
workfile-accesses) +READ WORK FILEno longer a phantom DB table (2026-07-19, WGEAGB0S deep API audit, Finding 2).READ WORK FILE n <buf>(sequential flat-file I/O) was matched by the(READ|FIND) <view>DB pattern in bothNaturalParserandNaturalCoarseScanner, creating a bogusDB_TABLE 'WORK'READS access (seen onUSIX052Nin the WGEAGB0S call tree). BothDB_READpatterns now negative-lookaheadWORK FILE, andREAD/WRITE WORK FILEare modelled as first-classWORKFILE+WORKFILE_ACCESSnodes (analogue ofDB_TABLE/DB_ACCESS), keyed by work-file number, with the record buffer on theREADS/WRITESedge and theDEFINE WORK FILE n '<name>'physical name on the node. NewGET /modules/{name}/workfile-accesses→[{workFile, physicalName, mode, recordBuffers, lineNos}], MCPworkfile_accesses, CLIac workfile-accesses. Covered byNaturalParserTest(READ + WRITE) and full-stackWorkfileAndCopycodeFunctionIT;mcp-api-usage/system-prompt docs updated. -
85.
/functionsitems carrysourceFile+viaCopycode(copycode-provided subroutine provenance) (2026-07-19, WGEAGB0S deep API audit, Finding 1). A subroutine pulled into a module viaINCLUDEwas listed withdeclaredIn=the including module and the copycode'sstartLine/endLinebut no file, so the lines pointed outside the module's own (shorter) file — e.g.ISIN0019(59-line file) reportedGET-FORMATat 90–120, which actually live inISIC0010.cpy. The FUNCTION node already stored the rightsourceFile; theMODULE_FUNCTIONS/MODULE_FUNCTIONS_OWN/_INHERITEDprojections just dropped it. NowInheritedFunction/FunctionInfoexposesourceFile(+ derivedviaCopycode = f.sourceFile <> m.sourceFile). Covered byWorkfileAndCopycodeFunctionIT; docs updated. -
Module
callersdefault is external-only; noMODULEself-loop (2026-07-19, WGEAGB0S deep API audit). Two coupled defects inCypherQueries.callers(scope): (1) the top-level main body'sPERFORMs originate at theMODULEnode, soscope=internal/default reported the module as its own caller (WGEAGB0S → WGEAGB0S), a self-loopcalleesnever mirrors — fixed withAND caller <> mon the internal scope; (2) the default (scope=null) merged external callers with intra-module PERFORM wiring, so a module's own subroutines showed up as its "callers" — the default now maps toexternal(genuine incoming CALLNAT/inheritance only).context/digest(both callcallers(…, null)) inherit the clean view;scope=internalstill exposes function→function PERFORM wiring; callees unchanged. Covered byModuleCallersSelfLoopIT(fails 2/3 before the fix). MCPcallerstool description + REST endpoint doc +agent-api-usage-ac-implementation.mdupdated. (Supersedes the earlier "Not a bug (verified):context.callersincludes internal PERFORM callers — noisy but accurate" note.) Ego-graphdirection=inforWGEAGB0Snow returns its dynamic callersW-LST-N0/W-MNT-N0(item 75). Payloaddirectionis alwaysREQUESTfor PDA-derived contracts (a single interface PDA doesn't encode direction) — a documented limitation, not a bug.- Follow-up (2026-07-19, WGEAGB0S call-tree closure re-audit): the "external-only" default was only
half-fixed. The default view still returned the calling
FUNCTIONnode (a subroutine/method) whenever the call originated inside a subroutine rather than the main body —ModuleCallersSelfLoopITmissed it because its fixture callerCALLNATs from the main body, where the edge already starts at theMODULE. Measured onupms: 46/56 modules in theWGEAGB0Sclosure hadFUNCTION-typed rows in the defaultcallers, 29/56 had duplicate rows, and hot utilities were unusable (CDRANGEdefaultcallers= 500 rows / 3 distinct FUNCTION names / 0 module callers). Fixed by making the external branch ofCypherQueries.callers(scope)roll every caller up to its owningMODULEvia(callerModule:MODULE)-[:CONTAINS*0..1]->(source)-[r]->(m)andcollect(DISTINCT …)— symmetric with howcalleesanchors its source side. Covered by newCallersRollupIT(caller invokes from inside a subroutine, twice → one rolled-up MODULE row with two aggregated sites, no FUNCTION leak); the oldModuleCallersSelfLoopIT,JavaWiringIT,JavaModulesExtendsFilterITstay green.
- Follow-up (2026-07-19, WGEAGB0S call-tree closure re-audit): the "external-only" default was only
half-fixed. The default view still returned the calling
-
100.
DEFINE DATA ... USING <member>binds by level-1 record name, not by member (file) name (found and fixed 2026-07-28, WGEAGB0S deep API audit tier 2 — 19 of 379USINGsites (5.0%) in the WGEAGB0S call-tree closure are wrong or unresolved)Symptom, two shapes.
- Wrong file.
GET /api/projects/upms/modules/WGEAGB0S/data-structuresreportsGround truth:W-WIF-A2 USING PDA fieldCount=10 src/manual/parameter_data_area/old/W-WIF-A7.pdaWGEAGB0S.nat:47saysPARAMETER USING W-WIF-A2, i.e. memberW-WIF-A2=new/W-WIF-A2.pda(5 fields:P-LINE-TYPE/LEVEL/KEY/VALUE) — exactly the fields the module uses at 386, 732–735 and 1286.old/W-WIF-A7.pdais a different member whose level-1 record was copy-pasted as1W-WIF-A2; it holdsP-REST-*, whichWGEAGB0Sreads fromW-WIF-A1(verified:W-WIF-A1carriesP-REST-FLAG,P-REST-LEVEL-IND,P-REST-POINT-KEY,P-LINE-START,P-LINE-END). So bothsourceFileandfieldCountare wrong. Same shape:BGEAGFN0/USIX052NUSING YFRAMBL0→old/ZFRAMBL0.ldainstead ofnew/YFRAMBL0.lda, andUSIX052NUSING YFRAMBL1→new/ZFRAMBL1.ldainstead ofold/YFRAMBL1.lda. Not cosmetic:YFRAMBL0/ZFRAMBL0andYFRAMBL1/ZFRAMBL1differ in the browse-array boundV(CONST<13>vsCONST<1000>), so an agent reading the wrong twin gets the wrong page size. - Never resolved at all.
USING VLAYERLA,USING USIX020L,USING USIX036LreportsourceFile: null, area: UNKNOWN, fieldCount: 0— although all three.ldafiles exist inside the project root and are ingested. 15 of the 19 affected sites are this shape (10×VLAYERLAinISI173N0/VMULTDN1/VMULTGN1/VMULTMN1..4/VMULTON1/VMULTSN2/ZINELEM1, 2×USIX020LinISI173N0/USIX021N, 3×USIX036LinYGARAMN0/YMODAMN0/YPRODMN0).search/identifiershows only thesourceFile: ""placeholder for each.
Cause, two cooperating places.
NaturalParser.parseDataArea(ac-parser-natural/.../NaturalParser.java, ~line 920) emits the member-named wrapper root only for the single-top-level case:A data area with several level-1 records therefore gets no node named after its member, so aif (topNames.size() == 1 && !topNames.get(0).equalsIgnoreCase(areaName)) { … }USINGof it can never resolve.VLAYERLA.lda(constants),USIX020L.lda(1#C-HM-FUNC, …) andUSIX036L.lda(1#V-ID-TRAN_TAB_FWD, …) are exactly that case. The existing comment claims "multi-top-group areas keep their existing shape (no wrapper, no name collision)" — that decision is what produces shape 2.CypherQueries.resolvePlaceholderTargets(ac-neo4j-store/.../CypherQueries.java, ~line 2650) matches a placeholder purely on(type, name, project). It already excludes module-owned groups (item 74) but has no preference for the node that is the member root of a data-area file, and no tie-break when two files declare the same level-1 name — soUSING W-WIF-A2matches the level-1 node insideW-WIF-A7.pdajust as well as the root ofW-WIF-A2.pda. Inupms42 level-1 names are declared in more than one data-area file, so this is not a one-off.
Fix. (a) In
parseDataArea, emit the member-named root whenever no level-1 record already carries the member name (!topNames.contains(areaName)), parenting every level-1 record under it — so each data-area file contributes exactly one node named after its member. (b) InresolvePlaceholderTargets, forDATA_STRUCTUREplaceholders prefer arealnode that is a data-area member root (real.sourceFileends.lda/.pda/.gdaand its basename equalsreal.name), falling back to the current global name match only when no such candidate exists — so coverage never regresses forUSINGs that have no matching file. Characterization tests:NaturalParserTestcase for a multi-top-level.lda(must yield a member-named root containing all top-level records), plus a Testcontainers IT with two fixture PDAs declaring the same level-1 name (theUSINGmust bind to the file whose member name matches).Done. Both halves landed as described.
NaturalParserTest.multiTopLevelDataAreaStillGetsAMemberNamedRoot…dataAreaWhoseTopLevelAlreadyMatchesTheMemberKeepsItsShapecover the parser;DataAreaMemberResolutionITcovers the graph end to end (DAOTHER.pdacarries a copy-pasted1DAMEMBERrecord,DAMULTI.ldahas several level-1 records and none named after the member).parsesLdaWithMultipleTopLevelStructureswas updated: its "top-level structures have noCONTAINSparent" assertion encoded exactly the behaviour this item changes.
- Wrong file.
-
101.
/data-structures/{name}/fieldssilently unions homonymous definitions from different files (found and fixed 2026-07-28, WGEAGB0S deep API audit)Symptom.
GET /api/projects/upms/data-structures/W-WIF-A2/fieldsreturns 15 fields — the union ofnew/W-WIF-A2.pda(5) andold/W-WIF-A7.pda(10) — with nosourceFileon any row and no parameter to disambiguate. An agent cannot tell that it is looking at two unrelated record layouts merged into one, and will happily "verify" a field that the module it is analysing cannot see.Cause.
CypherQueries.DATA_STRUCTURE_FIELDScollects every same-named definition intocanonandUNWINDs it:MATCH (s0:AstNode {type: 'DATA_STRUCTURE', name: $name, project: $project}) WITH collect(s0) AS defs WITH [d IN defs WHERE d.sourceFile <> ''] AS withFile, defs WITH CASE WHEN size(withFile) > 0 THEN withFile ELSE defs END AS canon UNWIND canon AS sThe Javadoc and
agent-api-system-prompt.mdboth claim the opposite ("scoped to the structure's own definition"), so the contract is documented as something the query does not do.Fix. Return
sourceFileon every row, accept an optional?sourceFile=filter, and when the parameter is absent and more than one definition exists prefer the member-root definition (item 100) rather than the union. Covered by an IT with two fixture areas declaring the same level-1 name.Done.
DataStructureFieldgainedsourceFile(sodb-tables/{name}/columns, which shares the record, returns it too); REST?sourceFile=, MCPdata_structure_fields(sourceFile)and CLIac data-structure-fields --source-filedelivered together. Covered byDataAreaMemberResolutionIT.dataStructureFieldsAreNotUnionedAcrossHomonymousDefinitions(fails before the fix: the decoy record's fields are merged in) and…CanBePinnedToOneSourceFile. -
102.
module_data_structurescollapses homonyms into one row with an arbitrary file and a blendedfieldCount(found and fixed 2026-07-28, WGEAGB0S deep API audit)Symptom. The
W-WIF-A2row onWGEAGB0SreadssourceFile = old/W-WIF-A7.pda, fieldCount = 10— a combination that is wrong even if you accept either candidate as the intended one, because the file and the count can come from different nodes.Cause.
CypherQueries.MODULE_DATA_STRUCTURESgroups byd.nameonly and then reportsWITH d.name AS name, rel, max(fc) AS fieldCount, head([sf IN collect(d.sourceFile) WHERE sf <> '']) AS sourceFilehead(collect(…))is order-dependent (nondeterministic across ingests) andmax(fc)is taken over all homonyms, so the reportedsourceFileandfieldCountneed not describe the same definition.Fix. Group by
(name, sourceFile)and return one row per resolved definition. With item 100 in place this normally collapses back to a single row; where it does not, the agent sees the ambiguity instead of a fabricated blend. Covered by the same IT as item 101.Done. A still-unresolved placeholder is reported only when no resolved definition exists for that
(name, relationship), so thesourceFile: null/area: UNKNOWNrow keeps its meaning. Covered byDataAreaMemberResolutionIT.moduleDataStructuresReportsOneRowPerUsedDefinition(before the fix:[DAMEMBER.pda, DAOTHER.pda]collapsed to one arbitrary row). -
103.
db-accessessilently truncates at the defaultlimit=50— a bare array with notruncatedsignal (found and fixed 2026-07-28, WGEAGB0S deep API audit)Symptom. Measured on
upms:GET /modules/WGEAGB0S/db-accesses?depth=10 → 50 rows GET /modules/WGEAGB0S/db-accesses?depth=10&limit=500 → 64 rowsThe 14 dropped rows hide 7 tables entirely —
VERSVW_MULTILIN,VERSVW_MULTTABL,VERSVW_PRODUCTO,VERSVW_RAMO,VERSVW_TABLAS,VERSVW_USERGRP,VERSVW_USUARIO_NEW. The response is a bare JSON array: no envelope, no total, notruncatedflag, so the caller cannot detect the cut. An agent following the documented Natural playbook (db-accesses?depth=3, nolimit) therefore gets a silently incomplete DB footprint — the single most damaging failure mode for a reengineering or impact analysis, because it looks like a complete answer.Cause.
AnalysisResource.dbAccessesapplieseffectiveLimit(limit), which defaults to 50.sql-statementson the same closure has no such cap (389 rows returned uncapped), so the two endpoints disagree about the same data.workfile-accessesshares the defaulted limit.Fix. Remove the default cap on
db-accesses/workfile-accessesso they matchsql-statements.Done — narrower than first proposed. Only the default cap was removed (
AnalysisResource.uncappedLimit- the same helper in
McpQueryTools); the{items, total, truncated}envelope was not added. Rationale: the defect is silent truncation, i.e. a cut the caller never asked for and cannot detect. An explicitlimitis neither — the caller chose it — so wrapping the response would change the array contract for every REST/MCP/CLI/UI consumer to signal something already known. If atruncatedflag is wanted anyway, it should be a separate item covering all paginated endpoints, not just these two. Covered byIncludeProvenanceAndAccessLimitIT.dbAccessesAreNotSilentlyTruncatedAtFifty(60-table fixture; returns 50 before the fix) and…explicitLimitIsStillHonoured.
- the same helper in
-
104.
includedAtis ambiguous for nested copycode includes — it can point into an intermediate file that the response never names (found and fixed 2026-07-28, WGEAGB0S deep API audit)Symptom.
GET /modules/ISI173N0/calleesreports theYFRAMN04callee asviaCopycode: 'YFRAMC01', includedAt: 27. Line 27 ofISI173N0.natis a comment. The real chain isISI173N0.nat:232 INCLUDE USIX050C 'YFRAMMC1' … USIX050C.cpy:58 INCLUDE &1& (parameterised) YFRAMMC1.cpy:27 INCLUDE YFRAMC01 YFRAMC01.cpy:12 CALLNAT 'YFRAMN04'includedAt: 27is a line inYFRAMMC1.cpy— an intermediate file that appears nowhere in the response. In the single-level case (WGEAGB0S → ADLML02,includedAt: 673= theINCLUDE ISIYESNOstatement) the same field is a line in the module's own file, so the field silently means two different things and the caller cannot tell which. (The resolution itself is correct — the parameterised 3-level expansion is followed properly; only the provenance reporting loses the chain.)Fix. Report the chain rather than one line: make
includedAtalways the line in the module's own file (232 here) and add a fullincludePath: [{sourceFile, lineNo}, …]on thesitesentries ofcallees/callers/db-accesses/workfile-accesses. Covered by an IT over a 2-level include fixture.Done.
CopycodePreprocessor.LineOriginnow carries the host include line unchanged through every nesting level plus the chain as a compactfile:line>file:lineedge property (Neo4j properties cannot hold a list of maps);CallSite/AccessSiteexpose it asList<IncludeStep>.sitesis a nested response object, so MCP and the CLI pick the new field up without a signature change. Covered byIncludeProvenanceAndAccessLimitIT.nestedIncludeReportsHostLineAndTheWholeChain(before the fix:includedAt = 2, the line inside the intermediate.cpy— theISI173N0shape exactly) and…directCallHasNoIncludeChain. -
106. A module's resolved
USINGedges are never reaped, so an old binding survives every refresh (found and fixed 2026-07-28 while re-verifying item 100 against the liveupmsgraph)Symptom. After the item-100 fix had landed and
upmshad been rebuilt and deep-refreshed,WGEAGB0S USING W-WIF-A2reported two rows — the correctnew/W-WIF-A2.pda(5 fields) and the pre-fixold/W-WIF-A7.pda(10 fields). The fix had not failed: the correct edge was created, the wrong one simply was never removed. Measured across the WGEAGB0S closure: the 19 wrong/unresolvedUSINGsites dropped to 5, and all 5 residuals were this shape — a stale edge sitting next to the right one (BGEAGFN0/USIX052N→YFRAMBL0,USIX052N→YFRAMBL1×2,WGEAGB0S→W-WIF-A2).Cause. Item 86 reaps a re-parsed Natural file's stale access edges, but only those pointing at a placeholder (
sourceFile = ""). AUSINGedge is resolved onto a realDATA_STRUCTUREnode byresolvePlaceholderTargets, and from then on nothing deletes it:MERGEonly ever adds. So any binding an older ingest made is permanent. This is not specific to item 100 — plain editing of a module'sDEFINE DATA ... USINGlist leaves the dropped data area attached forever, which is the more common everyday case.Fix.
DELETE_STALE_NATURAL_USING_EDGES, theINCLUDEScounterpart of item 86, run in the same spot (right before the fresh edges are merged). Deleting all of a re-parsed file'sUSINGedges is safe because the fresh parse always re-emits every one of them as a placeholder and the same finalize re-resolves them; unchanged ones round-trip identically. Covered byDataAreaMemberResolutionIT.anEditedUsingDropsTheOldBindingOnRefresh(editUSING DAMEMBER→USING DAOTHER, refresh, assert the old binding is gone — fails before the fix). Ordered last in that class because it mutates the fixture.Note: the fix prevents recurrence; it does not retro-clean a graph that already carries such edges — those disappear on the next refresh of each affected file, since the reap runs per re-parsed file.
upmsstill carried the 5 residuals at the time of writing and needs one more refresh. -
105. A fan-out query that surfaces a data-area file re-ingests it on every call —
search/identifiertook ~60-75 s per lookup (found 2026-07-28 WGEAGB0S deep API audit, root-caused and fixed the same day)Symptom.
GET /search/identifier?name=Xonupms:ZFRAMBL0(1 hit) 56.8 s / 62.2 s / 74.7 s across runs,YFRAMBL1(9 hits) 58.2 s. Both the system prompt and the Natural playbook recommend this endpoint for orientation and for confirming a candidate is a real module; at that latency it cannot be used in a loop, and a batch of lookups exceeds a 2-minute client timeout.Cause — not the query. The first suspicion recorded here (a missing/unusable
(project, name)index) was wrong, and the measurement that settled it is worth keeping:| call | result | |---|---| |
?name=NOSUCHNAME12345(0 hits) | 0.90 s | |?name=ZFRAMBL0(1 hit, a.lda) | 74.7 s |Identical scan work, 80× the latency — so the cost is not the scan. (The scan is a full label scan:
PROFILEshows 2,000,905 DbHits, because the name predicate is wrapped in aCASEthat strips a leading Natural sigil and so cannot useast_node_project_name. But that is ~1 s, and it is the same ~1 s in both rows above. Worth its own item if 1 s ever matters; it is not this bug.)The real cost is in
withFanoutWarm→DeepIngestCoordinator.ensureDeepMany, which deep-ingests the source files a result set surfaced.FULLY_INGESTED_SOURCE_FILESasks for aMODULEnode withingestDepth = 'FULL'— but a Natural data area (.lda/.pda/.gda) produces onlyDATA_STRUCTUREnodes, never aMODULE. So a data area can never be reported as fully ingested, is treated as pending on every call, and is re-warmed forever. Server log for one lookup:07:11:29,595 Ingesting DATA_STRUCTURE ZFRAMBL0, YFRAMBL0 ... [old/ZFRAMBL0.lda] <- 40 ms 07:11:29,635 Finalizing project 'upms' (1 files persisted, scoped deep to 0 modules) 07:11:29,635 Finalize upms (scoped-deep): 45 steps 07:12:31,175 <next request> <- ~60 s in finalizeThe warm itself is trivial; each one drags a whole-project 45-step finalize behind it and then reports "changed", so the caller re-runs its query on top. A repeat lookup re-ingested the same file again — it never converges.
Fix.
DeepIngestCoordinator.ensureDeepManyfilters.lda/.pda/.gdaout of the warm candidate set (isWarmable). Nothing is lost: a data area has no deep tier — both ingest tiers run the sameparseDataArea— so warming one can never add anything to the graph. Applies to everywithFanoutWarmcaller (search/identifier,callers,callees,call-tree), not just this endpoint.Covered by
DataAreaMemberResolutionIT.surfacingADataAreaDoesNotReIngestItOnEveryCall, which asserts the nodeidis stable across two lookups — ids are regenerated on every re-ingest, so a changed id is the re-ingest. It fails before the fix with two different UUIDs. Timing is deliberately not asserted: on a small fixture the finalize is fast, so a wall-clock bound would not reproduce the bug.Note: two further latency questions were surfaced by this and left open on purpose, not folded in: the 2M-DbHit label scan above, and why a scoped 1-file finalize runs all 45 project-wide steps (
scoped deep to 0 modules). The second is the larger prize and affects every incremental ingest.
Natural parser robustness
(Item 59 — shared field-declaration tokenizer — completed 2026-07-15; items 61 — comments parsed as
CALLNAT targets — 62 — data literals as false MODULE call targets — and 63 — CALLNAT matched inside a
string literal — completed 2026-07-16. All moved to x-docs/features.md. The unanchored-CALLNAT
family (#61 comments / #62 data literals / #63 string literals) is closed, and the shared lexical
helpers now live in NaturalLines + NaturalFieldTokenizer so a fix lands in both ingest tiers at
once. Items 120 and 121 below reopened this track: both sit in the copycode argument parser,
which the earlier work never touched. All three — 120, 121 and the 123 found while validating them —
completed 2026-08-07; measured after the deep refresh on 2026-08-09, they recovered 2202
copycode-derived call pairs (7319 → 9521, +30%) with none lost.
Item 123 is the one to re-read before writing the next lexical pattern: the grammar doc that
patterns are supposed to be written from was itself wrong, so following the convention correctly
reproduced the bug.)
-
120.
INCLUDEarguments on continuation lines are never bound — 263 call edges silently absent (found 2026-08-06 in theupmssource-vs-API cross-check, done 2026-08-07)Symptom. A Natural
INCLUDEmay spread its positional arguments over several lines. Only the first line is read, so every argument beyond it stays unbound,&3&survives substitution verbatim, and theCALLNAT &3&it feeds is dropped.VCOMIN50:1241 INCLUDE YFRAMBC8 '"AGNT-CHG-CMP-SP"' VCOMIN50:1242 '"YAGCHBN0"' 'YAGCHKEY' 'YAGCHROW' 'YAGCHPRI' ← &2& lives hereGET /modules/VCOMIN50/callees → YAGCHBN0 absent GET /modules/VCOMIN50/reaches?target=YAGCHBN0 → {"reachable": false} GET /modules/YAGCHBN0/callers → BAGCHFN0 only GET /dynamic-calls/unresolved → no VCOMIN50 entry eitherScale. Of 25 175
INCLUDEstatements inupms, 4697 have continuation lines. Checking every candidate pair against the live API: 263 of 264 module→module call edges are missing, across 154 calling modules and 89 targets. This lands hardest on the browse/access layer, because that is exactly where the idiom is used —YCARPBN1andYPOLIBN1report 0 callers,YCOMIBNHreports 1 while 29 files name it.Why it is the bad kind of wrong. The call is missing from
callees,callers,call-treeandreachesand from/dynamic-calls/unresolved, so nothing anywhere says "not analysed". Same failure class as items 103 and 107: unanalysable reads as "nothing found".Root cause.
CopycodePreprocessor.java:29—INCLUDE_STMTis anchored to a single line — and:74, which passes onlyinc.group(2)toparseArgs. Fix: collect following lines that consist solely of literals/tokens before parsing arguments. Notesubstitute()deliberately leaves unmatched&n&as-is; once continuation lines bind, a still-unbound ref should be reported (as an unresolved dynamic call), never silently dropped.Resolution (2026-08-07). Lookahead added, bounded by the member's own highest
&n&, so it stops as soon as the parameters are satisfied and cannot swallow a following statement's continuation; consumed lines are dropped from the emitted source.CALLNAT_DYNAMICnow accepts&n&, so a still-unbound parameter surfaces in/dynamic-calls/unresolvedinstead of vanishing.Correction to the scale figure above. The "263 edges" is the joint effect of this item with 121 and 123 — this item alone recovers exactly 0. Measured by simulating each fix over all 3587
upmsmodules before implementing: 121 alone+207copycode-derived CALLNAT pairs, 120 alone+0, 120+121+453, 120+121+123+2672, 0 lost in every combination. (A cleaner A/B run after implementation — old parser fromHEADvs new, identical counting rule — measured 7319 → 9521, +2202, 0 lost. The pre-implementation figures used a looser counting rule and a smaller baseline; the honest headline number is +2202.) The example quoted above is itself a 123 case — its target is written'"YAGCHBN0"'— so fixing only what this item describes would have left the motivating symptom untouched. That is why 123 exists; see it below.Guard.
CopycodePreprocessorTest(new — copycode expansion had no unit test at all, which is how two argument-parsing bugs survived: both are invisible unless you inspect the substituted text, and the existing IT only saw whether some edge came out) plusCopycodeExpansionIT#aBrowseIncludeWithMultiLineEscapedArgumentsResolvesItsCallend-to-end. -
121. Natural's doubled-quote escape shifts every positional copycode argument by one (found 2026-08-06 in the
upmssource-vs-API cross-check, done 2026-08-07)Symptom.
'''YAGCHBN0'''is one Natural literal ('YAGCHBN0'). The argument pattern splits it into three —'','YAGCHBN0',''— so&1&binds to an empty string and every later position is off by one. TheCALLNAT &2&then names whatever&1&was: an ADABAS sort key.DAGCHEN0:999 INCLUDE YFRAMBC8 '''AGNT-CHG-CMP-SP''' '''YAGCHBN0''' GET /modules/DAGCHEN0/callees → 'AGNT-CHG-CMP-SP' listed as a MODULE; YAGCHBN0 absentThis also creates a phantom
MODULEnode named after the sort key (sourceFile: "",unresolved: true).Scale. 2308
INCLUDElines across 260 modules use the escape. 37 of the 115 entries in/dynamic-calls/unresolvedare sort keys rather than genuine variables — i.e. a third of that list is this bug, not real dynamic dispatch.Mitigating. The phantom is at least flagged
unresolved: true, and/dynamic-calls/overridesoffers a manual correction path — so nothing here silently claims to be resolved.Root cause.
CopycodePreprocessor.java:33—ARG = '[^']*'|\S+. Needs'(?:''|[^'])*', plus unescaping''→'indequote(). Shares a root cause with item 120; both are the copycode argument parser and are best fixed and tested together.Resolution (2026-08-07). Exactly as diagnosed. In isolation this recovers +207 pairs (see the measurement table under item 120). The phantom
MODULEnodes named after sort keys are gone, which also removes 37 of the 115 entries from/dynamic-calls/unresolved— that list is now dynamic dispatch rather than a third parser artefact. -
123.
CALLNAT "X"— the double-quote string delimiter was never accepted, hiding 2219 calls (found 2026-08-07 while validating items 120/121, done 2026-08-07)Symptom. Natural delimits a string with
'or". TheCALLNATpatterns in both ingest tiers matched only', so a call whose target arrives as a double-quoted literal was not a call at all — and, unlike a dynamic call, it was not reported as unresolved either.How it was found — and why it matters procedurally. It was not found by reading code. Before implementing 120/121 the fixes were simulated over all 3587
upmsmodules, which showed item 120 contributing a delta of exactly 0. Item 120's own headline example passes its target as'"YAGCHBN0"'— the double-quote-inside-single-quote idiom, 7232 uses inupmsagainst 2521 for'''X'''. Without that check, both items would have been implemented, a long deep refresh run, and the motivating symptom found still broken afterwards.Scale. With 120 and 121 in place, accepting
"is what turns a+453recovery into the full one: measured A/B over all 3589upmsmodules, 7319 → 9521 copycode-derived CALLNAT pairs (+2202, +30%), 0 lost. Independently confirmed as a genuine Natural delimiter three ways: 3817 direct= "SRO"-style literals in hand-written modules; the browse idiom's 7232 uses; andNaturalLines.java:55, whose existing javadoc already documented both the"delimiter and the doubled-delimiter escape.The real lesson. Everything needed to write these patterns correctly was already in the repo, in two places, and both were bypassed.
NaturalLines.java:55documented the delimiter and the escape.natural-grammar.md§21.3 gives the correct EBNF — both delimiters, doubled-delimiter escape, repetition count. What misled wasnatural-grammar.md:118, a one-line simplification (character-string = "'" { any-character } "'") carrying only a quiet "refined in Section 21.3" pointer.ARGandCALLNATwere written from the summary line, not the refinement. The durable fix is therefore not new grammar text but making the summary line impossible to use by accident: it now names both delimiters and the escape inline and says to read §21.3 first.This is worth remembering as a research failure rather than a coding one: the authoritative answer existed, was correct, and was one cross-reference away.
Fix. One shared
NaturalLines.CALLNAT_LITERAL+literalTarget(Matcher), used byNaturalCoarseScannerandNaturalParser, so the tiers cannot drift — the same consolidation the #61/#62/#63 family got. Guarded byaCallnatTargetMayUseEitherStringDelimiterin both tiers' tests and byaCallnatInsideAStringLiteralIsStillNotACall, which keeps the widened delimiter from re-opening bug #63. -
124. A full deep refresh does not reap placeholder
MODULEnodes the parser no longer produces — fixed bugs keep answering from the graph (found 2026-08-09 while verifying 120/121/123, done 2026-08-09)Symptom. After the deep refresh that shipped items 120/121/123,
DAGCHEN0/calleesstill lists the item-121 phantomAGNT-CHG-CMP-SPas aMODULE— alongside the now-correctYAGCHBN0, from the same call site.GET /modules/DAGCHEN0/callees YAGCHBN0 CALLNAT YFRAMBC8.cpy:27 includedAt 999 ← correct, new AGNT-CHG-CMP-SP CALLNAT_DYNAMIC YFRAMBC8.cpy:27 includedAt 999 ← stale, unresolved:trueProof it is stale, not re-created. All 5 include sites feeding that copycode are byte-identical (
INCLUDE YFRAMBC8 '''AGNT-CHG-CMP-SP''' '''YAGCHBN0'''+ continuation), and running the currentCopycodePreprocessorover the realDAGCHEN0.natemitsCALLNAT 'YAGCHBN0' YAGCHKEYand zero lines mentioningAGNT-CHG-CMP-SP. The parser cannot produce this edge any more; afullrefresh (6311 files persisted) left it standing.Scale. 29 unresolved placeholder
MODULEnodes carry 511CALLSedges whoseoriginFileis a.cpy. Most are legitimate —NDBERR/NDBNOERR(427 edges) andUSR1009N/USR1023Nare real externals absent from the corpus. The bug's own residue is the sort-key family: 19 names, 45 edges, including one whose name still carries its quotes ('YCOTBMN0'). Small, but it is exactly the wrong 45: they are the artefacts of a bug that is now fixed, and they outlive the fix.Why it matters beyond the count. This is a meta-defect: it means fixing a parser bug does not fully take effect until someone notices the residue and rebuilds from scratch. Every parser fix from here on inherits it, and the residue is indistinguishable from a genuine unresolved call, so nothing flags it. Note the reap must not be naive —
&2&/&3&placeholders (11 edges) are the intended item-120 output and must survive.Where to look. The finalize pass has 48 steps and several reap steps (
delete-resolved-field- contains, the item-62 data-literal cleanup); none of them appears to drop a placeholderMODULEthat no longer has any producing call site in the re-parsed source.Resolution (2026-08-09). The gap was structural and had a precedent:
GraphRepositoryalready reaps a re-parsed Natural file'sREADS/WRITESto access placeholders (item 86) and itsUSINGedges (item 106) before merging the fresh ones. Nobody had done it forCALLS. Item 86's own javadoc states the general cause — "the target placeholder is never swept and the source node survives, so neitherDELETE_STALE_FILE_NODESnor the edge MERGE ever reaped it".Two parts:
DELETE_STALE_NATURAL_CALL_EDGES(per-file, gated onreconcile) andDELETE_ORPHANED_PLACEHOLDER_MODULES(finalize, theMODULEcounterpart of item 88's table sweep).Two boundaries the tests forced, both found by running rather than reasoning:
- Duplicate markers must survive the node sweep. Item 114 records a duplicate identity as a
placeholder, and an unreferenced one is a degree-0 placeholder — exactly the shape the sweep
targets.
MARK_DUPLICATE_IDENTITIESruns before finalize, so the sweep deleted the marker just written and the endpoints fell back to404instead of409 DUPLICATE_IDENTITY. Guarded onduplicatePaths IS NULL. Caught byDuplicateIdentityIT(3 failures). - Resolver-built dynamic edges must not be reaped. The proposal said "reap all of the file's
CALLS". That is wrong: 486CALLNAT_DYNAMICedges point at real modules and 392 of them carry no marker at all (nofolded, noresolvedBy), so nothing but the target's file distinguishes enrichment output from parser output. Reaping them broke the item-37a path-warm — a scoped finalize does not reliably re-resolve an edge whose far side is outside its scope, so the deletion was permanent and a dataflow trace stopped at the dispatch boundary. The predicate is thereforet.sourceFile = "" OR r.callKind <> 'CALLNAT_DYNAMIC': everything the parser re-emits is reaped, the resolvers' own output is not. Caught byAnalysisResourceIT#flowForwardPathWarmCrossesIntoDynamicallyDispatchedCallee.
Scope limit (deliberate). Keyed on the source node's file, so it misses the 605 call edges whose source subroutine is defined inside a copycode — 14 of them stale. Those nodes are MERGEd per
(type, name, sourceFile)and are therefore shared by every including module, so reaping them during one module's refresh would delete edges other modules contributed and never re-create them. That needs a per-module identity for copycode-resident nodes and is a separate item. The same hole exists in items 86 and 106, which use the same key.Guard.
StaleCallEdgeReapIT— three cases, two sabotage-verified. The fixture changes only theCALLNATtarget and keeps the enclosing subroutine, which is precisely why the bug hid behind a green suite:DerivedCallsModuleRefreshITdeletes the whole subroutine, so there theFUNCTIONnode disappears andDETACH DELETEtakes the edge along. - Duplicate markers must survive the node sweep. Item 114 records a duplicate identity as a
placeholder, and an unreferenced one is a degree-0 placeholder — exactly the shape the sweep
targets.
Agent API / MCP tooling gaps
(Done items 52, 53, 54, 56 moved to x-docs/features.md.)
-
122.
dispatch-tablerows carry a barelineNowith no provenance — copycode-derived rows point into the wrong file (found 2026-08-06 in theupmssource-vs-API cross-check, done 2026-08-07)Symptom.
callees,db-accesses,workfile-accessesandfunctionsall carrysourceFile/viaCopycode/includedAt/includePath.dispatch-tablecarries none of it — justlineNo— so a row spliced in from a copycode reports the copycode's local line number as if it were a line of the host module.GET /modules/VCOMIN50/dispatch-table → 26 of 44 rows report lineNo 18 or 20 VCOMIN50:18 * #05 20.10.2011 SAGAPI Erweiterung der Felder ... ← a change-history comment real sites: ISICINDE.cpy:18 and ISICINDI.cpy:20 (included 13× each)An agent that follows the line number lands on a comment in the module header and finds nothing — quietly, with no signal that the pointer was resolved against the wrong file.
Fix. Reuse the existing site/provenance shape rather than inventing a second one; the preprocessor already tracks
LineOriginfor every expanded line, so the data is present and only needs carrying through to the response DTO. Cheap, and it makes the endpoint's rows checkable.Unrelated to item 108 in cause (that one is about which dispatcher idioms are recognised at all), but both are
dispatch-tableand worth touching in one pass.Resolution (2026-08-07).
DispatchEntrygains the same quartet the other site-bearing endpoints already carry —sourceFile,viaCopycode,includedAt,includePath— withDISPATCH_TABLEusing the establishedcoalesce(w.originFile, m.sourceFile)idiom, so a host-local row is unchanged and a copycode-derived one names the.cpy. Additive: every existing field keeps its meaning.MigrationDossier.tsxgains a "from" column and a file-aware line link, so clicking a copycode row opens the copycode rather than the wrong line of the host.ac-clineeded no change —DispatchTableCommandprints the body verbatim (verified, not assumed). Guarded by three cases inNestedDispatchGuardIT, including one asserting the guard chain spans theINCLUDEboundary. -
109.
variables/{name}/writesgives the location but not the written value, so resolving a dispatch needs the source anyway (found 2026-08-02,upmswebservice-layer audit; done 2026-08-06)The data was already in the graph. 452 553 of 466 843 Natural
WRITESedges (97%) carryr.value; the query behind the endpoint simply never selected it. So the fix is one line of Cypher plus three DTO fields — the audit's workaround (open the file at the six line numbers the endpoint had just returned) was never necessary, only unreachable.Two corrections to this item's own fix text. It asked for the value "when the right-hand side is a literal, null otherwise". The property does not work that way: it holds the right-hand side as written —
'YVLOGBN0'(literal),*PROGRAM(system variable),#DISPLAY(1)(indexed),#SELECTED-KEY.NUM-CIS-GC(qualified reference). Filtering to literals would discard the majority, so it is returned raw, exactly asdispatch-tablealready reports the same field. And the indexed case is not an array index but a substring window (item 83), surfaced honestly asassignedSubstrPos/assignedSubstrLenrather than under an invented name.Java carries no value at all — 1125
WRITESedges inpur, 0 with one, because the Java parser never captures the right-hand side.assignedValueis therefore alwaysnullthere, and a barenullreads as "nothing is assigned" rather than "not captured for this language". Documented on the DTO and in the usage guide; the alternative would have been a new silent falsehood of exactly the kind items 107/114 exist to remove.Symptom. Working around item 108:
GET /variables/%23WT-OBJ-PROG/writes?module=WPOLIX0S → [{"function":"INIT-OBJECT-TABLE","sourceFile":"…/WPOLIX0S.nat","lineNo":772, …}, … 6 rows]Six correct write sites, and not one of the six assigned values. Answering "what does this dispatcher dispatch to" therefore requires opening the file and reading lines 772/775/777/780/783/785 — the API narrows the search to the right lines and then stops one step short.
dispatch-tablealready returnsassignedValuefor theDECIDEidiom, so the concept and the field name exist.Fix. Add
assignedValue(and, where the write is indexed,assignedIndex) to thewritesrows, populated when the right-hand side is a literal,nullotherwise. Cheap next to item 108 and useful far beyond it: "which constants does this module put into field X" is a routine question in a reengineering pass. -
110. No reachability query — "can A reach B?" has to be hand-rolled as ~100
callerscalls (found 2026-08-02,upmswebservice-layer audit; done 2026-08-06)GET /modules/{name}/reaches?target=A,B,C&direction=up|down&depth=N→{reachable, paths, truncated}, plusac reaches. The audit's question — does anyW*module reach the commission calculation — now runs as one bounded query:MATCH p = shortestPath((a)-[:CALLS_MODULE*1..6]->(b)) → 25 Pfade, 1,7 sPreviously ~100 HTTP round-trips and a hand-written path reconstruction.
The edge type is a correctness decision, not an optimization. The obvious
-[:CALLS*1..n]->is wrong: Cypher cannot constrain the intermediate nodes of a variable-length pattern, so the path would route throughFUNCTIONnodes and report module reachability where there is none. The materialized module-to-moduleCALLS_MODULE(item 68) is module-level by construction.The guard differs by direction, which is item 107 applied one level deeper: downward the answer comes from this module's own calls, so an un-ingested placeholder must not answer "not reachable" — that would be empty for want of data. Upward the routes are made of the callers' source and are genuine even when the target itself was never parsed.
What
reachable: falsedoes not mean.CALLS_MODULEis built only over resolved calls, so a route through an unresolved dynamicCALLNAT(item 82) is invisible. The answer is "no path over known edges" — stated on the DTO, because the unqualified reading is exactly the false-negative this codebase keeps having to remove.Bounded on purpose (item 75: 22 self-loops, 162 two-cycles in one project);
ReachabilityITcovers the witness path, the absent path, both depth sides of the bound, several targets in one request, the upward direction, a cycle, and the empty-target refusal.Symptom. The question was "does any
W*module reach the commission calculation (ISINCOMI/VCOMIN00/VVERAN50/VCOMIN55/VCOMIN57/VCOMIN50)?" — a yes/no with a witness path. There is no endpoint for it.call-treegoes downward from one root and returns a flat closure without paths, so it answers "what does A reach", never "who reaches B", and never "how". The workaround was a client-side breadth-first search upward over/callers, six seeds, depth 6: ~100 HTTP round-trips, 101 modules visited, and the path reconstruction written by hand.Fix (proposal).
GET /modules/{name}/reaches?target=<name>&direction=up|down&depth=Nreturning{reachable: bool, paths: [[module, …], …], truncated: bool}— or, more useful for this shape of question, a filtered variant ofcallers/call-treethat accepts a set of targets and returns only the witnesses. In Cypher this is one boundedshortestPath/variable-length match; done client-side it is 100 requests and an easy place to introduce a bug. Note the traversal must be bounded — see item 75 on theCONTAINScycles.Why it matters. "Who can trigger X" is the recurring question in legacy reengineering: which entry points reach a calculation, a table write, an external interface. It is the natural counterpart to
call-treeand currently the biggest hole in the query surface for that work. -
82. Manual override for unresolvable dynamic
CALLNATtargets (human/agent-settable) (proposed + implemented 2026-07-19, from the WGEAGB0S deep-API audit; REST + MCP +acCLI + Testcontainers ITs green). The dynamic-CALLNATresolvers cannot follow every name-assembly pattern — e.g.YGEAGGNH.nat:443 CALLNAT #GETSHORT-MODULwhere the name is built viaMOVE 'YGEAGKEY' TO #GETSHORT-MODUL+MOVE 'GN0' TO SUBSTR(#GETSHORT-MODUL,6,3)→YGEAGGN0(a real, ingested module). Such a call site leaves anunresolvedplaceholder (type=MODULE, sourceFile="", name = the variable). Add a REST + MCP +ac-CLI capability to list unresolved dynamic call sites and manually resolve a call site to one or more target modules (multiple targets = deliberate branches, each materialised as a realCALLScallKind=CALLNAT_DYNAMICedge, provenancemanual). Decisions (2026-07-19): callsite key =originFile + lineNo; overrides are persistent (own node type the refresh never deletes) and auto re-applied by an enrichment step after every refresh/deep-refresh; a reset endpoint clears manual overrides (one call site, or all). Related: this is the actionable counterpart to the Bug B consistency gap —callees/digestshould also surface theunresolvedflag thatgraphalready exposes. -
83. Auto-resolver for string-assembled dynamic
CALLNATtargets (SUBSTR/MOVE constant-folding) (proposed 2026-07-19, from the WGEAGB0S deep-API audit follow-up; implemented 2026-07-27). Many unresolved dynamic call sites are in fact statically foldable: the target name is built from literals only, e.g. theY…GNH"GetShort" family (~44 modules) —MOVE 'YxxxxKEY' TO #GETSHORT-MODUL+MOVE 'GN0' TO SUBSTR(#GETSHORT-MODUL,6,3)→YxxxxGN0— plus similar families (PXFRAA01.ST-PGMacross theP…MP0programs, theYFRAMBCK.cpy:27group). Add an enricher that constant-folds a chain ofMOVE <literal>andMOVE <literal> TO SUBSTR(var,pos,len)assignments feeding aCALLNAT varinto the effective target, and resolves the edge automatically when that target is a real ingested module — so item 82's manual override is only needed for genuinely runtime-dependent names, not the deterministic string idioms. Must respect precedence: a manual override (item 82) still wins over an auto-fold. Deliver with characterization ITs (a minimal fixture per idiom) + the usual REST/MCP/CLI-visible effect (fewerunresolvedsites; resolvedcallees/callers). Done:NaturalParserrecordsMOVE '<lit>' TO SUBSTR(var,pos,len)as aWRITESon the base var carryingsubstrPos/substrLen; newRESOLVE_DYNAMIC_CALLNAT_FOLD(+_SCOPED) folds base literal + ordered overlays (reduce/left/substring) → resolvedCALLSedge taggedfolded=true; the direct-literal/indirect/cross resolvers now skip partial-slice writes (substrPos IS NULL). Precedence honoured by guarding the fold against:DynamicCallOverridesites plus adelete-folded-overridden-dynamic-callnatstep beforeapply-manual. Characterization ITs inDynamicCallnatFoldIT(fold resolvesYABALKEY+GN0@6/3→YABALGN0; manual override wins); full dynamic-callnat regression 89/89 green. Docs inagent-api-usage-ac-implementation.md. -
26. MCP session reliability — RESOLVED BY REMOVAL (2026-08-04) (investigated 2026-07-07, reproduced 2026-08-02, never fixed).
mcp__agenticcode__*calls intermittently — and in the 2026-08-02 session, from the very first call — failed with"the first message from the client must be initialize: tools/call", forcing every playbook step to be re-expressed ascurl. The evidence pointed at the MCP client's reconnect handling rather than a server-side bug this codebase's config could fix, and the REST endpoints answered normally throughout. Decision 2026-08-04: the MCP server surface was removed entirely rather than debugged — see "MCP surface removed" inx-docs/features.md. REST +acCLI are now the only access paths; all 40 former tools had a REST twin, so no capability was lost.
Comments as graph data (item 141) — 2026-08-28
-
141. Comments and commented-out code are not indexed — so the documented way to find a Java counterpart returns a confident false negative
Symptom, part one: the counterpart lookup. The consuming project's convention (its
CLAUDE.md§3) is that a reengineered web service carries its Natural origin in a Javadoc block, and that the way to go Natural→Java is to search the program name inpur, where "an empty result means not yet reengineered". Measured:GET /pur/search/value?value=JX0034N0&contains=true → 1 hit, MultiTableImportJob.java:33 GET /pur/search/value?value=WPARTX0S&contains=true → [] GET /pur/search/value?value=WEXKYX0S&contains=true → []The batch job is found because
PROGRAM_IDENTIFIER = "JX0034N0"is a string literal.WPARTX0SandWEXKYX0Sare fully reengineered —PartnerControllerand its logic classes have existed for months — but their origin is recorded only in a Javadoc block:/** * ServiceEndpoint: partner.update.partnercs.Update * UPMSFunction: com.uniqagroup.upms.partner.svc.esp.UpmsPartnerCsUpdate * UpmsObject: PartnerCs, UpmsAdapter: Update */Comments are not nodes, so the search cannot see it. The API therefore answers "not yet reengineered" — the exact wording the consuming project derives from an empty result — for every web service that has already been reengineered. That is not a missing answer; it is the wrong one, delivered with no signal, on the single question every reengineering job opens with. The only reason it has not yet caused a duplicate implementation is that the agent happened to distrust it.
Symptom, part two: the semantics of
upmslive in the comments. Natural source in this corpus carries its change history, its business caveats and its disabled logic as comment text:GET /upms/search/value?value=Bug%20266&contains=true → []while
WAGNTX0S.nat:21reads* #01 09.05.07 VOVBJ03 Bug 266, and lines 22-23 record two regenerations with their ticket numbers. The#01…#05change markers correlate to--> #04/<-- #04blocks that delimit which statements a given change introduced — often the only record of why a branch exists. None of it is queryable. Every such question falls back to reading the file, which in the consuming project is explicitly the exception path and, for the Java side, blocked by a hook.Proposed shape. A
COMMENTnode per contiguous comment block, carryingsourceFile,startLine,endLine,text, and an edge to the nearest following declaration (module, function, field) — so/modules/{name}/commentsanswers "what does this module's header say" andsearch/valuereaches comment text like any other content. Three deliberate points:search/valuemust keep the two kinds separable. A comment hit and a code hit are not the same evidence. Suggestkind: "COMMENT"on the existing row shape (it already discriminatesASSIGNMENT/NODE) rather than a fourth search endpoint, plus a switch for callers that want today's behaviour. Folding comments into the default result set silently would move every existing completeness count — item 131's lesson.- Cost must be measured before it is defaulted on. Natural in this corpus is comment-dense
(
WAGNTX0S.natis ~25% comment lines in its header alone), and item 128 already showed a 3-5× deep-refresh cost for one new edge family. Measure persist time and store growth onupmsbefore deciding whether this belongs in Tier 1 or in the deep tier only. - Javadoc is structured, plain comments are not. The
ServiceEndpoint:/UPMSFunction:/UpmsObject:block above is a key-value list. Parsing it into properties is item 143's business; 141 only has to make the text reachable, and 143 should not be blocked waiting for it.
Implemented 2026-08-28. Comment blocks are graph nodes, opt-in everywhere.
NodeType.COMMENT+EdgeType.DOCUMENTS(ac-parser-core). One node per contiguous block, text invalue, propertiescommentKind/lineCount/truncated(text cut at 4 000 chars). Exactly oneDOCUMENTSedge per block, to the declaration immediately below it, else the one enclosing it, else the module — so a header banner documents theMODULEand a Natural/*on a field's own line documents that field.CommentBlocks/CommentPropertieshold the shared rule so the two parsers cannot drift.- Deliberately not
CONTAINS. That edge is walked by the functions listing, theSEARCH_BY_VALUEassignment arm, the ego-graph and the stale sweep; comments hung off it would leak into queries that never asked for them (item 128'sMENTIONSvsREFERENCESreasoning). - The name carries no prose. A comment node is named
comment@<startLine>;SEARCH_IDENTIFIER_COREmatches every node'snamewith no type filter, so text there would have turned everycontains=trueidentifier lookup into a full-text search. It additionally excludestype = 'COMMENT'outright. search/value?includeComments=true(CLI--include-comments) adds comment hits, labelledkind: "COMMENT". Default output is byte-identical to before — no existing completeness count moves.GET /modules/{name}/comments?kind=(CLIac comments) lists a module's blocks with their target declaration.**SAGdirectives are a separate kind and excluded unless asked for: they are generator metadata, andextractDescriptionalready mines them.- Deep-gated. Comments come from the full parse, not the Tier-1 coarse scan, so the endpoint
deep-ingests on demand and answers
409 NOT_DEEPLY_INGESTEDrather than[]— an empty list that means "not analysed" is the very failure this item was filed about. - Natural specifics. Full-line
*runs group into one block; a trailing/*comment is its own single-line block and is detected quote-aware, soMOVE 'A/*B' TO #Xis not a comment (the olderNaturalLines.stripInlineCommentis still not quote-aware — a separate, pre-existing precision bug in the strip path, filed as item 152). Comments are read from the module's own file, never from copycode-expanded lines: a.cpy's comments belong to the copycode's own module, once, with correct line numbers (items 75-B/75-C). - Tests:
CommentIndexIT(both languages, the opt-in guards, theSAGexclusion), plus parser unit tests for block grouping, adjacency-based targeting and the quote-aware inline scan.
Measured cost (2026-08-28, deep refresh of
upms+puron the real corpus). The item asked for this before defaulting the feature on, and the answer is: comments are not cheap.before after delta upmsnodes508 670 938 561 +429 886 (+84 %), of which every one is a COMMENTupmsdeep refresh977 s 1 727 s +77 % purnodes60 690 (call-graph depth) 83 801 (deep) +23 102 comment nodes Neo4j store (all projects) 2.0 G 2.2 G +10 % By kind:
upms242 454NATURAL_INLINE/ 136 480NATURAL_BANNER/ 50 952SAG(17.6 M chars total);pur15 421LINE/ 7 473JAVADOC/ 208BLOCK(2.8 M chars). 452 989DOCUMENTSedges, zero orphans, zeroendLine < startLine, zero empty texts, 157 blocks truncated at the 4 000-char cap.Consequences, stated rather than buried:
- Comments are 46 % of all
upmsnodes. Every unindexed substring scan (search/identifier,search/value) walks them even when it then filters them out — theMATCHis over:AstNode, the exclusion is a predicate. Measured onupms: acontainsname scan is 2.8 s in Cypher (2.5 s with the exclusion), ~10 s through the paged endpoint. An index or a separate label for comment nodes would fix the scan, and is the obvious follow-up if this becomes the complaint. - The deep tier only — the Tier-1 coarse scan emits no comments, so a plain
refreshpays none of this. That is why/modules/{name}/commentsis deep-gated rather than answering[]. NATURAL_INLINEalone is 242 k nodes (56 % ofupms's comment nodes) for 5 M chars — the per-field/* descriptionidiom. If the cost has to come down, dropping or coalescing those is the first lever, and the one that loses the least: field descriptions are also reachable via the field's own row.
Verified against the real corpus — the three probes the item was filed with:
GET /pur/search/value?value=WPARTX0S&contains=true -> 0 (code only, unchanged) GET /pur/search/value?value=WPARTX0S&contains=true&includeComments=true -> 3 (PartnerController's two origin Javadocs + PartnerCsUpdateLogic) GET /upms/search/value?value=Bug%20266&…&includeComments=true -> 6 GET /pur/search/value?value=WEXKYX0S&…&includeComments=true -> 8 GET /pur/search/identifier?name=UpmsObject&contains=true -> 2 (two real Java fields, no comment leakage) GET /upms/modules/WAGNTX0S/comments -> the header banner, the `#01 … Bug 266` change log, and per-field inline notes
Persist phase: instrumentation and the placeholder sweep — 2026-09-05
-
153. The persist phase was a single opaque number (2026-09-05)
The finalize phase has always logged per step (
runEnrichment: duration, nodes and edges created/deleted). Persist had nothing of the kind — a deep refresh ofupmsspent 724 s of 1 256 s there, spread over 32 batch lines without any breakdown. An extrapolation from the raw MERGE rate explained only ~234 s of it; ~460 s were unattributable.GraphRepository.mergeResultsruns 5-9 statements per batch in one transaction (node merges, up to 17 edge merges perEdgeType, five stale sweeps), none of which were consumed — so neither timings norSummaryCounterswere available.Built:
PersistStatsmeasures every statement (via.consume(), which also yields the counters) plus the Java-side preparation, and logs one aggregated line per batch:Persist deep 200 files: total 8740 ms = prepare 108, merge-nodes 1630, merge-positional-nodes 506, reap-table-access-edges 310, reap-using-edges 161, reap-call-edges 175, merge-edges 1479 (6 types), sweep-stale-file-nodes 132, sweep-stale-placeholders 1887, sweep-resolved-field-edges 278, commit 20Deliberately one line rather than one per statement: per statement it would be ~800 lines per refresh, burying the 51 finalize lines. The residual
commit= total − sum of labels is intentional — an unnamed residual is exactly what this instrumentation is meant to eliminate. The single-file path logs at DEBUG, because a fan-out warm runs through it hundreds of times.Result of the first measurement (upms, deep, 2026-09-05): parsing 18 s, persist 639 s, finalize 567 s.
sweep-stale-placeholders344.9 s = 54 % of the persist phase and 28 % of the whole run. The Java-side preparation, which I had suspected as a possible hidden cost block: 1.5 s. -
154.
DELETE_STALE_PLACEHOLDER_NODESran once per owner instead of once per batch (2026-09-05, found via item 153)The sweep (item 76) did
UNWIND $owners AS o MATCH (n {project, sourceFile: "", ownerModule: o}). No index coversownerModule, so every seek went throughast_node_project_sourcefile— and since all placeholders sharesourceFile = "", every seek returned all 14 551 placeholders of the project and discarded all but a handful. At ~200 owners per batch x 32 batches that is ~6 400 full passes through the same bucket.Fix: one pass, filtered against the list (
WHERE n.ownerModule IN $owners). Semantically identical (same node set;INdeduplicates repeated owners, which is inconsequential for a DELETE), one line of Cypher, no new index.before after sweep-stale-placeholders344.9 s 3.6 s (−99 %) persist phase total 639 s 332 s (−48 %) deep refresh upmstotal1 225 s 956 s (−22 %) Two wrong turns the measurement saved — both plausible, both false:
- For
resolve-bare-includedthe plan looked like "expand first, filter later". The index-first rebuild (seek on(project, name), then check containment) was slower and was still running after 10 minutes — field names are not selective enough project-wide. - For this sweep the obvious suggestion was an index on
(project, ownerModule). The reformulation counter-measured during validation was 3x faster without an index — and thus without a write surcharge on each of the 940 k node merges. An index remains open as the next lever, but has to prove itself against this shape.
The success prediction was wrong too, though in the favourable direction: ~100 s of savings were estimated (assumption: the
DETACH DELETEis irreducible), and it became ~341 s. The cost was almost entirely in the repeated scanning, not in the deleting.Next lever (measured, open): the finalize phase is now the larger item at 607 s, of which
resolve-bare-includedREADS+WRITES is 331 s andlink-args-to-params105 s (0 edges created). Together 436 s = 72 % of the finalize phase. - For
-
155. Enrichment steps can be profiled (
refresh?profile=true) (2026-09-05)The persist instrumentation (item 153) answered "which statement", not "on what within a statement". For the five expensive field steps it therefore remained open whether their time sits in matching or in writing — and a read-only reconstruction cannot answer it, because after the refresh it only sees the leftovers (the fallacy "the step finds nothing" lay ready exactly there).
runEnrichmentoptionally prefixes each step withPROFILEand logs, for every step from 5 s upwards, the five operators with the most db-hits (8 of the 51 steps rather than all). Opt-in per run viaPOST /refresh?profile=trueandac refresh --profile, both explicitly marked DIAGNOSTIC and listed in the agent guide only in the performance section, not in the endpoint table: a profiling switch is a tool for developers, not for agents.Measured overhead: none (908 s profiled against 905 s unprofiled). My warning about 10-30 % was not borne out — the numbers are directly comparable.
Result (upms, deep, 2026-09-05). Every expensive step shows the same pattern: expand wide, then throw almost everything away. Not a single write operator appears in the top 5.
step time dominant operator ratio resolve-bare-includedREADS/WRITES2x ~165 s Filter273 320 729 hits -> 43 753 rows1 : 6 246 VarLengthExpand(All)127 180 800 hits -> 63 017 519 rowslink-args-to-params105 s NodeIndexSeek65 711 041 hits -> 64 653 592 rows, filtered to 18 3711 : 3 518 resolve-field-placeholderREADS/WRITES2x ~50 s Filter50-52 M hits -> ~220 000 rows1 : 229 resolve-view-alias-*(3 steps)3x ~12 s Expand(All)37-43 M hitsFor
link-args-to-paramsthe individual seek is not the problem — it yields 11 estimated rows. It is merely executed millions of times, because the query forms a cross product over the include files of both sides times the argument list (UNWIND cvfilesxUNWIND pvfiles). An index therefore does not help there; the candidate set has to get smaller. -
157.
link-args-to-paramslooked the parameter up once per caller variable instead of once per argument (2026-09-05, found via item 155)The step cost 105 s and created zero edges on a re-refresh; not a single write operator appeared among the five most expensive in the profile. The fan-out measurement explained it:
call sites with arguments 27 240 argument slots 91 954 seeks caller side ( cv)2 509 135 seeks parameter side ( pv)28 868 877 The parameter at position i depends only on
(callee, i), but was looked up inside the caller loop — 28.9 M seeks for 18 371 result pairs.Fix: parameter lookup moved into an aggregating
CALLsubquery per(callee, i), caller side afterwards. In both variants (LINK_ARGS_TO_PARAMSand..._SCOPED, the latter in the interactive deep-ingest path).before after step link-args-to-params105 s 66.8 s (-36 %) deep refresh upmstotal905 s 831 s The read-only advance measurement had predicted 67.2 s — the implementation hit it to within 0.4 s. Of the 74 s total saving, however, only 38 s are attributable to the step; the rest lies in the run-to-run variance of ~10 % that was observable throughout this measurement series.
Two details that carry the reformulation:
- The subquery aggregates (
RETURN collect(pv)). It therefore always returns exactly one row, and an empty list when the parameter is missing, which the guardsize(pvs) > 0discards. A non-aggregating subquery would have swallowed the row — the same result today, but for the wrong reason, and wrong at the next change. - The scoped variant dropped
cmafter the firstWITH. A copy-paste of the whole-root version would have been unable to build the caller file list there — in the interactive path, where a silent null result is hard to notice.
Not done, with reasons: deduplicating the caller side as well. 91 954 argument slots stand against 50 610 distinct
(file, argument name)pairs — a factor of 1.8 for a considerably more complicated query with a re-join. The floor of today's structure is 56.5 s (caller side measured on its own), and the rebuild at 66.8 s is ten seconds away from it.Balance of the performance work (items 153-157): deep refresh
upms1 225 s -> 831 s (-32 %), achieved with two Cypher reformulations. Remaining distribution: persist ~310 s, finalize ~520 s, of whichresolve-bare-includedREADS+WRITES alone is 298 s. - The subquery aggregates (
-
158.
resolve-bare-includedexpanded every data area once per placeholder (2026-09-05, found via item 155)The most expensive enrichment step (147 s + 151 s for READS/WRITES) ran per
(module, placeholder)and expanded the complete subtree of every included data area down to depth 10. The subtree of a data area was thus traversed again once per placeholder of the including module. The profile showed 273 M db-hits for 43 753 result rows and 63 M intermediate rows from theVarLengthExpand.subtree expansions before 380 639 after (once per (module, include))41 377 Fix: the query is driven by the INCLUDE rather than by the placeholder — expand once per
(module, include)first, then match the placeholders onto it by name. In both variants (whole-root and scoped).before after resolve-bare-included READS147.4 s 75.8 s (-49 %) resolve-bare-included WRITES150.6 s 76.8 s (-49 %) deep refresh upmstotal831 s 695 s Edges created identical (
+23 788/+68 936as before), graph unchanged at 940 592 nodes and 2 186 496 edges — the equality evidence on the real corpus, in addition to 411 green ITs (among themGroupQualifiedLeafResolveITfor the qualified path from item 81).Why it is twice as fast, structurally and not by accident:
phis now bound with all four properties of the index(project, sourceFile, type, name). A composite index only applies when every property is bound — the old ordering had no name at that point and could therefore never use the index.The price, knowingly paid: modules with includes but without placeholders now expand in vain — 1 359 of the 3 489 including modules in
upms, ~11 618 additional expansions. Against 339 262 saved ones that is a good trade, and the share shrinks precisely when many placeholders are unresolved, i.e. when the step has real work to do.The advance measurement had estimated 80-120 s of savings; it became 145 s. As with item 154 the estimate was too low, because it treated the write share as irreducible.
Balance of the performance work (items 153-158): deep refresh
upms1 225 s -> 695 s (-43 %), achieved with three Cypher reformulations without a single new index. -
159.
resolve-field-placeholderexpanded the field subtree once per reference (2026-09-05, found via item 155)After item 158 the largest remaining cost in phase D. The qualified field reference (
CDPDA-M.SORT-KEY) was resolved reference-driven: for every one of the ~220 000(ph, phv, src, m)rows the module's INCLUDES list was scanned for the structure name and the whole field subtree ofrealre-expanded to depth 10 — ~50 MFilterhits for ~220 000 result rows.Fix: the resolution
(ph, phv) -> real -> realvruns first, over the project's 485 placeholder fields; the reference edges join on afterwards.realcomes from the(project, type, name)index. TheINCLUDESedge still decides which data area counts — it just no longer findsreal, it checks it.| | | |---|---| | subtree expansions before | ~220 000 (one per row) | | after (once per
(phv, real)) | 520 || | before | after | |---|---|---| |
resolve-field-placeholder READS| ~50 s | 18.2 s | |resolve-field-placeholder WRITES| ~50 s | 22.5 s | | deep refreshupmstotal | 695 s | 628 s |Two clauses are load-bearing, not cosmetic — both verified with
EXPLAIN, because the first attempt without them had no effect at all: theWITH DISTINCTis a planner barrier (without it the planner reverts to the reference-driven order and the hoist is undone), and the INCLUDES test is written asEXISTS { (m)-[:INCLUDES]->(real) }rather than aMATCH(as aMATCHthe planner sourcesmfromreal's includers and builds anm x srccartesian — measurably worse than the starting point).Equivalence evidence: 411 green ITs (among them
QualifiedFieldResolveIT,QualifiedGroupTargetResolveIT,QualifiedWriteReconcileIT,PlaceholderResolveNullLineIT) and, more directly, the old query form finds 0 remaining rows for both READS and WRITES on the graph the new one produced: the new order resolves exactly the same set. The by-name fallback (steps 26/27) stayed at+1/-1and+260/-220, so it picked up no leftovers either.The precondition that can tip: the order only pays while a placeholder's structure name stays selective enough that the index seek does not out-fan the include list it replaces — on
upmsat most 4 real structures share a placeholder name (mean 1.45). The result stays correct either way, since the INCLUDES test still filters; only the cost flips back.The
_SCOPEDvariant is deliberately unchanged — it starts from$namesand processes a handful of modules, where there is nothing to hoist. Whole-root and scoped therefore have different query shapes for the same result.Balance of the performance work (items 153-159): deep refresh
upms1 225 s -> 628 s (-49 %), achieved with four Cypher reformulations and no new index. -
160. The five per-file reap/sweep statements seeked once per
(file, owner)pair (2026-09-05, found by reading the item-153 persist instrumentation)With finalize down to ~316 s, persist (290 s) became the larger half and was measured for the first time. One batch stood out: the last 111 files cost 45.9 s, of which 32.4 s (71 %) went into the five reap/sweep statements — for 1 123 new nodes and 7 467 edges. Batch 3, with 48 154 edges, spent 1.4 s on the same statements.
Cause — the item-154 pattern again, one level up. All five ran as
UNWIND $files AS p MATCH (n {project, sourceFile: p.f, ownerModule: p.o}), but the index is(project, sourceFile). A copycode node carries its expansion site asownerModule, so one file holds many owners' nodes and every seek returned all of them:| | | |---|---| | copycode files (
ownerModule <> "") | 264 | |(sourceFile, ownerModule)pairs they carry | 20 343 | | owners per file: mean / max | 77 / 1 993 (YFRAMEC2.cpy) | | nodes in those files | 53 740 | | rows the seeks requested to reach them | 23 255 522 (1 : 433) |Fix:
$filescarries{f, os: [ownerModule, ...]}— one entry per file — and the query filtersn.ownerModule IN p.os. One seek per file instead of one per pair. No new index: the obvious alternative,(project, sourceFile, ownerModule), would tax the write side of all 940 k node merges, andmerge-nodesis the largest persist item at 69.5 s.Measured read-only on
YFRAMEC2.cpywith its full owner list before implementing: 7 944 098 db-hits pair-driven against 3 986 grouped, both matching the same 3 986 nodes. The objection this shape had to survive was whetherINre-scans the list per node — it does not, Neo4j hashes it.| | before | after | |---|---|---| |
reap-table-access-edges| 17.9 s | 3.5 s | |sweep-resolved-field-edges| 17.5 s | 4.3 s | |sweep-stale-file-nodes| 16.4 s | 1.4 s | |reap-using-edges| 16.3 s | 2.3 s | |reap-call-edges| 16.2 s | 2.3 s | | the five together | 84.3 s | 13.8 s (-84 %) | | persist phase | 289.8 s | 217.9 s | | deep refreshupmstotal | 628 s | 552 s |The 45.9 s outlier batch is gone; the most expensive batch is now 20.6 s and is dominated by
merge-nodesandmerge-edges.Equivalence evidence: every finalize step that consumes what these statements leave behind reports byte-identical counts to the previous run —
resolve-field-placeholder+167 536/-123 279and+206 414/-168 671,resolve-bare-included+23 788/-23 788and+68 936/-68 936,delete-resolved-field-contains-157 667,delete-resolved-placeholders-178 371. Had the reaps deleted too much or too little, these would move. Plus 411 green ITs.A test that asserted nothing was fixed on the way:
StaleFileSweepIndexITpassed{sourceFile, ids}— keys the query has never read — so it checked the plan of a query whose parameters were meaningless. It now passes the real shape and covers all five statements instead of one.Estimate vs. outcome: predicted 45-70 s, measured 72 s in persist / 76 s end to end. Third time in a row the estimate came in low.
Balance of the performance work (items 153-160): deep refresh
upms1 225 s -> 552 s (-55 %), achieved with five Cypher reformulations and no new index. -
162. The view-alias resolvers searched from the access, not from the alias (2026-09-06)
Three finalize steps (
resolve-view-alias-tablesREADS/WRITES andresolve-view-alias-access-nodes) cost ~32 s to make 355 edge changes — the worst ratio of any step.PROFILEshowed why: the planner entered at the 52 762INCLUDESedges and re-expanded each module's wholeCONTAINS+READSsubtree once per include.| operator | rows | db-hits | |---|---|---| | start at
INCLUDES| 52 762 | 158 286 | |(m)-[:CONTAINS*0..1]->(f)| 3 979 711 | 3 979 711 | |(f)-[r:READS]->(alias)| 11 206 268 | 37 650 118 | |Filter alias.type='DB_TABLE'| 111 566 | 11 429 400 |~53 M db-hits, of which the type filter discarded 99 %. The corpus holds only 100 alias declarations (100 distinct names, at most one each), so the fix is to start there and join the accesses on afterwards — the same inversion as items 158-160. The module scoping is untouched: the
INCLUDESedge still decides which data area counts, it is now checked rather than searched.| | before | after | |---|---|---| |
resolve-view-alias-tables READS| 10.9 s | 0.7 s | |resolve-view-alias-tables WRITES| 10.7 s | 0.9 s | |resolve-view-alias-access-nodes| 10.3 s | 0.8 s | | the three together | ~32 s | 2.4 s (-93 %) | | finalize | 298.8 s | 262.9 s | | deep refreshupmstotal | 517 s | 492 s |Equivalence evidence: all three steps report edge and property counts identical to the previous run —
+45/-45props 135,+117/-117props 675,+193/-193props 193 — plus 411 green ITs includingNaturalCrossFileViewAliasITandNaturalViewAliasDbAccessIT.Same two load-bearing clauses as item 159, verified with
EXPLAINfor all three queries:WITH DISTINCTas a planner barrier, and theINCLUDEStest as anEXISTSpredicate rather than aMATCH.An honest limit on the pre-measurement: on the resolved graph the new form short-circuits (no alias
DB_TABLEsurvives finalize), so the read-only 11.6 s vs 0.7 s comparison only proved that the old form pays full price even for zero hits. The real proof was the refresh.An inherited claim that could not be reproduced: the original comment justified module scoping with "
NEXT-VIEWis declared over 100 different tables across the corpus". Measured on upms: 100 declarations over 100 distinct names, none repeated. The scoping was kept exactly as it was precisely because the justification could not be re-verified.Balance of the performance work (items 153-162): deep refresh
upms1 225 s -> 492 s (-60 %), achieved with six Cypher reformulations and no new index. -
163.
resolve-bare-includedsearched from the include side, not from the placeholder (2026-09-06)The step was the largest single item left in finalize at 135.8 s, 52 % of it. Item 158 had already turned it once — from placeholder-driven to include-driven — which halved it. The remaining cost was structural and on the other side of the same trade-off:
| measurement on
upms| value | |---|---| |INCLUDESedges | 52 762 | | distinct data areas behind them | 2 583 | | subtree walks per data area | 20.4x | | rows out of(s)-[:CONTAINS*1..10]->(realv)| 3 264 895 (61 167 paths exist) | | rows surviving therealvfilter, each costing aphindex seek | 2 485 384 | | placeholders on the other side | 21 080, 1 023 distinct names, only 635 with a real field |So the join runs over 635 names but was driven from the 2.49 M side. Turned around:
realvis bound by the(project, name)index, andrealv.sourceFile IN includedFilesprunes the name hits to 6 973 before any subtree walk.Correction (item 167, same day): this entry first said "prunes 944 746 name hits". That was the count after the file filter. The name seek actually returns 9 363 762 rows, ~8.4M of them other modules' placeholders (
sourceFile ""), which the(project, name)index cannot exclude. The prune is a factor of 1 343, not 135 — and reading it as 944 746 is precisely why the file-driven seek of item 167 was not tried here straight away.| | before | after | |---|---|---| |
resolve-bare-included READS| 67.0 s | 43.7 s | |resolve-bare-included WRITES| 68.8 s | 46.5 s | | the step together | 135.8 s | 90.2 s (-34 %) | | finalize | 262.9 s | 225.7 s | | deep refreshupmstotal | 492 s | 457 s |Equivalence evidence: edge counts identical to the previous run in both verification runs —
+23 788/-23 788and+68 936/-68 936— plusBareFieldModuleScopeIT(3) andCopycodeExpansionIT(7) green.The prefilter is semantics-bearing, not tuning. It assumes a field lives in its data area's file. Verified in the sharp form rather than by its consequence:
upmshas zero cross-fileCONTAINSedges out of aDATA_STRUCTURE, project-wide. Reachable (46 057 nodes) is a strict subset of same-file (47 362), so it is a correct prefilter and the exactEXISTSstays — a pure file swap would have over-matched by 1 305 nodes. If the assumption ever breaks, this step resolves less and reports no error, which is why the check is written down here.This approach was recorded as falsified and it was worth re-testing. Item 156 stated that seeking the candidate by
(project, name)first was slower ("still running after 10 minutes, aborted"). That observation is correct and was reproduced today. What was too broad was the conclusion: the name alone is unselective, the name plus the module's include files is not. Two clauses make the difference, both verified withEXPLAIN— theWITH DISTINCTplanner barrier and the file prefilter. Without the barrier Neo4j plans the filter behind theSemiApplyand the old failure reappears exactly.Two measurement mistakes, recorded because they cost a wrong prediction. The read-only pre-measurement said 67 s -> 11 s; the truth is 67 s -> 43.7 s. On an already resolved graph the query short-circuits (6 973 surviving rows against ~92 700 during finalize), so it only ever proved that the old form pays full price for zero hits — the same caveat that applies to items 159 and 162. And the first verification refresh reported 114.8 s for the step and 554 s overall, worse than the 492 s baseline, because the host was loaded: persist, which this change cannot touch, was 65 s slower in it. The delta only became readable after a second run and after checking a step that had not been changed.
Balance of the performance work (items 153-163): deep refresh
upms1 225 s -> 457 s (-63 %), achieved with seven Cypher reformulations and no new index. -
164.
link-args-to-paramshad no index for the callee parameter, and lost every copycode call site ( 2026-09-06)The largest item left in finalize after item 163.
PROFILEshowed two operators carrying almost all of it, and only one of them was suspected beforehand:| operator | rows | db-hits | result | |---|---|---|---| |
NodeIndexSeek pv (project, sourceFile)thenFilter paramPosition| 62 048 947 | 63 068 522 | 29 484 | |NodeIndexSeek cm (project, sourceFile)thenFilter cm.type = 'MODULE'| 11 601 107 | 11 629 527 | 27 225 |The index.
paramPositionwas in no index, so every candidate file was read whole (~140 nodes) and filtered afterwards, once per argument slot per included file (94 379 x ~15.8). A(project, sourceFile, paramPosition)index fixes it, and it is nearly free to maintain because a composite index only holds nodes carrying every property: 1 988 of 938 746 nodes (0.2 %).The caller module.
cmwas found by matching aMODULEinsrc.sourceFile. That threw away 99.8 % of what it read, and it silently skipped every call site inside a copycode: a.cpyhas noMODULEof its own, so the lookup found nothing for 1 195 call sites. The javadoc claimed "a MODULE is 1:1 with its sourceFile" — true for.nat, false for copycode. Reachingcmthrough(cm)-[:CONTAINS*0..1]->(src)fixes both.Result, proven by A/B on the same code, graph and machine minutes apart:
| | with index | without index | |---|---|---| |
link-args-to-params| 25.2 s | 66.6 s | | finalize | 211.7 s | 256.6 s | | deep refreshupms| 458 s | 499 s | |resolve-bare-included(control, untouched) | 53.6 / 57.0 s | 54.2 / 57.2 s |The control step pairs the two runs, so the difference is the index and nothing else. The 66.6 s without it reproduces the original 66.7 s baseline to a tenth of a second. The whole gain is the index; the
cmchange contributes nothing to runtime — a first verification run accidentally shipped thecmfix without the index (the comment was written, theCREATE INDEXline forgotten) and came in at 69.5 s.Correctness.
ARG_TO_PARAMonupmsgoes 17 813 -> 17 871, matching a count computed from the graph before the change. 90 ITs green (AnalysisResourceIT85,AutoInvalidationIT3,JavaCrossClassFlowIT2). Thesize(cms) = 1guard is the conservative bound: a copycodeFUNCTIONcan be one node contained by several modules (item 76 counts 132 inupms), and resolving an argument against the includes of each of them is the cross-module misattribution item 77 had to fix elsewhere. None ofupms's 15 120 call-site nodes is ambiguous today, so the guard drops nothing — it is there so a future parser change cannot turn this into silent bad data.Item 156's note that "no index helps here" was wrong and is corrected there. It holds for the caller side, which already uses the 4-property index correctly; it never applied to the callee side. That is now the second blanket "already tried, does not work" entry this run of work has had to qualify rather than trust — see item 163 for the first.
Two mistakes worth recording. The prediction query bucketed by origin and put that bucket in the
DISTINCTkey, so an edge reachable both ways was counted twice and predicted 17 872 instead of 17 871 — the code was right and the prediction wrong. And a patch script opened the target file with mode"w"before encoding its content; an unencodable character then leftCypherQueries.javaat zero bytes. Encode first, write to a temporary file,os.replacelast.Balance of the performance work (items 153-164): deep refresh
upms1 225 s -> 458 s (-63 %), with eight Cypher reformulations and exactly one new index. -
165. The node merge key carried the copycode expansion site for every node, including the 76 % that never needed it (2026-09-06)
The persist phase was the largest remaining block, and the roadmap listed its three steps as never examined.
MERGE_NODESandMERGE_POSITIONAL_NODESkeyed on(type, name, sourceFile, project, ownerModule)while the widest index has four properties, so the seek stopped at(project, sourceFile, type, name)and aFilterdiscarded the rest. Measured onL4NCOPY.cpy: 75 076 rows read to keep 548.The cost is very unevenly distributed, which is what the previous attempt missed:
| non-positional nodes | keys | nodes | max per key | |---|---|---|---| | module-own | 712 104 | 712 104 | 1 | | copycode-resident | 1 305 | 17 198 | 916 |
So
ownerModuleis redundant in the key for 712 104 nodes and load-bearing for 17 198. The split follows that line: module-own nodes merge on the four-property key (exactly the existing index), copycode-resident ones on a newcopySiteproperty that is written nowhere else, so its index holds 66 445 of 938 746 nodes instead of all of them.Result:
| | before | after | |---|---|---| |
merge-nodes+merge-positional-nodes| 140.1 s | 54.3 s (-61 %) | | deep refreshupms| 458 s | 403 s |Correctness — the point that mattered most here. A mistake in a merge key does not drop an edge, it fuses two distinct nodes or duplicates them. Verified before the change: zero
(project, sourceFile, type, name)keys carry both a module-own and a copycode node, and the own-key is unique (712 104/712 104 non-positional, 160 197/160 197 positional withstartLine). The invariant is inGraphRepository.nodeOwner(). Verified after: node count, copycode node count and edge count identical to baseline (938 746 / 66 445 / 2 703 356) in both verification runs, and 411 ITs green.A migration was required and is easy to miss. The existing graph had no
copySite, so the first persist would have missed every pre-existing copycode node on its key and created a duplicate beside it.BACKFILL_COPY_SITEruns inensureSchema()before the first persist. Two traps found while verifying it:ensureSchema()is subscribed asynchronously, so Quarkus logs "started" while the backfill is still running, andCALL { ... } IN TRANSACTIONSdoes not propagate its inner transactions' counters, sopropertiesSet()reads 0 on a run that just wrote 66 445 values. The log line now reports the re-counted result instead.This index is not free, unlike item 164's. A/B on the same code, graph and machine:
| | with index | without | |---|---|---| |
merge-nodes-copy| 7.7 s | 47.4 s | |merge-positional-nodes-copy| 2.7 s | 39.1 s | |merge-edges| 55.0 s | 40.2 s | |commit| 36.8 s | 26.5 s | | deep refresh | 403 s | 460 s |Re-measured on a quiet host against that same no-index run — controls 49.2/52.7 s against 48.4/52.9 s, so the two pair tightly — the picture sharpens: node merges 120.8 s -> 48.8 s,
merge-edges40.2 -> 47.1 s,commit26.5 -> 32.8 s, refresh 460 s -> 374 s. The index costs ~13 s of write maintenance and saves ~72 s, net -86 s. The ~35 s first recorded here came from a loaded run and overstated the cost. It is still measurably not free — 66 445 indexed nodes are 34x item 164's 1 988 — so "a narrow index costs nothing" does not generalise.The cost is write maintenance, not page-cache pressure. The store had grown 2.4 -> 2.9 GB against an unchanged 2 GB cache, which made cache thrashing the obvious suspect. Measured instead of assumed: 11.2 MB read over an entire refresh, against ~120 MB per run when item 161 sized the cache at a 2.4 GB store. The cache is not the constraint and raising it to 3 GB would have taken a GB from the host for nothing.
Item 156's experiment is now explained rather than merely recorded. The five-property merge-key index made these steps much faster and everything else ~33 % slower. The mechanism: every node carries
ownerModule(""when module-own), so any index over it spans the whole graph. Splitting the key first is what makes the index affordable.A prediction that was wrong, recorded because it was wrong in a new way. From db-hits alone I predicted the split without the index would be worth ~4 % (1-3 s). Measured: 140.1 s -> 120.8 s, i.e. 19 s, or ~10 s once normalised for load. Counting rows read underestimates
MERGE, which also pays locking and comparison per candidate — db-hits are exact but they are not the whole cost model.Balance of the performance work (items 153-165): deep refresh
upms1 225 s -> 374 s (-69 %), with nine Cypher reformulations and two narrow indexes. -
167.
resolve-bare-includedsought the field by name alone, reading 9.4 M rows to keep 6 973 (2026-09-06)Item 163 made the placeholder name the driver, seeking
realvover(project, name). That entry recorded "944 746 name hits" — which was the count after the file filter. The seek actually returns 9 363 762 rows, about 8.4 M of them other modules' placeholders (sourceFile ""), which a name index cannot exclude. Misreading that number is why the obvious next step was not taken sooner.The module's include files are already known at that point, so
realvcan be sought per file over the existing(project, sourceFile, type, name)index instead:| | item 163 | item 167 | |---|---|---| |
realvseek | 9 363 756 db-hits | 6 973 | | following filter | 9 373 474 db-hits | gone | | all operators | ~21.4 M | ~3.6 M | | candidates | 6 973 | 6 973 |It trades 21 074 wide seeks for 660 816 exact ones; the empty ones are nearly free.
| | before | after | |---|---|---| |
resolve-bare-included READS| 49.2 s | 24.7 s | |resolve-bare-included WRITES| 52.7 s | 28.5 s | | the step together | 102.0 s | 53.1 s (-48 %) | | deep refreshupms| 374 s | 336 s |Equivalence: edge counts unchanged (
+23 788/-23 788,+68 936/-68 936), node and edge totals identical (938 746 / 2 703 356), full IT suite green. Control steplink-args-to-params25.2 s against 25.9 s, so the two runs pair despite a loaded start.Prediction wrong again, this time by two. 75-85 s was predicted, with "anything below 70 s is unlikely" stated explicitly; the result was 53.1 s. The error was assuming the ~636 000 property writes formed a large fixed block — the lookup dominated the old runtime as well.
A caveat that belongs with the change: the type list now drives the number of seeks rather than filtering a result, and the seek count is placeholders x include files x types, where one module can have 109 includes. A project with a wide include fan-out and few real hits could prefer the old shape.
Balance of the performance work (items 153-167): deep refresh
upms1 225 s -> 336 s (-73 %). -
169. After item 167 the whole read cost was the sheer number of seeks, and they were 6x redundant (2026-09-06)
With the seek made exact by item 167, measuring where the rest of the time went gave an unusually clean answer: full read side 8 045 ms, prefix alone 7 997 ms — the
EXISTScontainment check, thequalifierGroupcheck and the aggregation cost 48 ms together. Everything was in the seeks.And the seeks repeated: 330 408 (module, included file, placeholder name) triples over only 54 630 distinct (file, name) pairs, a factor of 6.05, because many modules include the same copycode and carry equally named placeholders. Resolving once per pair and joining the modules back on afterwards — the same inversion as items 159, 162 and 163, one level up:
| | before | after | |---|---|---| | read side | 7.5 s | 3.5 s | | db-hits | 6 234 128 | 4 160 044 | |
resolve-bare-included READS| 24.7 s | 18.4 s | |resolve-bare-included WRITES| 28.5 s | 21.7 s | | the step together | 53.1 s | 40.1 s (-24 %) | | finalize | 157.4 s | 146.0 s | | deep refreshupms| 336 s | 329 s |Equivalence was proven as set equality, not as a count. 6 839
(m, ph, realv)triples on both sides with zero difference in either direction — a count alone would not have caught a swap. Then the refresh: edge counts+23 788/-23 788and+68 936/-68 936, totals 938 746 / 2 703 356, 411 ITs green. Control steplink-args-to-params25.9 s against 26.2 s.db-hits understate this kind of saving. They fall by a third while the clock halves, because what is saved is mostly per-seek overhead. Item 165 made the same mistake in the other direction, predicting ~4 % from db-hits where the real figure was ~14 %.
The prediction landed at its pessimistic edge, and the stated caveat is why. 30-40 s was predicted by scaling the read side linearly with the placeholder count, with that assumption flagged as unverified. Result: 40.1 s. The read side halves on a resolved graph but only quarters during finalize.
Two constraints this shape carries, both in the code comment: the type list drives the seek count rather than filtering a result, and
collectmaterialises the hit list (~6 839 here, ~17 000 during finalize), so phase one must finish before phase two starts — a memory bound where there was none.Balance of the performance work (items 153-169): deep refresh
upms1 225 s -> 329 s (-73 %). -
170.
commitin the persist log was a residual, not a measurement (2026-09-06)The roadmap listed
commit(~40 s, the third-largest item) as an "irreducible floor, probably". It was in factMath.max(0, totalMs - attributed)— everything the labelled statements did not account for, which includes the transaction open and the driver overhead. A guess about a number that did not measure what its name said.Two
System.nanoTime()marks inside the transaction lambda split it three ways, keeping the managed transaction's retry semantics untouched. Result:| | | |---|---| |
commit| 31.6 s | |unattributed| 0.0 s | |tx-open| 0.0 s |A representative batch line reads
tx-open 1, unattributed 0, commit 275. The residual was genuine commit all along, so the roadmap's guess was right — it is now measured rather than assumed, and there is no lever here. Third negative result in a row after items 166 and 168.A wrong claim, retracted before it reached the code. Mid-analysis this entry was going to say "two thirds of the residual are unexplained", based on a bench transaction that committed ~270 ms for 60 000 relationships. That bench wrote ~6 MB where a real batch writes ~61 MB (1 955 MB per refresh over ~32 batches) — a factor of ten missed. At the ~49 MB/s the real numbers imply, the residual is exactly what commit I/O should cost. The measurement above confirms it.
A real bias found on the way and fixed:
PersistStatstruncated every individual measurement to whole milliseconds before summing, biasing all ~570 measurements per refresh downwards and pushing ~285 ms into the residual. It now sums in nanoseconds and rounds only when printing.Historical
commitvalues in these docs are not comparable with the ones printed from here on — they were the sum of all three parts. Noted in the javadoc as well.
Styling: theme tokens and the style inventory — item 196 (2026-09-22)
-
196. Styling: theme-token usage and sx/styled inventory
MUI v6 + Emotion: 244
sx={}, 27styled(), 11className, one theme inpur-ui-common/src/theme.ts, threeindex.css(fonts + body reset). Decided scope: theme-token usage and per-component inline inventory; plain.cssonly asMODULEwith aSTYLEper selector.theme.ts→DATA_STRUCTUREwith aFIELDper token (palette.primary.dark,spacing, …); eachsx/styled/styleblock →NodeType.STYLEunder the component with its CSS property keys andREFERENCESto the tokens it uses; hard-coded literals (#005CA9,16px) recorded asliterals. Answers: where is a token used, which tokens are dead, which components bypass the theme, which components overrideheight/zIndex.GET /modules/{name}/styles,GET /projects/{p}/styles/theme-usage?token=;ac styles,ac theme-usage. Static only — no cascade or rendered-layout claims.Implemented 2026-09-22. As planned, with the VALIDATE adjustments: three theme-root forms (
Theme-typed values,{ theme }styled parameters, the theme object imported under any name);usescounts project references only andunusedis documented as "no project reference", never "dead" (MUI consumes tokens itself); tokens the code reads that no theme declares keep their placeholder and are listed withdeclared=false(MUI defaults,spacing, typos);sx={props.sx}is adynamicblock; nested selectors flatten to&:hover.color; compound values yield their literal parts (1px solid #D2D2D2→1px,#D2D2D2). EndpointsGET /theme,GET /theme/{token}/usages,GET /styles; CLIac theme,ac theme-usages,ac styles; CSS rules from the Tier-1 scanner (STYLEper rule). Sidecar contract version 4 (themeTokens,styles,tokenRefs). Sidecar dry run onpur-ui-common: 129 tokens (111 paths, 18 constants), 169 style blocks (87 sx, 55 style, 27 styled; 66 with literals, 42 reading tokens), 89 token reads outside blocks. Verified inStylesITand onpurfe(server 326, recreate + deep refresh): 129 declared tokens, 12 undeclared ones the code reads, 304 style blocks (109 with hard-coded literals, 75 reading tokens),palette.primary.darkread 26 times; no placeholder left except the undeclared tokens. The item-195 inherited-field fix is confirmed on the same run (its 8 placeholders are gone).
Project rename — item 202 (2026-09-23)
-
202. A project cannot be renamed
POST /api/projects/{name}/renamewith{"newName": "..."}(CLIac project rename <old> <new>) rewrites the project key everywhere it lives:projecton everyAstNode(batchedIN TRANSACTIONSlike the delete, implicit transaction), on theDynamicCallOverrides, the entries of other projects'counterpartslists, and finally theProjectshell'sname. Edges carry no project.400 INVALID_REQUESTfor a blank or unchanged name,404for an unknown project,409 PROJECT_EXISTSfor a taken target. The metadata cache is invalidated for both names. Because the nodes move first and the shell last, an interrupted rename is finished by re-running it (the shell still answers to the old name until then). Test:ProjectRenameIT(nodes, callees, override and a peer's counterpart reference follow; old name 404; the three refusals).Found on the way:
DELETE /api/projects/{p}left the project'sDynamicCallOverridenodes behind (10 orphans after deletingupms2). The full delete now removes them; the recreate path (item 78) keeps them on purpose, they are configuration. Asserted at the end ofProjectRenameIT.Measured: renaming
upms(938 746 nodes, 10 overrides) toupms_alttook 97 s on server 333.
Override apply rebuilds CALLS_MODULE; single-INCLUDE programs — items 200, 201 (2026-09-23)
-
200. A dynamic-
CALLNAToverride is applied toCALLSat once, but the derivedCALLS_MODULEedges are not rebuiltFound while replaying the 10 manual overrides of
upmsinto a freshly ingestedupms2: the pinnedCALLSedges appeared at once, butJMIGRUN0kept 1CALLS_MODULEedge instead of 21 until the next refresh, soreachesandfield-flow(the item-68 consumers of the derived edge) did not see the pinned targets. NowupsertDynamicCallOverrideand the reset collect the calling modules of the site (CALLER_MODULES_AT_SITE, same includer fan-out as the apply) and run the scopedDELETE_CALLS_MODULE_SCOPED+BUILD_CALLS_MODULE_SCOPEDin the same transaction; a project-wide reset collects every module owning a manual edge before deleting them. Test:DynamicCallOverrideIT.overrideRebuildsTheDerivedModuleEdgesWithoutARefresh(reachesfalse → true after the override, false again after the reset, no refresh in between). -
201. A program that consists of a single
INCLUDEyields a secondMODULEnode named after the program with the copycode assourceFileCopycodePreprocessor.remapmapped every host node, the module node included, onto the origin of its first expanded line; forZDTSTBP6.nat(INCLUDE ZDTSTBC6/END) that is the copycode, so the deep parse producedMODULE ZDTSTBP6with the.cpyassourceFilenext to the Tier-1 shell keyed on the program file. The module node now always keeps the host file, line 1 to the host's own line count, and noviaCopycodetag. Test:NaturalParserTest.aHostThatStartsWithAnIncludeKeepsItsOwnFileOnTheModuleNode. A graph ingested before the fix keeps its stray.cpy-sourced module until the project is recreated (the item-58 sweep is keyed on pairs the fresh parse produces, and it no longer produces this one); the doc gives the one-line cleanup.
Styling robustness — item 199 (2026-09-22)
-
199. Styling review findings: several themes, repeated token reads, CSS scanner edge cases
From the code review of item 196. (1) The sidecar read only the first
createThemeper file and the resolver demanded exactly one declaring token, so a light/dark pair in one file lost the dark tokens and a pair of theme files left every shared token an unresolved placeholder with?unused=truereporting used tokens as unused. Now every call is read (one fact per token and file, first value wins,variantscounts the themes), and the placeholder resolver redirects a read onto every real token of the name;theme/{token}/usagesandstyles[].tokensdeduplicate the fan-out. (2) Two reads of one token on one line of a style block merged into one edge that kept only the last key; the parser now groups them andpropertyis the comma list. (3) The CSS rule scanner: quotes are tracked so acontent: "{"no longer unbalances the depth counter and drops every following rule; a;at depth zero ends a block-less at-statement so@importno longer leaks into the next selector; a nested rule head inside an at-rule body is stripped before declaration matching (a:hover {was read as propertya); a selector repeated on one line (minified CSS) gets a:colsuffix instead of collapsing.Tests.
TypeScriptCoarseScannerTest.cssRulesSurviveAtStatementsStringsNestingAndMinification, the dark theme and the doubledPRIMARYread in the parser fixture (facts regenerated, contract still 4 —variantsis optional),StylesITwith a second theme file asserting both rows count the read and neither is unused.
Function-level callers across modules — item 197 (2026-09-22)
-
197.
functions/{fn}/callerscannot see cross-module calls — Java and TypeScript alikeA cross-module call (Java cross-class, TypeScript import + call, an item-193 endpoint call from a thunk) is a
MODULE -CALLS-> MODULEedge carryingcallerFnandcalleeMethod; the function-level callers query followed only directFUNCTION -CALLS-> FUNCTIONedges (NaturalPERFORM, same-class Java), so…/AgstammLogic/functions/handleMerge/callersanswered[]although the module-levelcallerslistedAgstammController. (The roadmap's own Java example, a REST controller method, has no Java callers because it is the HTTP entry point; the gap showed on the logic class it calls.)Fix. Query only, no enrichment:
FUNCTION_CALLERSgained a secondUNIONbranch that joins the module edges into the target module oncalleeMethod = callee.nameand resolvescallerFnto the FUNCTION of the calling module, honouringmanualHidden; the same-module branch is untouched. Rows keep theCallRefResponseshape (edgeKind= the edge'scallKind, sites fromlineNo+ the caller module's file). Name matching over-approximates overloads, and a call from top-level code with no enclosing function has no row here (the module-levelcallersstill shows it). REST path andac function-callersunchanged.Test.
FunctionCallersCrossModuleIT: a Java method called from its own class and from another class lists both callers with their lines; a TypeScript function called from a component in another module lists the component.AnalysisResourceIT(Natural PERFORM callers, zero-caller case) stays green.Verified on
pur/purfe(server 328, no re-ingest):AgstammLogic.handleMerge→mergeBrokerat line 98;GeneralAgreementUiControllerEndpoint.createNew→ the thunk ingeneralAgreementSliceat line 92 — both empty before.
Stale parsed edges are reaped on a deep refresh — item 198 (2026-09-22)
-
198. Stale edges to placeholder modules survive a re-parse for Java and TypeScript
Found while verifying item 193 on
purfe: after the sidecar fix that mapspur-ui-common/dist/xto…/src/x, a deep refresh still showed 1 022CALLS/REFERENCESedges into 46distplaceholders next to the freshsrcedges. The per-file reconcile (item 58) sweeps stale nodes of a re-parsed file, but the stale-edge reaps were Natural-only, so an edge from a surviving module that the new parse no longer produces lived forever — and its placeholder, having an edge, escaped the placeholder sweep. Only recreating the project cleared it.Fix.
mergeEdgesBatchstampsr.ingestGen = $ingestGenon every parser-emitted edge (a re-emitted edge is re-stamped through its MERGE key). A new language-agnostic stepreap-stale-parsed-edges(CypherQueries.DELETE_STALE_PARSED_EDGES) runs aftermerge-edgesand beforesweep-stale-file-nodes, keyed on the item-160(sourceFile, ownerModule)pairs of the re-parsed files, and deletes every edge from those nodes whose stamp is older than the run's. Deep only (reconcile), like the node sweep. Finalize-built edges carry no stamp unless a resolver copied it from a parser edge, and the deep finalize that follows rebuilds those. The three Natural reaps stay (they run before the merge and gate on statement kinds). Pre-existing edges without a stamp are never reaped: the first deep refresh after the upgrade stamps, the second reaps — no project recreation needed any more; the usage doc's "recreate after a parser change" note is retired.Test.
StaleParsedEdgeReapIT: a TypeScript component retargeted frombtoclosesapp/src/bfromcalleeswhile an untouched file keeps it; a retarget onto a missing./missing/dmints a placeholder that disappears once the import is retargeted again; the same for a Java class switching its call target fromBtoC. Existing reap ITs (StaleCallEdgeReapIT,StaleTableEdgeReapIT,RefreshReconciliationIT) and the whole server IT suite stay green.
DTO field bindings — item 195 (2026-09-22)
-
195. DTO field binding: which component reads/writes which backend field
No mappers exist:
*UseCaseDTOs sit verbatim in Redux asSvcResult<T>. Bindings are typed path expressions (<SmartInput field={AgstammUseCaseField.broker.ebene}/>, generatedFieldsclasses) resolved by the sidecar to the dotted pathbroker.ebeneand linked to theFIELDof the DTO interface;pur-r-vbuchbinds lodash paths into the whole state.SmartInput→WRITES,SmartOutputand plain reads →READS, from the componentFUNCTIONto theFIELD. With 193'sCOUNTERPART_OFon theDATA_STRUCTUREthis answers "which page editsAgstammUseCase.broker.ebene" across the frontend/backend boundary.GET /data-structures/{name}/fieldsgainsboundBycounts;ac data-structure-fieldsfollows.Implemented 2026-09-22. As planned, with these decisions from VALIDATE: the target is the generated interface's
FIELD(declaring DTOBroker, not the root), reached through abinding=trueplaceholder resolved exactly by module → structure → field (the generic resolver never resolves module-owned structures, item 74); every hop is typed by the checker (XFields<TRoot, TSelf>), list hops are the call's result type; a prop-rooted expression ispartialand its carrier prop is not part of the path;kind=prefixrecords a handed-on sub-object as a read of the container field; WRITES for tags matchingInput$|Dropzone$|Editor$. One endpointGET /bindings(with the field'sCOUNTERPART_OFcolumns) instead of per-field endpoints;data-structures/{dto}/fieldsgainedboundReads/boundWritesand — a bug found on the way — now returns TypeScriptFIELDs at all (the query filtered them out; item 193's doc claim was wrong). A second 193 flaw surfaced on real data: interfaceFIELDs were named by the bare member, and the node identity is type + name + file, sovidwas ONE node under six interfaces of the generated file (sixCOUNTERPART_OFtwins, six-fold binding rows). Fields are now<Interface>.<member>with propsfield/owner;data-structures/{dto}/fieldsandbindingsreport the bare member,counterpartsthe qualified name. Sidecar contract version 3 (bindings). Sidecar dry run onpur-ui: 204 bindings (174 field, 30 prefix; 50 partial) over 22 DTOs — 61SmartInput, 64SmartOutput, 23fieldTermForRowData. Verified inBindingsIT(frontend + backend, counterpart columns, counts, no placeholder left) and onpurfe(server 322, recreate + deep refresh): 257 sites over 23 DTOs, all linked topur;counterparts?kind=field744 fields with exactly one twin each. A leaf inherited from a base interface (datStartonAbstractHistorizedDO) resolves to the declaring interface — the single-hop case was fixed after that run and is covered by the next deploy.
The Redux store — item 194 (2026-09-22)
-
194. Store:
STORE_SLICE+FIELD,READS/WRITES/CALLSfrom reducers, selectors, dispatchRedux Toolkit: 16
createSlice/createAppSlicefiles, thunks viacreateAppAsyncThunknamed<slice>/<op>, status viaisSlicePending/Fulfilled/Rejectedmatchers, three-hop access slice → facade hook (useAgstamm,useAgstammSelector) → component;pur-r-vbuchselects by lodash path into the whole state. NewNodeType.STORE_SLICEper slice withFIELDchildren named slice-qualified (agstamm.agstammUseCaseSvcResult) so/variables/{name}/reads|writesandflow-forwardwork unchanged. Reducer assignments →WRITES, selectors →READS,dispatch(action)→CALLSto the reducer/thunkFUNCTION. The store field is not flattened into the DTO: it holdsSvcResult<X>andUSES_TYPEthe DTODATA_STRUCTURE.GET /projects/{p}/store,GET /projects/{p}/store/{slice}/fields/{field}/reads|writes;ac store,ac store-reads,ac store-writes.contextextended with store reads/writes.Implemented 2026-09-22. Deviations from the plan above, decided while reading the real store: the slice is named by its reducer key (
gruppenprovision), not the RTK name (generalAgreement) — the key is what every selector path starts with; the store mapping is traced by the sidecar throughxReducer = xSlice.reducer/export default, across workspaces. Reducers areFUNCTIONs named by the action type they handle (schluesseltabelle/updateX,schluesseltabelle/suche/fulfilled,…/matcher:isSlicePending(sliceName)), so a dispatched action'scalleeMethodis the reducer. Reads come from three forms: selector arrows (incl. destructured results), wrapper hooks (useSchluesseltabelleSelector, resolved to their base path) andgetState()chains (81 sites onpur-ui). Instead of one endpoint per field, oneGET /store(slices with fields and counts) and oneGET /store/{slice}/accesses?field=&mode=; CLIac store,ac store-accesses.USES_TYPEfield → DTO deferred to 195 (the field'sdataTypealready saysSvcResult<X>);contextis not extended (the accesses endpoint andvariables/<slice>.<field>/reads|writesanswer). Sidecar facts contract bumped to version 2 (slices,store,stateAccesses,calls[].actionType). Cross-file reads are placeholders<key>.<field>(store=true) resolved by a new cheap finalize step in every mode; the item-74 stale-edge sweep covers store targets on a deep re-ingest. Verified on the fixtures (sidecar, parser, reader tests) and end-to-end inStoreIT; onpurfe(server 318, recreate + deep refresh, 29 s): 9 slices, 43 reducers, 216 reads / 82 writes, no placeholder left; the reducer key ≠ slice name case (gruppenprovision/generalAgreement) resolves correctly. Not modelled: a slice created inside a factory function (filetransferSlice), which is also not mounted in the store.
Web-service calls and the counterpart link — item 193 (2026-09-22)
-
193. Webservice calls: outbound
rest-endpointson the frontend andCOUNTERPART_OFtopurTwo generated client generations exist: legacy
generated/*endpoints.tsclasses (one per backend controller,baseUrl+get/postmembers,AgstammControllerEndpoint.saveBroker.post(...)) and, inpur-r-vbuch, a hey-apisdk.gen.ts(URL literal per function, consumed via TanStack Query hooks and thunks). Both becomeFUNCTIONs carryingrestPath(composed frombaseUrl+ member path),httpMethod,outbound=true,requestType,responseType— the same properties the Java parser writes (item 130), soGET /rest-endpointslists the frontend's outbound calls with no new endpoint. Thunks and components reach them through ordinaryCALLS.New
EdgeType.COUNTERPART_OFand a counterpart enrichment step (Cypher inCypherQueries, driven like the other steps — there is noGraphEnricherSPI): outbound frontendFUNCTION→purhandler matched onhttpMethod+restPath; generatedDATA_STRUCTURE→ Java class by simple name;FIELD→FIELDby name. Per-file reconcileDETACH DELETEs incoming edges, so the step re-runs at the end of every refresh of either project; a project settingcounterparts: [..]names the partner. First cross-project MERGE in the codebase.GET /projects/{p}/counterparts? module=&limit=&offset=+ac counterparts. Same edge serves item 143.Implemented 2026-09-22. Sidecar (
extract.mjs, contract v1 + additiveendpointsanddeclarations[].members): the legacy generator's Endpoint classes (baseUrl+get/postmembers, URL template reconstructed frombuild<Backend>URL(…),{param}for substitutions, request/response/params types from the member's type arguments) and hey-api sdk functions (url:literal, verb from the client call). Parser:TypeScriptRestPathssplits an application base (/<name>/v<n>) off and strips the query string; endpointFUNCTIONs carryrestPath/httpMethod/outboundexactly like Java handlers, sorest-endpointslists them with no query change beyondoutbound = f.outbound = 'true'; interface members becomeFIELDs; a member call on an Endpoint instance is retargeted to the endpoint function (calleeMethod = AgstammControllerEndpoint.saveBrokeron the module-to-module edge; note thatfunctions/{fn}/callersdoes not follow such edges for Java either — item 197). Store:EdgeType.COUNTERPART_OF, project propertycounterparts(create/update/get/list, RESTProjectRequest.counterparts, CLI--counterpart, self-reference →400 COUNTERPART_SELF), enrichment stepsdelete-counterpart-edges/link-counterparts-rest(path shape key:{param}→{}, class + method@Pathcomposed asREST_ENDPOINTSdoes — keep in step) /link-counterparts-dto/link-counterparts-field, run after every finalize for the project and for every project listing it (COUNTERPART_HOLDERS), the first cross-project MERGE in the codebase.GET /projects/{p}/counterparts+ac counterparts(--kind,--unmatched,--count-only, paging).CounterpartsITbuilds a Java backend and a TypeScript frontend as two projects and pins: outbound rows, call → handler incl.{vermnr}shape,unmatched= the one call nothing serves, DTO and field links, survival of a backend refresh, the setting and the two 400s. Scope note: the registered frontend ispur-ui+pur-ui-common(backendpur, base''), so the base-stripping matters only for the excluded vstamm/vbuch clients. Verified on the deployed server (version 312) 2026-09-22:purfedeep refresh 29 s (sidecar 6.1 s + 8.6 s for the two workspaces), 285 files, 0 failures;rest-endpoints51 outbound rows, 51 of 51 linked topurhandlers, 243 DTOs and 632 fields linked; the 52 unmatched DTOs are mirrors of JDK/framework types (Class,Comparable,Annotation, …) with no class inpur. Two defects found on real data and fixed the same day: imports ofpur-ui-commonresolved into itsdisttypings (now mapped to thesrctwin) and transitive packages (immer,redux) became placeholders (now external imports; Tier-1 readsnode_modulesdirectory names as externals). After the fix (server 314, project recreated because of item 198): 0 edges intodist, 7 placeholders left (from the Tier-1 pass before the node_modules rule), counterparts unchanged 51/51.
TypeScript / React analysis — item 192 (2026-09-22)
-
192.
ac-parser-typescript: language wiring, coarse scan, sidecar, import/call graph, LoCNew Maven module
ac-parser-typescriptimplementingLanguageParser,CoarseScannerandLineCounterfor languagetypescript(.ts/.tsx) pluscss(.css). Frontend registered as projectpurfe. DeliversMODULEper file,FUNCTIONper exported function / React component / hook (propertykind),CALLSandREFERENCESfrom imports and call sites,DATA_STRUCTURE+FIELDper exported interface/type (generated ones carrygenerated=trueandjavaCounterpart=<simpleName>), and LoC/SLoC soGET /loc?language=typescriptworks.Validated design decisions (do not re-derive):
- Sidecar =
ac-parser-typescript/sidecar/(ownpackage.json, pinnedtypescript, built in anode:24-slimimage stage and copied intoDockerfile.jvmtogether with thenodebinary). The JVM starts it per workspace withProcessBuilder, a timeout and--max-old-space-size=1024, reads a per-file JSON facts document from stdout, and the process ends.noEmit, noincremental(the project root is mounted:ro).excludeDirs(node_modules,distby default) applies to the file walk only; the sidecar still reads the project'snode_modulesfor library typings. - Project context like
CopycodeLibrary: aTypeScriptFactsbuilt once per ingest and passed into every per-fileparse(); the Tier-1 coarse scanner is pure Java regex and needs no Node. SourceFiles.LanguagegainsTYPESCRIPTandCSS, and every== JAVA/ "else Natural" branch (AstIngestService.parse/coarseScan/count,ProjectIngestService~223 / ~1054, the sevenlanguage: 'natural'|'java'literals inCypherQueries) becomes an exhaustive switch.ProjectResource.SUPPORTED_LANGUAGESgainstypescript.- Anonymous nodes (a
sxblock, a store write) use thestartLineMERGE variant likeDB_ACCESS, named<owner>@<line>. docker-compose.yml:mem_limit4g → 5g, frontend root mounted:ro.- Fixtures under
ac-parser-typescript/src/test/resources/fixtures/typescript/; the JSON facts format is the tested contract between the two halves, so Java unit tests need no Node. - Accepted: a
changedOnlyrefresh re-runs the sidecar over the whole workspace but re-persists only the changed files; unchanged files' facts may lag one refresh (same class as item 46a).
Implemented 2026-09-22. New module
ac-parser-typescript(TypeScriptCoarseScanner,TypeScriptParser,TypeScriptLineCounter,CssLineCounter,TypeScriptModuleNames,TypeScriptProject,TypeScriptFacts+TypeScriptFactsReader,TypeScriptSidecar) and the sidecarac-parser-typescript/sidecar/extract.mjs(pinnedtypescript5.9.3,npm ci; the JSON facts contract v1 is documented at the top of the script and pinned by the checked-infacts-pur-r-vstamm.json, whichTypeScriptSidecarTestregenerates live and compares). Server:SourceFiles.LanguagegainedTYPESCRIPT/CSSwith exhaustive switches inAstIngestService(parse/coarseScan/count);SourceFiles.ingestedBygates TypeScript/CSS totypescriptprojects;DEFAULT_EXCLUDE_DIRS=target, node_modules, dist;TypeScriptSidecarServicebuilds the per-ingest context (Tier-1: package.json only; deep: sidecar per workspace, failures reported assidecar:<workspace>pseudo-paths); the copycode stand-down of item 129 is now Natural-only by name.ProjectResource.SUPPORTED_LANGUAGESandac project create --languageaccepttypescript.Dockerfile.jvmis a two-stage build (node + sidecar copied fromnode:24-slim), the compose build context moved to the repository root with a root.dockerignore,mem_limit4g → 5g, the frontend root mounted:ro. Configagenticcode.typescript.*inapplication.properties(%prodpoints into the image). Measured on the real frontend: 413 files, 15 s for all four workspaces, imports resolved 99.5 % (the rest:index.css, two deepmomentlocale paths). Decision taken while implementing: a CSS module keeps its.cssin the identity, becauseindex.cssnext toindex.tswould otherwise collide on…/src/indexand be skipped as a duplicate. 35 unit tests in the module, 232 across the build. - Sidecar =
Closing the performance campaign (items 173-175) — 2026-09-06
-
173. Where the deep refresh stands after items 153-172, and what is left (written 2026-09-06)
Deep refresh of
upms: 1 225 s -> 301 s (-75 %) (item 175, two clean runs: 301 s / 290 s). Nine reformulations landed, two narrow indexes were kept, five investigations produced no lever and are recorded so nobody repeats them.Where the time sits now. Re-measured 2026-09-06 (item 175) with two full
rebuild-and-refresh.sh upmscycles, no instrumentation attached, both figures from the runs' own log lines: run 1 = 301 s, run 2 = 290 s (spread 3.7 %; run 2 is the faster one because it starts against an already resolved graph). Both columns below are run 1 / run 2.| persist (148.7 / 139.0 s) | | finalize (136.8 / 135.1 s) | | |---|---|---|---| |
merge-edges| 47.6 / 45.9 s |resolve-field-placeholderWRITES | 23.8 / 23.9 s | |commit| 34.1 / 29.2 s |link-args-to-params| 23.7 / 23.2 s | |merge-nodes| 24.0 / 22.9 s |resolve-bare-includedWRITES | 20.7 / 20.2 s | |merge-positional-nodes| 13.3 / 14.2 s |resolve-field-placeholderREADS | 19.6 / 19.3 s | |merge-nodes-copy| 7.8 / 7.0 s |resolve-bare-includedREADS | 17.2 / 17.0 s | |sweep-*(3 steps) | 9.6 / 9.1 s |delete-resolved-placeholders| 7.1 / 7.1 s | |reap-*-edges(3 steps) | 8.3 / 7.4 s |stamp-unresolved-placeholders| 4.3 / 4.2 s | |prepare+tx-open| 1.2 / 1.1 s | remaining 44 steps | < 2.6 s each |Parsing is only ~9 s: parse and persist are interleaved, so the refresh is essentially persist plus finalize in equal halves. The estimates this table previously carried held up well at step level (every one within ~1 s of the measurement); only the two totals were off, ~162 s / ~140 s against a measured 148.7 s / 136.8 s.
Already investigated without a lever — do not re-open without a new idea:
merge-edges(166, the Eager costs ~5 %, removing it would freeze the property model),resolve-field-placeholder(168, write-bound: 2.6 M property writes in 45 s),commit(170, measured as genuine commit I/O once the residual was split),merge-nodes(171, two suspicions both refuted),link-args-to-params(172, seek-count-bound; the one index that helps costs 15 s to save 10).Candidates for a next attempt, honestly ranked by what is actually known:
Fewer, larger persist batches.Measured and rejected 2026-09-06 (item 174) --- see below. The fixed per-batch cost is ~70 ms, so halving the batch count would save ~1.1 s out of 331 s.delete-resolved-placeholders(7.2 s) and the other sub-5 s finalize steps. Small, and the effort-to-payoff ratio is now clearly worse than it was at item 153.- Nothing on the read side. After items 163/167/169 the field resolvers are seek-bound, and item 172 showed the only index that would cut seeks costs more than it saves.
What is NOT worth doing, so the next person does not spend a run finding out: raising the page cache (item 166 measured 11.2 MB read over an entire refresh against a 2.9 GB store — the cache is not the constraint), and any further whole-graph index (item 172: write cost depends on how many written nodes touch the index, not on index size).
Method notes worth keeping, each of them learned the hard way today: db-hits systematically mislead on seek-heavy and MERGE-heavy work — measure the clock as well; a single refresh is not a verdict, always check a step you did not change before believing a delta; and a read-only measurement against an already resolved graph short-circuits the field resolvers and flatters every prediction.
Housekeeping still open: commit
d6924a1is a pure deletion ofCypherQueries.java(the file was destroyed by a patch script and restored from56314d5), andCypherQueries.java/GraphRepository.javacontain mojibake bytes that makegreptreat them as binary, so recursive searches silently skip them. (Two further points listed here on 2026-09-06 — the half-stagedCypherQueries.javaand item 161's "not yet measured" title — were resolved the same day.) -
174. Larger persist batches --- measured, no lever (2026-09-06)
Item 173 listed "fewer, larger persist batches" as the only remaining idea with a plausible mechanism. It was measured before it was attempted, and the mechanism does not exist. No code change;
agenticcode.ingest.batch-sizestays at 200.How it was measured. One deep refresh of
upmsat batch size 200 (32 batches, 6 311 files), with thePersistStatsinstrumentation from item 170 and a sampler pollingSHOW TRANSACTIONS YIELD estimatedUsedHeapMemoryon the Neo4j side every ~3.8 s. The run took 331 s rather than the 302 s baseline; the sampler spawns acypher-shellJVM per sample and accounts for that ~10 %. Timings below are from the run's own log lines, not from the wall clock.The fixed per-batch cost is ~64 ms. Summed over all 32 batches:
prepare1.2 s,tx-open0.008 s,unattributed0.000 s --- 37 ms per batch of genuinely fixed work. The only other fixed component is the commit floor, visible on the batches that wrote nothing at all (commit25-28 ms). Going from 200 to 400 files removes 16 batches, i.e. ~1.0 s of ~300 s (0.3 %) --- inside run-to-run noise, which item 175 measured at 3.7 %. Everything else in persist scales with the data written; per-label figures are in item 175's table.Correction (item 175). The persist figures first written here came from the sampler run and were inflated by it: persist 159.3 s instead of 148.7 s (+7 %), and the persist/finalize split was given as 159 s / 163 s when it is really ~149 s / ~137 s --- the sampler cost finalize ~19 %. The per-label numbers above were withdrawn for the same reason. The conclusion is unaffected and in fact slightly stronger: the fixed per-batch cost is 64 ms, not 70 ms.
The earlier batch-400 failure was blamed on the wrong limit. Item 173 recorded that the attempt "failed on
dbms.memory.transaction.total.max". That attribution was unfounded and is withdrawn: the measured peak Neo4j transaction heap at batch size 200 is 313 MB against a 1.40 GiB limit (default 70 % of the 2 GiB Neo4j heap), i.e. 4.5x headroom --- Neo4j was never the binding constraint. The likelier constraint is the server JVM:JDK_JAVA_OPTIONS=-Xmx2500m, and the run logs show it sitting at 2 219 / 2 500 MB (89 %) throughout, with a whole batch of parsed ASTs held in memory before persisting. The original failure log is gone (the server container has since restarted), so this is the likely cause, not a proven one. Either way the upside does not justify finding out.Where persist time actually goes. Persist is ~149 s and finalize ~137 s (item 175) --- parsing is only ~9 s, because parse and persist are interleaved. Persist is therefore still half the refresh, but all of it is per-row write work in
merge-edges/merge-nodes/commit, all three of which are already recorded as investigated without a lever (items 166, 171, 170). -
175. Closing measurement of the performance campaign (2026-09-06)
Every figure in item 173's table was an estimate carried over from mid-campaign runs, and item 174 had published persist figures taken from a sampler-perturbed run. Both are now replaced by measurements from two full
rebuild-and-refresh.sh upmscycles run back to back with nothing else touching the stack. Docs only, no code change.| | run 1 | run 2 | |---|---|---| | deep refresh, end to end | 301 s | 290 s | | persist (32 batches) | 148.7 s | 139.0 s | | finalize (51 steps) | 136.8 s | 135.1 s |
Run-to-run spread is 3.7 %, and it is not noise: run 2 starts against a graph the preceding run already resolved, so the field resolvers find less to do. Any future A/B has to compare a first run with a first run. Take 301 s as the campaign's closing figure, since that is the condition every earlier baseline was measured under.
The old estimates were better than expected at step level --- every single step in item 173's table came within ~1 s of its measurement (
merge-edges46.1 estimated vs 47.6 measured,link-args-to-params23.6 vs 23.7,resolve-bare-includedR+W 38.1 vs 37.9). Only the two totals were wrong, and the sampler run's figures in item 174 were wrong by more (+7 % persist, +19 % finalize). Instrumentation that polls the database distorts finalize far more than persist --- worth remembering before attaching a sampler to a run whose timings are meant to be quoted.