Files
agenticCode/x-docs/roadmap.md
Ingo Schnabel 7e0d75cc5e Roadmap
2026-09-23 12:39:15 +02:00

77 KiB
Raw Blame History

AgenticCode Roadmap — Open Tasks

The remaining work. Completed items live in x-docs/features.md with their full implementation notes — when an item is finished, move it there rather than marking it [x] here. Findings that turned out to be wrong go to "Not a bug" at the end, so nobody re-files them.

High priority: items 142–143 — a dead Natural subroutine is indistinguishable from a live one, and the cross-project counterpart relation has no edge. (141, comments not indexed, is implemented — 2026-08-28, see features.md; both remaining items build on it.) Postponed: 51 (M6) (Web UI scale & polish, 2026-07-15) and 133 (reference-index cost, accepted 2026-08-20). Everything else is open at normal priority.

Restructured 2026-08-28: 59 completed items moved to x-docs/features.md, sections consolidated from eleven to six. No item text was changed, except that a second item numbered 91 became 151.

High priority — comments, dead code and the counterpart relation (2026-08-27)

Three gaps reported after the VermittlerController reengineering round (nine UPMS services, four of them built in one session). They share one root: the graph holds structure, and the UPMS→PUR work lives half in the things that are not structure — Natural comments, commented-out code, and the Java↔Natural correspondence that is recorded only as a Javadoc convention. All probes below were run 2026-08-27 against a freshly refreshed graph (upms ingested 08:37Z, pur the same morning) and are reproducible as written.

Item 141 — the one that made the API return a wrong answer rather than a missing one — is implemented (2026-08-28) and moved to features.md. Comment blocks are now COMMENT nodes with a DOCUMENTS edge to the declaration they document, reachable via /modules/{name}/comments and search/value?includeComments=true. The two items below build on it.

  • 142. functions cannot tell a live subroutine from a dead one — an auskommentierter PERFORM is invisible

    Symptom.

    GET /upms/modules/WAGNTX0S/functions → 40+ rows, every one { name, declaredIn, sourceFile,
                                            viaCopycode, startLine, endLine, kind: null }
    

    There is no field saying whether anything still performs this subroutine. In Natural of this age, logic is not deleted — it is commented out, usually only at the PERFORM site, leaving the DEFINE SUBROUTINE block fully intact and fully parseable. The graph sees a declared function with a body and reports it as one.

    Why this is the expensive one. The consuming project names it as such in its own instructions: "A subroutine whose PERFORM is commented out is dead; treating it as live is a common and serious error." The failure mode is not a wrong query result the agent can notice — it is a subroutine faithfully reengineered into Java, reviewed, tested and shipped, implementing behaviour the legacy system stopped executing years ago. Nothing downstream catches it, because the Java is a correct translation of code that really is there.

    Proposed fix. Two derived properties on the FUNCTION row: performSites (count of live PERFORM/CALLNAT references reaching it) and live (performSites > 0, or the module's entry point). A dead subroutine is then one field on a response the agent already fetches, instead of a file read it has to remember to do. Note the ordering dependency: counting live sites means knowing which PERFORMs are commented out, which is item 141's parse work — done as of 2026-08-28: both parsers recognise comment lines and persist them as COMMENT nodes, so a commented-out PERFORM is now visible as comment text rather than absent from the graph.

    Report live: null rather than true where the analysis cannot decide — a PERFORM behind an unresolved dynamic dispatch (item 145), for instance. A false true restores exactly the failure this item exists to remove.

  • 143. The Natural↔Java↔caller correspondence has no edge, and "what is not reengineered yet" cannot be asked

    Symptom. Three projects hold three views of one service, and nothing joins them:

    project holds how the link is recorded
    upms WAGNTX0S.nat —
    pur AgentProvisionUpdateLogic UpmsObject: Javadoc block (invisible, item 141)
    app UpmsAgentUpdate named in the same Javadoc block

    Answering "what is the counterpart of this" today means a search/value per project, a guess at the naming convention, and — for web services — a fallback to reading files, because of 141. It is the first question of every job, and it costs several calls plus one known-unreliable heuristic.

    The question that cannot be asked at all. "Which upms web-service programs have no pur counterpart yet?" is the project's central planning question — what is left to do — and there is no formulation of it against the API. It is currently answered by a human keeping a list.

    Proposed shape. Parse the two conventions the codebase already follows — PROGRAM_IDENTIFIER / the // XXXXXXXX.nat comment, and the ServiceEndpoint:/UPMSFunction:/UpmsObject: Javadoc block — into a first-class cross-project edge:

    pur:AgentProvisionUpdateLogic --REENGINEERED_FROM--> upms:WAGNTX0S
    pur:AgentProvisionUpdateLogic --SERVES-->             app:UpmsAgentUpdate
    
    GET /upms/modules/WAGNTX0S/counterparts       → the pur and app sides, with the evidence line
    GET /upms/counterparts?missing=true           → the backlog, as data
    

    Three things to get right.

    • This is the first edge that crosses a project boundary. Every existing query is project-scoped, and the traversal queries (callers, callees, call-tree) must not start following it. Item 128's MENTIONS-vs-REFERENCES decision is the precedent, and the reason it was the right one.
    • The evidence must ride along. Return the sourceFile:lineNo of the Javadoc block or of the constant that produced the edge. A correspondence asserted without its evidence is unusable in a project whose reports require file+line for every claim.
    • A missing counterpart is not the same as an unparseable one. missing=true must distinguish "no pur module claims this program" from "a module claims it in a form the parser did not recognise", or the backlog list silently pads itself with the parser's own gaps.

Known bugs

  • 203. Synthetic inheritance CALLS edges carry an arbitrary line and are never reaped (found 2026-09-23 comparing a fresh pur ingest against the 2026-09-11 one)

    Symptom. LINK_CALLS_TO_IMPLEMENTATIONS merges one CALLS {resolvedVia: 'INHERITANCE'} edge per caller/implementation pair and sets lineNo ON CREATE from whichever originating base call the match meets first. With several call sites the line is arbitrary and changes between runs: SearchResultLogic -> ResultRepository was line 95 in one graph and 97 in the other, both real calls. callers/callees sites for such an edge therefore point at one call chosen at random. The edge also carries no ingestGen, so the item-198 reap never removes it when the base call is gone; only a project recreate does (pur: 75 stale ones next to 235 new ones in unchanged files).

    Fix. Key the synthetic edge on the originating call (lineNo + originFile, one edge per base call site, like the parsed edge it derives from) or aggregate all originating lines into a list; and delete the project's (or the scoped modules') inheritance edges before re-deriving them in the finalize, as DELETE_CALLS_MODULE does for CALLS_MODULE.

  • 152. NaturalLines.stripInlineComment is not quote-aware — a /* inside a string literal truncates the statement (split out of item 141 on 2026-08-28, pre-existing)

    Every Natural scan path strips a trailing comment with line.indexOf("/*"), ignoring string literals: MOVE 'A/*B' TO #X is cut to MOVE 'A, and whatever the statement said after the literal is invisible to the parser. Call sites: NaturalLines.stripInlineComment (shared by NaturalParser and NaturalCoarseScanner) plus the private copy in CopycodePreprocessor.

    Found while implementing item 141, which needed the opposite direction — deciding whether a /* starts a comment node — and therefore has its own quote-aware scanner (NaturalParser.inlineCommentStart). It was deliberately not folded back into the strip path in the same change: there the bug merely truncates a line, and swapping the rule underneath every Natural scan is a behaviour change across the whole corpus that deserves its own measurement (how many lines actually contain a quoted /*, and what starts resolving once they stop being truncated).

    Fix shape: move the quote-aware scan into NaturalLines, use it from all three call sites, delete the CopycodePreprocessor copy, and report the ingest delta on upms before and after.

  • 73. A NONE/ANY branch is reported under its enclosing guard alone — the condition is a negation no guard chain can express (found 2026-07-17 while fixing item 72; item 72 does not fix this). NaturalParser's VALUE_RESET clears a DECIDE's active value on NONE/ANY, so an assignment inside such a branch is attributed to the enclosing guard chain only. Under VALUE 'TABL' → DECIDE ON #FIELD-NAME → NONE → MOVE ..., item 72 now reports guards = [#SHORT-VIEW='TABL'], which reads as "happens for all of TABL". The truth is "#SHORT-VIEW = 'TABL' AND NOT (#FIELD-NAME = any of the branch's VALUEs)". This is the same complaint item 72 makes — a condition served as complete when it is not — so item 72's guards is not a total answer, only a strictly better one. A guard chain is a conjunction of equalities by construction; expressing this needs a negated link (e.g. a DispatchGuard with negated=true carrying the sibling branches' values), which is a model change, not a parser tweak. (Unmeasured: how many NONE branches actually contain assignments — many are IGNORE. Worth counting before investing.)

(Fixed bugs 55, 57, 58, the dossier field-ordering fix, and the raw-hop depth family — 65, 67 (call-tree), 68 (field-flow) — plus 69/70 (edge identity + copycode provenance) moved to x-docs/features.md. #71 is a known imprecision left behind by item 67 rather than a wrong answer; #72 is a real wrong answer, found by the 2026-07-17 VMULTMN4 audit.)

  • 137. 1437 nodes have endLine < startLine (split out of item 75 on 2026-08-23; the symptom was first noted 2026-07-17 there). A copycode that opens a block it does not close carries the defect into every host: YFRAMBC0.cpy opens FOR at line 65, the END-FOR lives in the includer, and the node ends up 65 -> 15. Item 75's per-site identity did not cause this and did not fix it — it only stopped hiding it, so the count tracks the number of expansion sites (383 -> 1007 -> 1437) rather than the number of defective constructs. Mostly DB_ACCESS and CONTROL_FLOW. Not yet assessed: whether any consumer reads these ranges (line-range filters, snippet extraction, the UI's source highlighting) and what it does with an inverted one. Measure that before deciding between "clamp endLine to startLine", "end the node at the copycode's last line", or "leave it and document it".

  • 138. Same-named declarations inside one file collapse onto one node, and its type is whichever declaration was persisted last (found 2026-08-23 while corpus-verifying item 75; measured 2026-08-23). The node merge key is (type, name, sourceFile, project, ownerModule), and VARIABLE/DATA_STRUCTURE are not in POSITIONAL_NODE_TYPES, so startLine is not part of it. Every declaration of a name within one file therefore becomes one node carrying one dataType and one startLine. Not Natural-specific — Java is the worse case. JavaParser.java declares a parameter named type in 22 methods with five different types (NodeType 42/47, EdgeType 53/57, String 208/619, TypeDeclaration<?> 262/354/410, ClassOrInterfaceDeclaration 631/646/652). The graph holds one node, dataType: "NodeType", startLine: 42. GET /api/projects/ac/search/identifier?name=type returns exactly one row for that file — the wrong type for 20 of the 22 methods, with nothing marking the answer as ambiguous. The identifier index is not occurrence-based, so there is no second row to disambiguate from.

    Measured — how often the collapse actually conflicts (declarations compared against the source; the scanner is a regex, not the parser, so treat the Java figures as ±noise and the "unmeasured" column as genuinely unknown rather than clean):

    project collapsed nodes identical decls conflicting unmeasured
    upms (natural) 9480 5037 93 4350
    ac (java) 217 157 51 9
    pur (java) 2738 2330 374 34
    app (java) 2216 1470 325 421

    Natural is largely benign: a field repeated across many views is genuinely the same field (COD-EMISOR in YAGNTBNH.nat, 55 declarations, all (A8)), and a good share of the 93 are spelling variants of one format (N8 vs N8,0, A03 vs A3). The real Natural conflicts are things like #TEXT in JB0001N1.nat declared A45/A50/A60/A90. Java conflicts are ordinary and unavoidable: depth as IngestDepth/String/int, PartnerAdresse vs PartnerAdresse[]. Total multi-parent nodes: upms 10,315 (worst 55 parents), pur 3125, app 2233, ac 339 (worst 81). Re-verified 2026-08-27 after all four projects were recreated on server 254 (so the counts no longer rest on the pre-75-C node identity): upms 10,315 (worst 55) and app 2233 (worst 54) are unchanged, pur moved 3125 -> 3138 (worst 54), ac 339 -> 302 (worst 81). The conflict/identical split in the table above was measured before the recreate and has not been re-scanned; the multi-parent totals barely moved, so it is unlikely to shift much, but it is not re-verified.

    Second consequence, edges — upper bound only. READS/WRITES/USES_TYPE/RESOLVED_FIELD edges touching a collapsed node: upms 27,328, pur 590, ac 216, app 110. These are only contaminated where the declarations genuinely differ, so this bounds the damage from above; it does not measure it. Where they do differ, field-flow and variables/reads|writes join two unrelated variables into one flow — an invented edge, not merely an imprecise one.

    Cost of fixing — lower bound. Keying per declaration site would take upms 508,670 -> ~537,248 (+5.6%), pur +14.5%, app +13.4%, ac +13.7%. Lower bound because two declarations under the same parent share one CONTAINS edge (edge identity, item 69) and are counted once here. Not assessed, and the reason this is not "just add startLine to the key": every lookup that resolves a field without knowing its line — the bare-field resolvers (item 77), USING resolution, the call/table enrichers — would have to pick among N nodes where it previously found one. That is the actual work, and it is unexamined. A node-identity change also forces a recreate of every project.

    Two corrections to the original entry (2026-08-23). (1) It claimed "any query that walks up from a field to its structure gets an arbitrary one of 56 answers". There are exactly 5 upward CONTAINS traversals in all of CypherQueries, and every one of them is anchored on FUNCTION or DB_ACCESS parents, never on a field. The damage is the collapsed attributes and the contaminated edges, not an upward lookup. (2) The first count (14,843 nodes, worst 56) counted relationships, not distinct parents: Java emits one CONTAINS edge per occurrence line, so the project node in GraphRepository.java shows 746 edges from 81 methods across 663 distinct lineNo values. The corrected figure is 10,315 for upms. (That per-occurrence CONTAINS is worth its own look: it makes the edge count of CONTAINS a reference count rather than a containment count, which is not what the schema doc claims it is.)

  • 139. search/value cannot find string content that is not part of a single-line expression — Java text blocks are entirely invisible (found 2026-08-23 while measuring item 138; scoped 2026-08-27). Looking for the Cypher source of a query, GET /api/projects/ac/search/value?value=MATCH (n:AstNode&contains=true returns [], although that text occurs 19 times in ac-neo4j-store/.../CypherQueries.java. ownerModule likewise returns []. The fallback was grep.

    What the index actually holds. Every hit comes back as kind: "NODE" with value set to the source text of a single-line expression (properties.store(out, "AgenticCode CLI configuration")). String literals are therefore findable only incidentally, when they happen to sit inside such an expression on one line. Two classes are missing:

    • Text blocks ("""…""" assigned to a static final String) — the entire CypherQueries class, and by extension any embedded SQL, Cypher, JSON or HTML held the same way.
    • Multi-line expressions — "MISSING_VALUE" at AnalysisResource.java:838 is a plain call argument and still returns [], because the call is wrapped across lines.

    Why it matters here. AgenticCode's own purpose is making unfamiliar code searchable; embedded query text is exactly what an agent asks about ("where is this table read?", "which query builds this projection?"), and today that question can only be answered with grep. Not yet decided: whether to index literal content as its own value kind (kind: "LITERAL") or to widen expression capture to multiple lines. The first is the more useful shape but adds nodes; the second is cheaper and fixes only half. Neither the node cost nor the Natural side (does the same gap exist for long MOVE/COMPRESS text?) has been measured.

  • 181. callees / callers truncate silently at 50 rows — no header, no body flag (found 2026-09-17, PartnerCopy Java↔Natural verification)

    GET /upms/modules/DPARTFN0/callees              → 50 items (17 MODULE), no X-AC-Truncated header
    GET /upms/modules/DPARTFN0/callees?limit=1000   → 58 items (25 MODULE)
    GET /upms/modules/DPARTFN0/callees?scope=external → complete
    

    The unscoped default drops YPARTBN0, YPARTGNH, YPARTMN0, YPARTMNH, YPHONBNH, YPHONMNH, ZINCLGET, ZINERR01 — among them the module that actually writes the partner. digest lists them, and callers of YPHONMNH does include DPARTFN0, so the analysing agent reported it as an inconsistency between endpoints; it is the 50-row page, and nothing in the response says so.

    Item 131 added X-AC-Total-Count / X-AC-Truncated to the search endpoints (item 135 to two more); callees/callers return an object with sourceFiles/items and carry neither the headers nor a truncated field. A wrong answer, not a missing one: the caller cannot tell a complete fan-out from a cut one. Either send the headers (and a truncated field, as call-tree does since item 67) or return all rows when limit is absent, as db-accesses does.

Agent API gaps

  • 108. dispatch-table only understands the DECIDE dispatcher, not the dispatch-table idiom — the one the endpoint is named after (found 2026-08-02, upms webservice-layer audit)

    Symptom. Two dispatcher idioms are common in upms, and the endpoint covers one of them.

    GET /modules/WSUBPX0S/dispatch-table  → 25 rows   (DECIDE ON VALUE 'supl_fin' → #W-ACT-PROG := 'WNSUPD0S')
    GET /modules/WPOLIX0S/dispatch-table  → []
    GET /modules/WGARCX0S/dispatch-table  → []        (likewise WACOMX0S, WCLAIX0S, WOBJPX0S)
    

    The five empty ones dispatch through an array built from literals:

    DEFINE SUBROUTINE INIT-OBJECT-TABLE
      ASSIGN #WT-OBJ-PROG (1) = 'WXSPOD0S'
      ASSIGN #WT-OBJ-PROG (2) = 'WXSDAD0S'
      …
    …
    #W-ACT-PROG := #WT-OBJ-PROG (#I-OBJ)
    CALLNAT #W-ACT-PROG …
    

    Every target is a literal in the source; nothing is runtime-dependent. These are 10 of the 116 entries in dynamic-calls/unresolved, all with variable: #W-ACT-PROG.

    Relation to item 83. Item 83 constant-folds MOVE '<lit>' / MOVE '<lit>' TO SUBSTR(var,pos,len) chains. This is the same class — literal assignment feeding a CALLNAT var — but through an indexed array element rather than a scalar, so the fold does not apply. The set of possible targets is exactly the set of literals assigned to any element of that array; which index is live at runtime is not statically known, so the honest model is n candidate edges (as item 82 already allows for a manual override with multiple targets), not one.

    Fix (proposal). Extend the fold to indexed writes: collect ASSIGN <array>(<n>) = '<literal>' for the array feeding CALLNAT, and emit one CALLS edge per literal that names a real ingested module, tagged folded=true + something like viaTable=true. Independently, dispatch-table should report them, keyed by index instead of by guard value — the response already has guardValue/assignedValue, so guardValue: "(3)" or a dedicated index field would fit. Caveat: in upms most of these targets are not ingested (see item 107), so the edges would resolve to nothing there — the value is in no longer silently reporting [].


    Correction (2026-08-10) — the call-graph half of this item is already implemented, and the fix proposed above was written and then reverted as redundant.

    RESOLVE_DYNAMIC_CALLNAT_INTRA_INDIRECT already resolves exactly this idiom, and its javadoc gives the identical example (ASSIGN #TBL(1) = 'WPARTD2S' → #W-ACT-PROG := #TBL(#I) → CALLNAT). For the assignment writing the dispatch variable it follows the READS on that same line to the array, then takes every string literal written to the array as a target — the same candidate-set over-approximation this item proposes. A fixture (DispatchTableFoldIT) confirms it end-to-end, including a table built across an IF/ELSE.

    Why upms still shows nothing: the target modules are absent from the checkout entirely — WXSPOD0S, WXSDAD0S, WXSCMD0S, WCARLD0S are not in the graph at all, not even as placeholders. A literal naming no ingested module correctly yields no edge. That is item 107's problem, not a parser or resolver gap. This item's own caveat predicted it; what the item got wrong is the conclusion that the idiom is unsupported.

    How it was caught: the new resolver's tests passed with the new resolver disabled — the sabotage check, not the green run, is what exposed the redundancy. Without it the duplicate would have shipped, adding a second enrichment step doing the same work by a different route.

    What actually remains open is narrower than the title suggests: the dispatch-table endpoint does not report these tables as rows. It has the fields for it (guardField/guardValue/ assignedField/assignedValue), and the source carries a real key — a parallel array (#WT-OBJ-NAME(2) = 'GARC' beside #WT-OBJ-PROG(2) = 'WXSGAD0S'), so guardValue: "(3)" as proposed above would needlessly discard it. Pairing needs the array index, which is not in the graph: lookupVariable resolves #WT-OBJ-PROG (1) to the base node and drops (1), so the edge carries only ["value", "originFile", "lineNo"]. Adding an assignedIndex edge property is the prerequisite — the ASSIGN_COMPUTE pattern already captures the bracket. Note 9 of the 65 modules build the table in an IF/ELSE, where one index carries different values per branch, so such rows must be reported as multiple candidates, and the IF guard itself is not on the edge (guard props come from DECIDE context only) — recovering it would need a CONTAINS traversal, which item 75 makes hazardous.

Normal priority. Same session as items 141-143, same probe date; none of these produced a wrong answer, they cost calls or forced a file read. Three candidates from the original list were dropped after probing and are recorded at the end, so nobody re-files them.

  • 144. db-accesses is table-granular — "who writes this column" cannot be asked

    db-accesses?depth=N reports READS/WRITES/DECLARES per table. The expensive defect of this session was one column deep: VERSVW_AGNT_COMI is a periodic group bound to a version through V_ID_REC, and the historisation base class wrote the version row without carrying the group forward — so every agent change silently dropped the commission schemes. Both sides appear in db-accesses: the base class writes VERSVW_AGENT, some other class writes VERSVW_AGNT_COMI. At table granularity there is nothing to notice.

    A ?columns=true variant returning {table, column, mode, via, sourceFile, lineNo} — or /db-tables/{name}/writers from the table side — would have made "which code writes V_ID_REC" a single call. The Java side is derivable without new parsing: @Column names are already stored (item 130 resolves constant references for @Path the same way), and the entity→table mapping is already the DECLARES row.

  • 145. call-tree does not say how many unresolved dynamic CALLNATs are inside the subtree it just returned

    The response carries sourceFiles, items and truncated — so a traversal cut by the budget is honest, which item 67 fixed. A traversal that is complete as far as the graph knows but crosses three CALLNAT #SUBPROGRAM sites is not distinguishable from one that crosses none. /dynamic-calls/unresolved lists every open site project-wide (upms: BGARMFN0:1213, BGARMSN0:657, …), but nothing correlates that list with a given call tree, so checking costs a second call plus a manual intersection against sourceFiles — and only an agent who already suspects the problem will make it.

    Suggest an unresolvedDynamicCalls count (and the sites) alongside truncated. Same principle: a subtree that looks complete and is not is worse than one that admits it was cut.

  • 146. flow-forward / field-flow stop at the CALLNAT parameter boundary

    Natural passes data positionally: CALLNAT 'YXXXMN0' #A #B #C binds to the callee's DEFINE DATA PARAMETER list by position. The mapping is mechanical and the graph has both halves — the call site's argument list and the callee's parameter declaration — but nothing joins them, so a field's life ends at the call. Following a value through a three-deep CALLNAT chain is a manual positional count per hop, done by reading both files, and it is where the reengineering evidence chain actually lives ("which request field ends up in this column").

    Item 68 already introduced the derived CALLS_MODULE edge to make field-flow cross subroutine boundaries; this is the same class of derivation one level out.

  • 147. Test coverage is not graph data

    "Which endpoints of this controller have an active test, and which only a @Disabled stub?" was a literal task this session, and it was answered by reading the test class top to bottom. Everything needed is already in the graph after item 128: @Test/@Disabled are annotations (search/annotation), the test→subject relation is a CALL or MENTIONS edge, and item 130 knows which method serves which REST path.

    Something like GET /rest-endpoints?withTests=true adding testSites and disabledTestSites per endpoint would answer it in one call. Modest scope, and it is the question that decides what to work on next in a reengineering push.

  • 148. upms reports 6315 examined / 6311 persisted / 0 failures — the four are unaccounted for

    GET /api/projects/upms → ingest: { filesExamined: 6315, filesPersisted: 6311,
                                       filesFailed: 0, failures: [], incomplete: false }
    

    Item 126 delivered exactly the metadata that makes this visible, and it immediately shows a gap it does not explain: four files were walked, produced no persisted node, and are not failures. Probably benign (an empty file, a file producing no module shell — cf. item 132), but "examined − persisted ≠ failed" is silently non-zero, and the consuming project's instructions still carry a rule saying a 404 from upms is no proof, with a filesystem cross-check as the documented workaround.

    Either account for the difference (a filesSkipped with reasons) or state that examined-minus- persisted is expected and why. Cheap, and it retires a standing distrust rule in a downstream project.

  • 149. Message numbers are findable per project but not across the pair

    Better than expected: GET /upms/search/value?value=4080 returns the ##MSG-NR assignment sites (DACOMEN0:40, DACTBEN0:40, …), which is most of what a validator job needs. What is missing is the other half — the same number in pur lives as a LiteralId/ErrorField enum constant, so "where is 4080 raised on the Java side, and does it exist at all yet" is a second search in a second project with a different result shape.

    Falls out of item 143's cross-project work if the counterpart edge exists; not worth its own mechanism otherwise. Filed so the connection is not lost.

  • 150. /modules/{name}/source works well, and the guide steers agents away from it

    Probed and correct on both languages, including the encoding question — upms files are ISO-8859 and come back as proper UTF-8 (…DB-Tabelle AGNT ändern…), ?file= reaches copycode that has no module node (ISICINDE.cpy:15-20), Java line ranges are exact. Two small things:

    • agent-api-system-prompt.md:18 says to read files from disk and use the source endpoints "only when you have no filesystem access". For a consuming project whose own rules make the graph the primary source and file reads the exception — and which enforces that with a hook on the Java tree — that guidance is backwards. It is a documentation change, not a code change: say that the endpoint is the appropriate path when the caller's policy prefers it, and note the ISO-8859 handling explicitly, since that is the reason an agent would otherwise reach for iconv.
    • A copycode is reachable by path but not by name (/upms/modules/ISICINDE/source → 404 MODULE_NOT_FOUND), while viaCopycode/includePath on other responses hand back exactly that name. Accepting a copycode name would close the loop.

Normal priority. Reported 2026-09-17 after verifying PartnerController.insertPartnerCopyHauptwohnsitz (pur) against the Natural service it replaces, WPARTX1S → WPARTD1S → DPARTFN0/DPARTEN0/YPARTMNH plus the nested ADDR level WADDRX0S → WADDRD1S → DADDRFN0/DADDREN0 (upms), with the old adapter UpmsPvwPartnerHauptwohnsitzInsert (app). Four agents compared mapping, validation, main address and sample-partner copy; every gap below forced a fall-back to reading source. Two further gaps from the same run are not re-filed: "which Java method implements which Natural rule" is item 143, and "is message 4136 raised on the Java side at all" is item 149 — both were needed in this run (two missing 4136 rules were found only by reading PartnerValidator). The silent 50-row cut on callees is item 181 under Known bugs.

  • 182. No call order and no guard conditions — "does X run before Y, and only on ADD?" cannot be asked

    The two most consequential findings of the run were ordering questions:

    • DPARTFN0.DO-ACTION → PROCESS-OBJECT → BEFORE-ET → COPY-ADDR-COMM-BANK-DATA runs inside the PART object call, i.e. before WPARTX1S.CALL-NEXT-LEVEL-PUT inserts the main address. The Java does it the other way round, so DADDREN0.CHECK-DOUBLES (8012) sees different data.
    • Are ACCESS-MDCL, CHECK-RELEASE, VMDCLN01, YMDCLMN0 reached on #ADD? They are not — DPARTFC0.cpy sends C-MOD-ADD to WHEN NONE → CALL-MAINTAIN — but digest, callees and call-tree list them unconditionally.

    callees returns a set with lineNos; call-tree a depth-annotated set. Neither says in which order a function performs its callees, nor under which DECIDE/IF branch. The DECIDE guard already exists on dispatch-table rows and CONTROL_FLOW nodes exist (P2-a), but no query joins them to a call edge.

    Proposed shape. GET /modules/{name}/functions/{fn}/sequence → ordered [{ lineNo, kind: PERFORM|CALLNAT|INCLUDE, target, guards: [{ construct, condition }] }], and an optional ?guard=CDAOBJ2.#FUNCTION=C-MOD-ADD on call-tree that prunes branches whose literal guard contradicts the value. Item 111e (per-statement CONTROL_FLOW) and item 75 (CONTAINS traversal cost) are the constraints; item 142 (live) is the sibling question for commented-out PERFORMs. Report a guard that cannot be evaluated statically as null, never as satisfied.

  • 183. Transaction boundaries are not graph data

    "Is the partner rolled back when the address fails?" needed four files: WPARTX1S:230 forces P-OPT-NO-ET := 'X', WPARTD1S:356 passes it on, USIX044C.cpy suppresses END TRANSACTION in DPARTFN0 (#OMIT-ET) and YPARTMN0, and WPARTX1S:235-240 finally does one END TRANSACTION or BACKOUT TRANSACTION for the whole PUT. No endpoint returns ET/BT sites, let alone the flags that suppress them.

    Proposed shape. GET /modules/{name}/transactions?depth=N → [{ module, function, lineNo, kind: END|BACKOUT, guards, viaCopycode, via }], guards as in item 182. The Java counterpart (@RunInTransaction, @Transactional) is the annotation half of item 190.

  • 184. Data flow between modules called one after another, and qualified PDA names, return an empty 200

    • W-WIF-A1.P-NIF-PERSONA is written in WPARTD1S:374 and read in WADDRD1S:1056. Both are called in sequence from WPARTX1S with the same PDA by reference; neither calls the other. field-flow finds no edge, because it pairs a producer with consumers reachable via CALLS from the producer.
    • MSG-INFO.##MSG-NR travels ISINSOIN (writes 4218) → DPARTEN0 → DPARTFN0 → YPARTMNH (RESET MSG-INFO). This chain is how an invalid social insurance number is silently accepted by Natural — the single most surprising finding of the run — and it was found by reading.

    Probes (all 200 []):

    GET /upms/variables/NIF-PERSONA-SAMPLE/flow-backward?module=DPARTFN0
    GET /upms/variables/%23%23MSG-NR/flow-forward?module=ISINSOIN
    GET /upms/variables/MSG-INFO.%23%23ERROR-FIELD/reads?module=DPARTEN0   (unqualified ##ERROR-FIELD → 21 rows)
    

    Item 146 covers the positional argument→parameter hop; this is the case beside it — sibling calls sharing a by-reference PDA, where the producer's caller is the consumer's caller. Two asks: (a) derive producer→consumer pairs across sibling CALLNATs of one caller (order-imprecise is acceptable if flagged, as field-flow already is); (b) a qualified name GROUP.FIELD that resolves to nothing should answer 404 or a hint, not an empty 200 that reads as "nobody reads it".

  • 185. Validation rules are not extractable — per action: condition, message, error field, return code

    Rebuilding DPARTEN0's rule list for #ADD (≈ 40 rules) and DADDREN0's (≈ 12) was the bulk of the run, entirely by reading. The rules live half in copycode — ISI173C1 (4080/4216), ISICMAND (126), USIX058C (8011) — and dispatch-table shows only some of the expanded assignments. Item 46a already splices copycode at ingest with &n& substitution, so the graph has the material; there is no view on it.

    What decides correctness is not the message number alone but what the caller does with it: DPARTEN0:217-222 raises only IF MSG-INFO.##ERROR-FIELD NE ' ', so ISINSOIN, which sets ##MSG-NR without an error field, never stops the insert. A rule list without the "error field set / return code set" columns would have missed it.

    Proposed shape. GET /modules/{name}/rules?action=ADD → [{ lineNo, viaCopycode, includedAt, condition, msgNr, errorField, returnCodeSet, escape }], plus GET /modules/{name}/source?expandCopycode=true&includedAt=<line> for the substituted text of one include. Pairs naturally with item 149 (the Java side of the same message number).

  • 186. Assignments carry no formats — implicit conversions are invisible; MOVE BY NAME is not resolved

    • DPARTFN0:480 YPARTMA1.NIF-PERSONA := VNUMEGAA.P-NUMBER assigns N12 to A14. Whether the new partner id is 000000012345 or 12345 decides whether the Java port (String.valueOf(…)) is correct — rated critical in the run, but left a hypothesis for lack of exactly this information. variables/NIF-PERSONA/writes returns assignedValue: "VNUMEGAA.P-NUMBER" and nothing about either format.
    • DPARTFN0:1063/1114/1165/1216 MOVE BY NAME Y…ROW TO Y…MA1 copies every same-named field, including the sort keys VAL-SK-* and LOG-*-END that the Java recomputes or resets. Which fields match had to be worked out by diffing two data-structures/…/fields responses by hand.

    Proposed shape. On writes: sourceDataType, targetDataType and conversion (N→A zero-padded, A→N, truncating, …) where both sides resolve. For MOVE BY NAME: GET /variables/{struct}/move-by-name?module=&lineNo= → matched field pairs, plus later writers of each target field in the same module.

  • 187. XML web-service routing and the tag contract are not queryable

    P1-m resolves W-MNT-N0 → KDWWIFN0 → Wxxxx*S into edges, but the routing key is lost: GET /upms/modules/WPARTX1S/callers → [W-LST-N0, W-MNT-N0] (CALLNAT_DYNAMIC), with no trace that objects PartnerCopy and PartnerAddress both route here (KDWWIFN0.nat:888-891). The file x-docs/kdwwifn0-prog-routing.csv already holds this mapping outside the graph.

    The second half crosses projects: the old adapter (app) emits tags such as cod_salut_id, ind_copyable (always '0'/'1' when a Vermittler exists — the root of a Java 8011 regression) and cod_addrtype_id = "1"; WPARTD1S consumes them in DECIDE ON #W-TAG. Which request field reaches which PDA field was reconstructed from both sources.

    Proposed shape. (a) the routing key on the dynamic edge or dispatch-table rows ({ objectName, adapter, target }); (b) GET /modules/{name}/xml-tags → consumed tags with target field and line; the producing side in app rides on item 143's cross-project edge, not a traversal.

  • 188. sql-statements cannot be filtered to one call site

    "Which statement, and which ORDER BY, does the browse at DPARTFN0:1048 run?" (YADDRBNH, two leading key components). sql-statements?depth=2 on DPARTFN0 returns 555 statements. Suggest ?via=YADDRBNH and ?callSite=DPARTFN0:1048 (or ?function=COPY-ADDR-COMM-BANK-DATA), returning only the statements reachable from there. The ordering mattered: Natural browses by NUM_ADDRESS, DAT_START, V_ISN, the Java repositories by number only.

  • 189. Java call-tree and db-accesses on pur are too noisy to use for "what does this logic reach"

    GET /pur/modules/PartnerCopyLogic/call-tree?depth=3   → 178 items, all type MODULE, among them var,
                                                            String, Math, Integer, LOGGER, java.util.List,
                                                            constants (COUNTRY_AT, PARTTYPE_LEGAL, …)
    GET /pur/modules/AddressLogic/db-accesses?depth=3     → 203 rows over 157 tables: 184 DECLARES,
                                                            14 READS, 5 WRITES
    

    Both agents on the Java side gave up and read the classes. Suggest excluding JDK/java.lang types, local-variable type names and constant holders from call-tree by default (a ?includeTypeRefs=true escape hatch), and restricting DECLARES at depth>0 to entities actually used by a reached repository method.

  • 190. Annotation semantics — which interceptor runs, and what it does — are not reachable

    @RunInTransaction(type = READ_WRITE) and @RequiresRole(roles = {}) on the endpoint decided two findings: whether an IllegalArgumentException after the partner store rolls it back, and whether the open role list matches Natural's ZINXSEC (it does). search/annotation finds the annotation sites but not the @InterceptorBinding → interceptor class → @AroundInvoke method chain. Suggest GET /pur/annotations/{name}/interceptors → [{ interceptor, aroundInvoke, sourceFile, lineNo }]; the rollback rule itself stays a source read, but finding the class should not be.

  • 191. Reference data behind the code is out of reach (low priority — may be out of scope)

    Four findings stay hypotheses because they depend on DB content, not code: key-table formats (IN-FORM-CLAVE, IND-CLIENT — does CODES-TO-INT zero-pad this code?), CONSTDAT values per client (NATPERS, AUSTRIA, NONTRADE, UNKTAXOF, MAINADDR, hard-coded in the Java), whether '00000001' IS (N1) holds, and the INSTLDA routine suffix for number range PART-NR. Not a graph problem; filed so the boundary is explicit. If a reference-data export exists, a read-only GET /reference-data/{table} would turn four hypotheses into lookups.

Frontend — React/TypeScript analysis (2026-09-22)

The pur backend (Java, ingested) is driven by a React/TypeScript frontend at /home/ingo/deve/uniqa/pur-sources/frontend (npm workspaces pur-ui, pur-ui-common, pur-r-vstamm, pur-r-vbuch; 253 .tsx + 213 .ts, ~31k hand-written lines). Scope decided 2026-09-22: only pur-ui and pur-ui-common are registered (project purfe, excludeDirs pur-r-vstamm, pur-r-vbuch); both use the legacy generated client against backend pur. Nothing of it is in the graph. The questions that cannot be asked today: which component reads/writes which store field, which backend endpoint does this page call, which DTO field is bound where, which theme token is used by whom. Items 192–196 add a typescript language with a Tier-1 regex coarse scanner in Java and a Tier-2 Node sidecar on the TypeScript compiler API (whole-program, type-aware; measured 3–5 s and ~0.5 GB heap per workspace). Decided 2026-09-22 after PROPOSAL + VALIDATE. 192 (parser module, sidecar, wiring, import/call graph, LoC), 193 (web-service calls in rest-endpoints, COUNTERPART_OF to pur), 194 (the Redux store), 195 (DTO field bindings) and 196 (styling) are implemented — 2026-09-22, see features.md; the validated design notes live there.

Ingest performance

  • 180. The DOCUMENTS edge carries two strings and nothing else — 20.2 % of all edges (found 2026-09-07 while measuring item 179)

    The finding. MODULE_COMMENTS is the only query that touches comments, and it finds them by sourceFile, not by traversal — its own javadoc says so. The DOCUMENTS edge appears there once, as an OPTIONAL MATCH, purely to fill two scalars in the response:

    MATCH (c:AstNode {type:'COMMENT', project:$project, sourceFile: m.sourceFile})
    OPTIONAL MATCH (c)-[:DOCUMENTS]->(t:AstNode)
    RETURN ..., t.name AS target, t.type AS targetType
    

    DOCUMENTS occurs zero times in GraphRepository and nowhere else in CypherQueries. Nothing traverses it in either direction. So 430 075 relationships — 20.2 % of the entire edge population of upms — exist to carry two strings into one endpoint. That is a modelling error regardless of what it costs.

    Lever A — replace the edge with two properties. targetName and targetType on the COMMENT node, set at parse time, where CommentBlocks.documentedNode(targets, startLine, endLine) already computes the target. Measured cost of the edges (item 179's A/B removes exactly these edges from merge-edges and nothing else):

    pair 1 pair 2
    DOCUMENTS edges 4.9 s 8.0 s
    comment nodes (merge-nodes) 18.4 s 27.2 s

    Removes 430 075 relationship merges and ~860 k endpoint index seeks, makes MODULE_COMMENTS faster (two property reads instead of an OPTIONAL MATCH), and costs no fidelity: same nodes, same text, same line numbers, same response shape, no API change, no test rewrite. Work needed: parser sets the two properties; MODULE_COMMENTS drops the OPTIONAL MATCH; DOCUMENTS removed from EdgeType (or deprecated); a reap for the existing edges on first refresh.

    What it gives up: a comment is then no longer reachable from the declaration it documents in Cypher. Nothing does that today, but a future "show me the comments on this field" would go by (sourceFile, line range) or want an index on targetName.

    Lever B — one COMMENT node per documented target (430 075 -> 105 160; comments attach to only 105 160 distinct targets, mean 4.1, and 61 716 targets have exactly one). This is the larger half of the cost, but all 17.6 MB of comment text stays, so it saves the per-node overhead and not the writes — realistically 10-15 s of the measured 18-27 s. It costs the /comments response shape, its tests, and per-block node identity.

    Recommendation: do A, leave B. A is a strict improvement — less data, less work, a faster query, nothing given up — and worth doing even at zero time saved. B trades real fidelity for a few seconds and should wait until someone needs them. Caveat on A's number: part of the 5-8 s is the endpoint seeks, so the realised saving may land at the lower end.

    See item 179 for the measurement and the method (reversed arm order, ratios rather than seconds).

  • 179. What comment nodes cost: measured, 17 % of a deep refresh (2026-09-07)

    Item 141 made comments graph data. They are now 45.8 % of the upms node population (430 075 of 938 746) and 60.4 % of everything merge-nodes processes, plus 430 075 DOCUMENTS edges — 20.2 % of all edges, and 17.6 MB of text, the largest write payload in persist. This measures what that costs. No model change was made; comments stay exactly as they are.

    Instrument. agenticcode.ingest.comments.enabled (default true) on AstIngestService, applied in both persist and persistBatch — the single-file path would otherwise have kept comments while the batch path dropped them. Filtering happens at the persist seam, not in the parser, so parse cost stays in both arms and the delta isolates persist. Deliberately a config property and not a query parameter: an experiment gets no REST or CLI surface.

    Two pairs were needed, and the second one is why this entry is trustworthy.

    OFF as % of ON pair 1 (OFF ran first) pair 2 (ON ran first)
    persist 77 % 77 %
    merge-nodes 33 % 31 %
    merge-edges 91 % 87 %
    commit 82 % 88 %
    finalize 121 % (artefact) 84 %
    refresh total 98 % (contaminated) 83 %

    Result: comments cost ~17 % of a deep refresh, of which persist is the larger and best-established part at 23 %. merge-nodes drops to roughly a third, matching the 60.4 % share of what that statement processes. Finalize also benefits (84 %) — fewer nodes, less to scan — which was not predicted.

    The prediction was wrong, and low. 25-30 s was estimated from proportional arithmetic; the real figure is 40-75 s depending on machine speed. The node/edge counts, by contrast, were predicted exactly (938 746 -> 508 671, 2 123 868 -> 1 693 793), which is what made the timings worth reading at all — the filter was verified before the clock was.

    The cold-first-run artefact — the reason arm order must be reversed. Pair 1 showed finalize 21 % slower without comments, concentrated in link-args-to-params at 47.4 s against 27.5 s. That step provably cannot see a comment node: comments carry exactly one edge type, DOCUMENTS, no CONTAINS and no incoming edges at all, while the step traverses CONTAINS and seeks on (project, sourceFile, paramPosition). Reversing the order settled it — in pair 2 the step is 25.6 s vs 27.9 s, i.e. equal. The 47.4 s was the first measured run of the session, not the comment-less arm. Any future A/B here must run both orders, or discard its first run.

    Ratios, not seconds. The machine ran ~25 % slower during pair 2 (load average 3.8 -> 6.7 on 8 cores, a VirtualBox VM alone at 163 % CPU), so absolute seconds were not comparable across pairs — but the persist ratio came out at 77 % in both. Expressing an A/B as a ratio of arm to arm within one run cancels machine drift; this is the method to use when the host cannot be quiesced.

    If someone wants to spend this 17 %: the obvious shape is one COMMENT node per documented target instead of per block — 430 075 -> 105 160 nodes, since comments attach to only 105 160 distinct targets (mean 4.1 each, 61 716 targets have exactly one). That should capture most of the persist saving while keeping comments queryable, at the cost of reworking /comments, its response shape and its tests. Not attempted.

    Raw data (durable, outside the repo): /home/ingo/ac-measurements/item179/ and item179b/ — per-arm server logs, counts, and run logs with load averages.

  • 156. The five field-resolution steps expand wide and filter late (measured 2026-09-05 with refresh?profile=true, item 155)

    After the placeholder sweep fix (item 154), finalize at ~600 s is the larger half of an upms deep refresh (persist: ~310 s). Five steps account for 88 % of it, and all show the same pattern — the cost is in the seeking, not in the writing:

    step time measurement
    resolve-bare-included READS/WRITES ~2x 165 s done 2026-09-05 (item 158): 298 s -> 153 s, INCLUDE-driven instead of placeholder-driven
    link-args-to-params 105 s done 2026-09-05 (item 157): 105 s -> 66.8 s by hoisting the parameter lookup
    resolve-field-placeholder READS/WRITES ~2x 50 s done 2026-09-05 (item 159): 99 s -> 40,7 s, resolution-first instead of reference-driven

    What has already been tried and rejected, so nobody repeats it: deduplicating the INCLUDES lists gains nothing — 52 995 edges stand against 52 975 distinct files, i.e. 20 duplicates in the entire project. For resolve-bare-included, looking the candidate up by an index on (project, name) first and checking containment afterwards was, in its naive form, slower (still running after 10 minutes, aborted) — field names are not selective enough project-wide. Refined 2026-09-06 (item 163): the diagnosis was right, the conclusion too broad. The name alone really is unselective (9 363 762 hits on upms — item 163 recorded this as 944 746, which was the count after the file filter); selectivity only comes from combining it with the module's include files (6 973). With realv.sourceFile IN includedFiles as a prefilter and a WITH DISTINCT barrier in front of it, this was the fastest shape known until item 167 replaced it with a per-include-file seek — without the barrier Neo4j plans the filter behind the SemiApply and reproduces exactly the failure recorded above. For link-args-to-params no index helps: the individual seek is efficient at ~11 rows, it simply runs millions of times, because the query forms a cross product over the include files of both sides times the argument list. Falsified 2026-09-06 (item 164). It holds for the caller side (cv), which already uses the 4-property index correctly — but not for the callee side: paramPosition was in no index, so pv sought on (project, sourceFile) and filtered afterwards, reading 62 048 947 rows for 29 484 hits. An index (project, sourceFile, paramPosition) takes the step from 66.6 s to 25.2 s by A/B and costs nothing measurable at persist, because a composite index only holds nodes carrying every property — 1 988 of 938 746. That is the difference from the merge-key index below, which spanned the whole graph.

    Status 2026-09-05: all three items are done (157, 158, 159), and the persist phase has since been measured and fixed as well (item 160: 290 s -> 218 s). The deep refresh is at 552 s instead of 1 225 s (-55 %).

    | what is left | time | note | |---|---|---| | resolve-bare-included READS/WRITES | ~~~136 s~~ | done 2026-09-06 (items 163, 167, 169): 135.8 -> 90.2 s name-driven, -> 53.1 s per-include-file seek, -> 40.1 s by resolving once per distinct (file, name) pair instead of 6x per module | | link-args-to-params | ~~~61 s~~ | done 2026-09-06 (item 164): 66.6 s -> 25.2 s via a (project, sourceFile, paramPosition) index | | merge-nodes | ~~~70 s~~ | done 2026-09-06 (item 165): merge key split by node kind; with merge-positional 140.1 s -> 54.3 s | | merge-positional-nodes | ~~~52 s~~ | done 2026-09-06 (item 165), see above | | merge-edges | ~~~47 s~~ | investigated 2026-09-06 (item 166): no structural lever, see below | | resolve-field-placeholder READS/WRITES | ~~~45 s~~ | investigated 2026-09-06 (item 168): write-bound, no lever — see below | | resolve-view-alias-* (3 steps) | ~~~33 s~~ | done 2026-09-06 (item 162): 32 s -> 2.4 s, alias-driven instead of access-driven | | commit | ~~~32 s~~ | measured 2026-09-06 (item 170): it really is the commit — unattributed and tx-open 0.0 s each. No lever. |

    Item 172 (2026-09-06): link-args-to-params — cost located, remedy measured and rejected. The step is essentially all read (read side 23.5 s of a 26.2 s step), and the cost is the number of index seeks: the planner expands cv.type IN [3 types] into three seeks per include file, so 729 317 x 3 = 2.19 M. Diagnostic with a single type: 23.6 s -> 14.5 s, i.e. ~6.2 us per seek.

    Three remedies measured, all without effect:

    • Hoisting the include-file list per module instead of per slot (29 484 -> 2 800 expansions): 2.4 %, inside the noise.
    • Anchoring on the callee instead of on every project node (saving 938 746 node reads and 3.28 M db-hits in the expand): no effect at all. db-hits mislead here, as they did in items 165 and 169.
    • Cross-module deduplication of (file, argument name) pairs (factor 6.53, as in item 169): fails on the join-back. Item 169 could rebuild it from the graph through an index; here one would have to find call sites by an argument name at a given position, which is not indexable, and an in-memory join would be a cross product of 29 484 x 19 473.

    The index (project, sourceFile, name) works, but costs more than it saves. A/B on a quiet host, control step resolve-bare-included READS 17.1 s against 17.3 s:

    | | with index | without | |---|---|---| | link-args-to-params | 13.7 s | 23.6 s | | finalize | 134.5 s | 140.3 s | | merge-edges | 55.6 s | 46.1 s | | commit | 39.9 s | 36.8 s | | merge-nodes | 26.9 s | 24.1 s | | deep refresh | 314 s | 302 s |

    Net +12 s, so rejected. Worth carrying forward for index decisions: the copySite index (item 165) covers 66 445 nodes and costs ~13 s; this one covers all 938 746 and costs ~15 s. Write cost does not scale with index size — it depends on how many written nodes touch the index.

    Item 171 (2026-09-06): merge-nodes re-measured, no further lever. After item 165 the step sits at 26.5 s. Two suspicions checked, both refuted:

    • The plan contains two Eager operators (property map and dynamic label), each materialising the whole batch. Dropping either changes nothing measurable.
    • Item 165 had moved node.ownerModule = '' out of the MERGE pattern into a SET, so it now runs on every merge rather than only on create — the suspicion being a self-inflicted regression. The ON CREATE variant was in fact slower by median.

    Measured over 60 000 re-merges, five rounds per variant; medians 2 422 / 2 532 / 2 352 / 2 359 ms with ranges that overlap throughout (e.g. 2 298-2 864 against 2 307-2 587). No difference is separable from the noise.

    Item 168 (2026-09-06): resolve-field-placeholder investigated, no worthwhile lever. The trick from item 167 (seek per known file over the 4-property index instead of searching broadly) does not transfer: the walk (real)-[:CONTAINS*1..10]->(realv {name}) is already bounded to one data area, not project-wide. Two rewrites measured, neither an improvement — seek plus EXISTS instead of the walk: 1 978 ms against 1 736 ms (worse); one walk per real with an in-memory join: 1 567 ms against 1 728 ms (inside the noise).

    Why no lever is to be expected: the step is write-bound. It creates/deletes 373 950 and 291 950 edges and writes 2 608 773 properties in 44.9 s. After item 167 resolve-bare-included writes only 636 104 properties and takes 53.1 s — this step does four times the write work in less time. The read side is not the bottleneck here.

    Item 166 (2026-09-06): merge-edges investigated, no worthwhile lever — recorded so nobody re-opens it. Two findings, both measured:

    • The plan contains an Eager, caused by SET r += e.properties: Cypher cannot rule out that the map contains lineNo, which the MERGE reads, so it materialises the whole batch. Neither merging the two SET clauses nor ON CREATE/ON MATCH removes it; only purely scalar assignments do. Cost in an isolated benchmark (60 000 edges, three alternating rounds): 3 371 ms against 3 202 ms, i.e. ~5 %, or 2-3 s. The price would be replacing the dynamic property map with fixed fields — not worth it for a twentieth.
    • The rest is genuine write work: 2.7 M edges, endpoints sought exactly over (ingestGen, nid), mean out-degree 2.26, so MERGE's existence check walks very short chains.

    Page-cache hypothesis tested and refuted. The store had grown to 2.9 GB (item 161: 2.4 GB) against an unchanged 2 GB cache, which made eviction the obvious suspect. Measured: 11.2 MB of read I/O over an entire refresh, against ~120 MB per run at the time of item 161. The cache is not the constraint; 3 GB would have taken a gigabyte from the host for nothing.

    The write side of persist was profiled on 2026-09-05 and turned out to be half lookup, not write: MERGE_NODES keys on five properties (type, name, sourceFile, project, ownerModule) while the widest index has four, so the seek hits (project, sourceFile, type, name) and a Filter discards the rest. For copycode that rest is enormous — 53 740 nodes sit behind only 804 distinct (sourceFile, type, name) keys (mean 66.8, max 3 542), so the seeks deliver 54 M rows to place 53 740 nodes. Measured: 161 owner lookups on one key cost 570 262 db-hits.

    Experiment, run and reverted 2026-09-05 — a net loss. Do not repeat it. Adding (project, sourceFile, type, name, ownerModule) did exactly what it promised locally and cost more than it saved globally.

    Explained 2026-09-06 (item 165), and a way around it found. The mechanism is index size: every node carries ownerModule ("" for module-own ones), so an index over it spans the whole graph. For 712 104 of the nodes the property is redundant in the key — (project, sourceFile, type, name) is already unique there — and load-bearing only for the 17 198 copycode nodes. Splitting the merge by node kind and indexing a property that only copycode nodes carry takes merge-nodes + merge-positional-nodes from 140.1 s to 54.3 s; re-measured on a quiet host the index costs ~13 s of write maintenance against ~72 s saved, net -86 s.

    The original measurement of the discarded attempt:

    | | without index | with index | |---|---|---| | merge-nodes | 69.5 s | 41.4 s | | merge-positional-nodes | 52.1 s | 18.4 s | | merge-edges | 44.9 s | 59.1 s | | commit | 32.9 s | 40.2 s | | persist total | 217.9 s | 182.1 s | | finalize total | 316.4 s | 429.6 s | | deep refresh total | 552 s | 628 s |

    The two MERGE lookups gained 62 s, as predicted. But every finalize step got ~33 % slower — resolve-bare-included 73.1 -> 97.1 s, link-args-to-params 67.4 -> 90.6 s, resolve-field-placeholder 18.1 -> 24.1 s — a uniform surcharge, not one bad step. The cause is not the query shapes: server.memory.pagecache.size was 1 GiB against a 2.4 GB store with 1.2 GB of indexes. A fifth wide index (carrying the 59-char sourceFile, cf. item 111d-2) multiplies page-cache misses. Note the cost of a miss here is not disk I/O — item 161 measured 237 MB of block reads across two whole refreshes — but the syscall-plus-copy path through the OS file cache. The index was dropped again.

    What that opens up, and it may be the largest lever in this whole list: nothing here was ever tuned for memory. Raising the page cache costs no code and no schema change. Tracked as item 161.

    Status 2026-09-06: the deep refresh is at 492 s, down from 1 225 s (-60 %). Every finalize step above 30 s has now been rewritten once (items 157-160, 162); what remains is either already optimised or documented as a dead end.

    Levers (unproven, to be measured in this order): make resolve-bare-included's qualifierGroup clause cheaper (it doubled the read share from 22 s to 47 s, measured before the rewrite); merge-edges (~43 s) is the only large step with no index support at all — all nine indexes are node indexes — though whether a relationship index even applies to a MERGE between two already-bound nodes is unknown and should be settled with EXPLAIN before anything is created; and — the biggest but riskiest lever — restrict finalize to changed modules, for which the *_SCOPED variants already exist but a whole-root refresh does not use them.

    Falsified 2026-09-06 — raising the ingest batch size. Do not retry it. agenticcode.ingest.batch-size (default 200) looked like the cheapest lever: pure configuration, no code, and commit costs 26.3 s across 32 batches. Doubling it to 400 failed the refresh: Neo4j rejected the transaction with The allocation of an extra 2.0 MiB would use more than the limit 1.4 GiB — dbms.memory.transaction.total.max threshold reached, the driver retried, the retry failed too, and the run aborted after 4 of 16 batches with 0 of 51 finalize steps — leaving the graph half-updated (1 031 039 nodes instead of 938 746; the ~92 k surplus were placeholders finalize would have resolved away). A full deep refresh repaired it.

    The ceiling is not the ac-code-server heap (which peaks at 1 924 MB of its 2 500 MB cap and was the wrong thing to watch) but dbms.memory.transaction.total.max = 1.4 GiB, Neo4j's default of 70 % of the 2 GiB heap. At 200 files a batch already sits close enough to that limit that doubling breaks it — so 200 is not a conservative default, it is just under a hard wall.

    Raising the wall means more Neo4j heap, which means a bigger container limit, and heap competes directly with the page cache whose whole measured worth is 6 % (item 161). An intermediate value (250-300) stays possible, but the prize is small: halving commit entirely would be ~13 s of 517 s (2.5 %), against the risk that just materialised.

    Also already falsified, from item 159: simply reordering the MATCH clauses is not enough — the planner reverts to the old order unless a WITH barrier pins it, and expressing an existence constraint as a MATCH instead of EXISTS { } lets the planner source the wrong driving node and build a cartesian. Both shapes must be re-checked with EXPLAIN after any edit to these queries.

  • 161. Neo4j's page cache holds 42 % of the store — raised 1 G to 3 G, measured: 314 s -> 302 s (2026-09-05, fallout from the reverted index experiment in item 156)

    | | | |---|---| | store on disk | 2.4 GB (1.2 GB range indexes) | | server.memory.pagecache.size | 1 GiB -> raised to 3 G | | coverage | ~42 % -> full store with room to grow | | mem_limit | 5 g -> 7 g (heap 2G + cache 3G + ~1G reserve) |

    Why this is believed to matter: the index experiment in item 156 slowed every finalize step by a uniform ~33 % (resolve-bare-included 73.1 -> 97.1 s, link-args-to-params 67.4 -> 90.6 s). A uniform surcharge across unrelated queries is not a query-shape problem; the fifth index cost cache pages that traversals needed. If eviction can cost 33 %, coverage should be able to buy something back. That is an inference from one experiment, not a proof.

    The counter-argument, which has not been ruled out: the host holds ~11 GiB in buff/cache, so Linux may already keep the whole store in its own file cache. A Neo4j cache miss would then cost a read() syscall plus a copy rather than disk I/O, and the gain would be small. Both observations can only be reconciled by assuming the syscall path is itself expensive enough to matter.

    How to measure it honestly: Community Edition cannot pre-warm the page cache, so the first refresh after the container restart runs cold while the 552 s baseline had a 12-hour warm cache. A result of "552 s, no change" would therefore prove nothing. Run one warm-up refresh, then measure. Compare against persist 217.9 s / finalize 316.4 s / total 552 s.

    If it works, item 156's index verdict has to be re-taken — that index gained 62 s on merge-nodes + merge-positional-nodes and lost 113 s in finalize purely to eviction. With the store resident, the gain might survive without the penalty.

    Measured 2026-09-05 (warm-up run then measurement run, as prescribed above): -33 s, -6 %.

    | | 1 G cache | 3 G cache | |---|---|---| | persist | 217.9 s | 207.1 s | | finalize | 316.4 s | 298.9 s | | deep refresh total | 552 s | 519 s |

    Even the cold warm-up run came in at 540 s. Within finalize the gain sits entirely in the big traversal steps and is uniform but small: resolve-bare-included 73.1 -> 68.7 s and 75.3 -> 70.8 s, link-args-to-params 67.4 -> 61.4 s (each -6 to -9 %); the short steps did not move. In persist only commit gained clearly (32.9 -> 26.4 s); merge-nodes barely (69.5 -> 66.5 s) and merge-positional-nodes not at all.

    Why the gain is small — measured, not guessed. server.memory.pagecache.directio is false, so Neo4j reads through the OS file cache; the store is cached twice. The Neo4j container's block I/O over both refresh runs was 237 MB read (20 988 read IOs) against a 2.4 GB store — essentially nothing ever came from disk, at 1 G of cache or at 3 G. There was no disk I/O to save. What a page-cache miss actually costs is a read() syscall, a kernel-to-userspace copy and Neo4j's eviction bookkeeping: real, but nanoseconds rather than milliseconds. Those are the 6 %.

    This retracts the explanation first given for the item-156 index failure. That the fifth index "evicted pages traversals needed and pushed the load onto disk" cannot be right — there was no disk load. The same miss path explains both directions instead: far more misses cost 33 %, far fewer buy 6 %. Same cause, no thrashing hypothesis needed.

    Settled at 2 G / 6 g, and then measured too (2026-09-06): 517 s — indistinguishable from 3 G.

    | page cache | total | persist | finalize | |---|---|---|---| | 1 G | 552 s | 217.9 s | 316.4 s | | 3 G | 519 s | 207.1 s | 298.9 s | | 2 G | 517 s | 208.0 s | 298.8 s |

    So the whole gain is already had at 2 G; the third gigabyte buys nothing. That matters here because the machine also runs a 4 GB VirtualBox guest (Windows11, 2 vCPU) alongside, and 2 G still covers nearly all of the 2.4 GB store. Consistent with the mechanism: the OS file cache absorbs the misses either way, so what is being bought is the syscall-plus-copy path, and covering ~85 % of the store removes as much of it as covering 100 %.

    The bigger lever for that machine is not this setting at all: vm.swappiness is at the Debian default of 60, and /proc/vmstat shows 0.69 GiB swapped out over 14 hours with allocstall = 0 — i.e. the kernel evicts anonymous pages of idle processes with no memory pressure whatsoever, which is exactly what makes a backgrounded IDE feel sluggish later. Lowering it to 10 addresses that directly. And while the VM runs, ./manage-ac.sh stop frees 7-9 GiB, more than any cache tuning.

    Item 156's index verdict stays as it was. The hoped-for reconciliation (index gain survives once the store is resident) is not supported: with the store now effectively resident the whole refresh only gained 33 s, so there is no 113 s of eviction penalty to recover.

  • 111d-2. sourceFile — 59 chars, 96% over the inline threshold, ~51× duplicated (found 2026-08-05 while sizing container memory)

    neostore.propertystore.db.strings is 385 MB, the largest file in the 2.0 GB store (plus 321 MB property store). Average string bytes per node:

    | property | ⌀ chars | note | |---|---|---| | sourceFile | 59 | only 11,080 distinct values across 570,739 nodes — every path stored ~51× | | id | 36 | a UUID the docs themselves call meaningless across ingest generations | | name | 13 | genuine | | language / project | 6 / 4 | constant per project | | value / dataType | 8 / 4 | genuine |

    With id retired by 111d-1, sourceFile is what remains — and it is now the single largest redundant property left: 59 chars on 96% of nodes, but only 11,080 distinct values across 570,739 nodes, i.e. every path stored ~51 times over.

    Options: a (:SourceFile {project, path}) node with an IN_FILE edge, or a per-project integer index plus a lookup table. Either way every response that exposes sourceFile (most of them) needs a join or a resolution step, which is the real cost of this one.

    Shortening the property chain compounds with 111a/b/d-1: the ~36 dbHits it takes to reach type are a function of chain length, so removing a property speeds up reads of all the others.

    Far more invasive than 111d-1 — sourceFile appears 426 times in CypherQueries (against ~10 for the id), sits in the node MERGE key, leads two indexes, drives the per-file reconciliation sweep, carries the sourceFile = "" placeholder convention, and is a field on most response DTOs and throughout the web UI. Needs its own proposal and a hard look at whether the ~150 MB is worth that surface.

  • 111e. Is a persisted CONTROL_FLOW node per statement worth it? (raised 2026-08-05)

    147,483 CONTROL_FLOW nodes = 26% of all nodes, each carrying the same ~139 chars of mostly redundant metadata. Sources are always readable to a human or agent that knows the location, so the question is which queries genuinely need persisted statement-level nodes rather than a location plus an on-demand re-parse. Largest single reduction available, but open-ended: what depends on them has to be established first.

  • 132. changedOnly has a permanent re-parse floor: a file with no hashed shell is always "changed"

    Symptom. Two consecutive POST /ac/refresh?changedOnly=true runs with nothing edited in between: the first examines 40 files, the second still examines 34.

    Cause. sourceHash is stamped on a file's MODULE/DATA_STRUCTURE shell (item 41), so a file that produces no such shell — package-info.java, a .java test fixture, an empty file — has no stored hash at all. ac walks 509 candidates but only 475 carry a hash. The item-129 skip treats an unknown hash as changed, which is the safe direction, but those 34 files are then re-parsed on every incremental refresh forever. That is a ~7% floor on this project and will differ per project.

    Fix. Record the content hash per file rather than per shell — on the (:Project) or a small per-file node — so "unchanged" is answerable for files that contribute no module. Until then the floor is harmless but real, and it makes changedOnly's numbers look wrong to anyone who counts.

  • 133. The item-128 reference index costs 3-5x deep-refresh time (POSTPONED 2026-08-20 — the cost is accepted; deep refresh at this speed is fine. Revisit only if the loop becomes painful.)

    Measured 2026-08-19 (same machine, same corpus, before → after item 128): ac 19s → 106s, app 37s → 106s, pur 29s → 125s, upms 1230s → 1496s.

    The index is genuinely useful (PartnerController: 65 sites, incl. two files a rename pass had missed). Decision 2026-08-20: the trade is accepted — a 2-minute Java refresh and a 25-minute upms refresh are tolerable for what the index gives back, and item 129's ?paths= / ?changedOnly= cover the fast inner loop. Kept open, not closed, because the analysis below is still the right starting point if the cost ever does become painful: measure which emission dominates — imports (~26.6k on pur) vs declared type positions (~13.4k) vs annotation usages — then consider a narrower default (imports + annotations only, declared types behind a flag), or emitting mentions only in the deep tier rather than in the Tier-1 scan.

Web UI — code understanding & navigation

A React/TypeScript web UI (ac-ui/) for navigating the AgenticCode graph to migrate legacy Natural to Java. Full vision/architecture: x-docs/ui-proposal.md. Backend prereqs 48–50 and frontend milestones M0–M5 (+ the M6 explorer regex filter) are DONE — see x-docs/features.md. Only M6 scale/polish remains:

  • 51 (M6). Scale & polish (POSTPONED 2026-07-15) — virtualisation/large-graph performance, multi-project, auth, theming, export (SVG/PNG/report). (The rest of item 51 is complete; see x-docs/features.md for M0–M5 detail.) Key open risks carried from the earlier milestones: auto-ingest latency on large projects (deep-ingest concurrency cap defaults to 2); the STALE_SOURCE warm side-effect (a warming query re-ingests + re-hashes a module and thereby clears its staleness — treat ingest status as "true at query time"); node-id instability after re-ingest (the UI keys on name + sourceFile); auth/multi-user and the deploy model still open.

Not a bug — retracted findings

Probed, found to be wrong or already delivered, and recorded here so the same finding is not re-filed. Kept in this file rather than moved to x-docs/features.md: these are not implemented features, they are guards against repeating work.

  • Not a bug — retracted 2026-08-18 (was #127): "an ambiguous simple module name is resolved silently to one candidate". The probe picked a name that is not ambiguous: pur holds exactly one module with simpleName = Builder (…KeyBasedIterator.Builder) out of 5,119, so the 200 was the correct answer. Item 115's refusal works and was verified on names that really do collide — GET /pur/modules/AuthorizationInterceptor/context and /pur/modules/VoidConverter/context both answer 409 AMBIGUOUS_NAME with candidates and qualifiedNames in details. The guard is withModule/withIngestedModule, which reject on state.ambiguous(); MODULE_INGEST_STATE collects every real candidate by name or simpleName, so nested classes are covered.

    The item's one genuine complaint — "whether pur in fact holds more than one Builder could not be checked from the API" — was real, and is what item 125 now answers: GET /pur/search/identifier?name=Builder&type=MODULE lists every module sharing a short name.

  • Not a bug — retracted 2026-07-17 (was #71): "external subroutine calls do not count as a module hop" rested on a bad measurement of mine, not on the code. I counted 4,808 "genuine external subroutine calls" in upms with (ma)-[:CONTAINS*0..]->(src)-[:CALLS]->(f), applying the single-owner filter to the target but not to the source: for a copycode-shared source function (L4N-ENTER is CONTAINSed by 137 modules) that enumerates all 137 as ma, each differing from the target's owner. The correct test — source and target single-owner — returns 0. And it must: resolvePlaceholderTargets filters ph.type IN ['MODULE', 'DATA_STRUCTURE'], so a FUNCTION placeholder is never resolved across modules; a PERFORM to another module's subroutine stays an unresolved placeholder (6 in upms) and never becomes a cross-module CALLS edge. A path therefore cannot enter a module at a FUNCTION node, which was #71's entire premise. (The real, tiny gap: those 6 unresolved placeholders. Natural external subroutines are a language feature this parser does not resolve — worth its own item if the corpus ever needs it.)

Probed and dropped — do not re-file. 409 AMBIGUOUS_NAME already returns candidates and qualifiedNames in details (verified on pur/modules/AuthorizationInterceptor/context) — items 115 and 125 covered it. Incremental refresh is item 129 and delivered (?paths=, ?changedOnly=, ingest.incomplete). Serving decoded source by line range is delivered, see item 150 above.