Roadmap: vier Findings aus dem upms-Webservice-Audit (107-110) + Datenpunkt zu 26

Gefunden beim Beantworten der Frage, ob ein W*-Webservice-Modul eine
Provisionsberechnung ausloesen kann.

107 (Known bug, der folgenreichste): jeder /modules/{name}/-Endpunkt
antwortet 200 mit einer leeren Huelle fuer ein Modul, das gar nicht
existiert - digest fuer WXSPOD0S und fuer NOSUCHMOD123 sind identisch.
search/identifier liefert korrekt []. Dokumentiert ist 404. Das hat bei
mir zu einem falschen Ergebnis gefuehrt: call-tree auf 13 nicht
ingestierte Module las sich als "analysiert, nichts gefunden" statt
"nicht analysierbar". Gleiche Klasse wie 103.

108: dispatch-table versteht nur den DECIDE-Verteiler, nicht das
Dispatch-Tabellen-Idiom (ASSIGN #WT-OBJ-PROG(n) = 'MODUL' + CALLNAT
#W-ACT-PROG) - 10 der 116 offenen dynamischen Aufrufe. Verwandt mit 83,
aber ueber ein Array-Element statt einen Skalar.

109: variables/{name}/writes liefert den Ort, nicht den geschriebenen
Wert - fuehrt bis vor die Zeile und hoert dort auf. assignedValue gibt
es bei dispatch-table bereits.

110: keine Erreichbarkeitsabfrage. "Wer kann X ausloesen" musste als
Client-BFS ueber /callers gebaut werden: ~100 Requests fuer eine Frage,
die in Cypher ein gebundener shortestPath ist.

Zu 26: MCP-Fehler reproduziert, diesmal ab dem allerersten Tool-Aufruf
der Sitzung - spricht gegen Idle-Timeout, fuer eine nie zustande
gekommene Session. REST lief parallel normal.
This commit is contained in:
Ingo Schnabel
2026-08-02 19:24:46 +02:00
parent 38d5fafdc8
commit 3db477f4a3

View File

@@ -77,6 +77,40 @@ before. The parked "parallel parse phase" idea was implemented 2026-07-18 (item
## Known bugs
- [ ] **107. Every module endpoint answers `200` with an empty shell for a module that does not exist —
indistinguishable from a real but empty module** (found 2026-08-02, `upms` webservice-layer audit;
contradicts the documented `404 MODULE_NOT_FOUND`)
**Symptom.** Measured against the live server, `upms`:
```
GET /modules/WXSPOD0S/digest → 200 {"name":"WXSPOD0S","description":null,"functionCount":0,
"callers":{},"callees":{},"dbTables":[],"dataStructures":[]}
GET /modules/NOSUCHMOD123/digest → 200 {"name":"NOSUCHMOD123", … identical shell … }
```
Byte-identical answers apart from the echoed name — for a module that exists nowhere in the graph and
for a name typed at random. `callees`, `call-tree` and `context` behave the same (all `200`).
`search/identifier?name=WXSPOD0S&type=MODULE` correctly returns `[]`, so the graph *knows*; only the
module endpoints invent the row. `agent-api-system-prompt.md` promises
`404 NODE_NOT_FOUND / MODULE_NOT_FOUND — unknown id / module name`.
**Why it matters — this produced a wrong analytical result, not just an ugly response.** The question was
"can any `W*` webservice module reach the commission calculation?". Five dispatchers
(`WPOLIX0S`, `WGARCX0S`, `WACOMX0S`, `WCLAIX0S`, `WOBJPX0S`) dispatch dynamically to 13 targets
(`WXSPOD0S`, `WXSDAD0S`, `WXSCMD0S`, `WCARLD0S`, `WXSCLD0S`, `WXSCLD2S`, `WXSFUD0S`, `WXSGAD0S`,
`WACOMD0S`, `WCLAID0S`, `WOBJPD0S`, `WCATEX0S`, `WCATED0R`). `call-tree?depth=4` on each returned
`200` with 0 modules and 0 provenance hits, which reads as **"analysed, nothing found"**. The truth is
**"not analysable"** — all 13 source files are absent from the checkout (`find` finds none). An agent
that trusts the `200` concludes "these paths trigger no commission processing"; the honest answer is
"unknown". Same failure mode as item 103: an incomplete answer that looks complete.
**Fix.** Return `404 MODULE_NOT_FOUND` when no `MODULE` node exists for `(project, name)` — the check
`search/identifier` already performs. If a node exists but its source is not ingested (placeholder,
`sourceFile = ""`), that is a *different* state and deserves an explicit marker in the payload
(`placeholder: true` / `sourceFile: null`) rather than an all-zeros body. Worth checking which other
`/modules/{name}/…` endpoints share the shell (`functions`, `db-accesses`, `data-structures`,
`dispatch-table` all plausibly do — `dispatch-table` returning `[]` for a non-existent module is
currently indistinguishable from item 108's real gap).
- [ ] **75. `CONTAINS` is not acyclic — 22 self-loops and 162 two-cycles in `upms`**
(found 2026-07-17 while root-causing item 74; **cause NOT established — do not treat the notes below as
settled**). The containment hierarchy that dozens of queries traverse with `CONTAINS*` contains cycles:
@@ -809,6 +843,82 @@ lands in both ingest tiers at once.)*
*(Done items 52, 53, 54, 56 moved to `x-docs/features.md`.)*
- [ ] **108. `dispatch-table` only understands the `DECIDE` dispatcher, not the dispatch-*table* idiom —
the one the endpoint is named after** (found 2026-08-02, `upms` webservice-layer audit)
**Symptom.** Two dispatcher idioms are common in `upms`, and the endpoint covers one of them.
```
GET /modules/WSUBPX0S/dispatch-table → 25 rows (DECIDE ON VALUE 'supl_fin' → #W-ACT-PROG := 'WNSUPD0S')
GET /modules/WPOLIX0S/dispatch-table → []
GET /modules/WGARCX0S/dispatch-table → [] (likewise WACOMX0S, WCLAIX0S, WOBJPX0S)
```
The five empty ones dispatch through an **array built from literals**:
```natural
DEFINE SUBROUTINE INIT-OBJECT-TABLE
ASSIGN #WT-OBJ-PROG (1) = 'WXSPOD0S'
ASSIGN #WT-OBJ-PROG (2) = 'WXSDAD0S'
…
…
#W-ACT-PROG := #WT-OBJ-PROG (#I-OBJ)
CALLNAT #W-ACT-PROG …
```
Every target is a literal in the source; nothing is runtime-dependent. These are **10 of the 116**
entries in `dynamic-calls/unresolved`, all with `variable: #W-ACT-PROG`.
**Relation to item 83.** Item 83 constant-folds `MOVE '<lit>'` / `MOVE '<lit>' TO SUBSTR(var,pos,len)`
chains. This is the same class — literal assignment feeding a `CALLNAT var` — but through an **indexed
array element** rather than a scalar, so the fold does not apply. The set of possible targets is exactly
the set of literals assigned to any element of that array; which index is live at runtime is not
statically known, so the honest model is *n* candidate edges (as item 82 already allows for a manual
override with multiple targets), not one.
**Fix (proposal).** Extend the fold to indexed writes: collect `ASSIGN <array>(<n>) = '<literal>'` for
the array feeding `CALLNAT`, and emit one `CALLS` edge per literal that names a real ingested module,
tagged `folded=true` + something like `viaTable=true`. Independently, `dispatch-table` should report
them, keyed by index instead of by guard value — the response already has `guardValue`/`assignedValue`,
so `guardValue: "(3)"` or a dedicated `index` field would fit. **Caveat:** in `upms` most of these
targets are not ingested (see item 107), so the edges would resolve to nothing there — the value is in
no longer *silently* reporting `[]`.
- [ ] **109. `variables/{name}/writes` gives the location but not the written value, so resolving a
dispatch needs the source anyway** (found 2026-08-02, `upms` webservice-layer audit)
**Symptom.** Working around item 108:
```
GET /variables/%23WT-OBJ-PROG/writes?module=WPOLIX0S
→ [{"function":"INIT-OBJECT-TABLE","sourceFile":"…/WPOLIX0S.nat","lineNo":772, …}, … 6 rows]
```
Six correct write sites, and not one of the six assigned values. Answering "what does this dispatcher
dispatch to" therefore requires opening the file and reading lines 772/775/777/780/783/785 — the API
narrows the search to the right lines and then stops one step short. `dispatch-table` already returns
`assignedValue` for the `DECIDE` idiom, so the concept and the field name exist.
**Fix.** Add `assignedValue` (and, where the write is indexed, `assignedIndex`) to the `writes` rows,
populated when the right-hand side is a literal, `null` otherwise. Cheap next to item 108 and useful far
beyond it: "which constants does this module put into field X" is a routine question in a
reengineering pass.
- [ ] **110. No reachability query — "can A reach B?" has to be hand-rolled as ~100 `callers` calls**
(found 2026-08-02, `upms` webservice-layer audit)
**Symptom.** The question was "does *any* `W*` module reach the commission calculation
(`ISINCOMI` / `VCOMIN00` / `VVERAN50` / `VCOMIN55` / `VCOMIN57` / `VCOMIN50`)?" — a yes/no with a
witness path. There is no endpoint for it. `call-tree` goes **downward from one root** and returns a
flat closure without paths, so it answers "what does A reach", never "who reaches B", and never "how".
The workaround was a client-side breadth-first search upward over `/callers`, six seeds, depth 6: **~100
HTTP round-trips**, 101 modules visited, and the path reconstruction written by hand.
**Fix (proposal).** `GET /modules/{name}/reaches?target=<name>&direction=up|down&depth=N` returning
`{reachable: bool, paths: [[module, …], …], truncated: bool}` — or, more useful for this shape of
question, a filtered variant of `callers`/`call-tree` that accepts a **set** of targets and returns only
the witnesses. In Cypher this is one bounded `shortestPath`/variable-length match; done client-side it
is 100 requests and an easy place to introduce a bug. Note the traversal must be bounded — see item 75
on the `CONTAINS` cycles.
**Why it matters.** "Who can trigger X" is *the* recurring question in legacy reengineering: which entry
points reach a calculation, a table write, an external interface. It is the natural counterpart to
`call-tree` and currently the biggest hole in the query surface for that work.
- [x] **82. Manual override for unresolvable dynamic `CALLNAT` targets (human/agent-settable)**
(proposed + implemented 2026-07-19, from the WGEAGB0S deep-API audit; REST + MCP + `ac` CLI +
Testcontainers ITs green). The dynamic-`CALLNAT` resolvers cannot
@@ -859,3 +969,13 @@ lands in both ingest tiers at once.)*
handling on a new SSE stream, not a server-side bug this codebase's config
can fix. Left open pending either a `quarkus-mcp-server-sse` version bump
with related fixes, or evidence this is in fact server-triggered.
**2026-08-02 — reproduced, and worse than "intermittent" in this session.** Every MCP call failed with
the same message from the *first* call onwards (`mcp__agenticcode__module_dispatch_table`, then
`mcp__agenticcode__list_projects`), so the whole `upms` webservice audit ran over REST via `curl`
instead. Two observations that may narrow it: (a) it was not a degradation after a run of successful
calls — the very first tool call in the session failed, which argues against an idle/reconnect timeout
and for the session never being established; (b) the equivalent REST endpoints answered normally
throughout, so the server and the graph were healthy. Practical impact for agent work: the documented
MCP surface is unusable in these sessions and every playbook step has to be re-expressed as `curl`,
which is why the REST examples in `agent-api-system-prompt.md` earn their keep.