Files
agenticCode/x-docs/finalize-steps.md
2026-09-06 20:21:04 +02:00

19 KiB
Raw Permalink Blame History

Finalize pipeline: the 36 enrichment steps

Numbering caveat (2026-09-06). This document describes a 36-step pipeline. The pipeline as it runs today has 51 steps, so the step numbers below no longer line up one-to-one with the numbers in the Finalize <project> [n/51] log lines. The descriptions and the recorded constraints are still accurate; only the numbering is stale. Content is otherwise unchanged from the German original.

What each step of a finalizeProject run does — with the emphasis on its limitations and constraints. Source of the ordering: GraphRepository.enrichmentSteps(dataflow, resolveFields, scoped). The numbering below corresponds to a full deep run over the whole root (recreate?deep=true / refresh?deep=true): dataflow = true, resolveFields = true, scoped = false. In other modes steps are skipped (see "mode gating").

Modes and what they switch on/off

The step list is not fixed — three flags control which steps run:

Flag When true Switches on
(always) every run placeholder resolution at call-graph level (1-16), Java DB, inheritance, link-calls-to-implementations, cleanup/projection (31-36)
resolveFields deep ingest only field level: 17 (indirect), 18-26 (field/bare-field resolution + the associated deletes)
dataflow deep ingest only argument->parameter dataflow: 27-29, and building on it 30 (cross-module dynamic CALLNAT)
scoped per-program deep ingest replaces project-wide variants with …_SCOPED (only modules in $names), same ordering

Fundamental constraints of the pipeline as a whole:

  • Every step runs in its own transaction and is idempotent (all MERGE/SET, deletes are conditional). A step may safely be repeated — but a half-run deep refresh leaves the graph partially enriched: earlier steps are committed, later ones are not. Do not abort halfway.
  • The ordering is binding. Many steps assume that an earlier one has already resolved/built edges ( dependencies noted per step below). Reordering breaks correctness.
  • CONTAINS is not acyclic (item 75: self-loops on ","-named parser-artefact DATA_STRUCTUREs + shared copycode CONTROL_FLOW nodes). Every step with an unbounded CONTAINS* over module subtrees is therefore at risk of exploding. Where that was critical, it has been switched to sourceFile owner joins (16, 17, 27, 30) or bounded to a fixed depth (18-24).
  • Shared placeholder nodes (item 76: one node per (name, project), sourceFile="") are referenced by hundreds of modules. That makes cardinality-heavy steps (above all 19) expensive and limits the discriminating power of resolution.

Phase A — placeholders at call-graph level (steps 1-7)

For every edge type in RESOLVABLE_EDGE_TYPES the cross-file placeholder (sourceFile="") is re-pointed at the real target node.

1. resolve-placeholder CALLS — PERFORM/CALLNAT '…' targets onto the real module. 2. resolve-placeholder INCLUDES — USING copybooks/PDAs onto the real DATA_STRUCTURE. 3. resolve-placeholder USES_TYPE — type references. 4. resolve-placeholder EXTENDS, 5. IMPLEMENTS, 6. INJECTS, 7. REFERENCES — Java inheritance/DI edges.

  • Constraint: resolves only if a real definition with a matching name exists in the project. If the target is missing ( external module, copybook not ingested), the placeholder stays and is only marked "unresolved" in step 34 — not deleted (so that a later scoped deep ingest can still resolve it).
  • Constraint: purely name-based. Targets that share a name but differ semantically are not distinguished.
  • Always project-wide (no scoped variant) — including in call-graph-only mode, because it is cheap.

Phase B — Java-specific resolution (steps 8-15)

8. resolve-panache-inherited-entity — sets repositoryEntity on concrete Panache repositories whose entity is inherited via a project base class.

  • Dependency: needs resolved EXTENDS (after step 4). Constraint: project base classes only; inheritance external to the JAR is invisible.

9. resolve-java-db-access — JPA/Panache DB_ACCESS candidates onto the entity's DB_TABLE.

  • Constraint: needs persisted MAPS_TO edges (entity->table); without a mapping there is no resolution.

10. resolve-java-query-jpql / 11. resolve-java-query-native-sql — @Query JPQL and native SQL respectively onto DB_TABLE.

  • JPQL constraint: needs MAPS_TO. Native SQL: matches only a literal table name — dynamically assembled SQL is not recognised.

12. build-implemented-by / 13. build-overridden-by — materialises interface->impl and base-method->override edges respectively.

  • Dependency: needs resolved IMPLEMENTS/EXTENDS.

14. link-references-to-subclasses / 15. link-injects-to-subclasses — wires REFERENCES/INJECTS declared in ancestors onto concrete subclasses.

  • Constraint across 8-15: purely Java. For a pure Natural project these are no-ops (0 edges), but they run anyway (cheap).

Phase C — dynamic CALLNAT resolution, intra-module (steps 16-17)

16. resolve-dynamic-callnat-intra — CALLNAT <var> where <var> is assigned a string literal in the same module becomes a real CALLS edge (tagged CALLNAT_DYNAMIC). 17. resolve-dynamic-callnat-intra-indirect — like 16, but the program name is copied indirectly from another variable/lookup table (#PROG := #TBL(#I)).

  • Constraint (over-approximation): every literal the dispatch variable can take becomes a target. Correct for genuine dispatchers, but more edges arise than a single call really hits.
  • Constraint: literal assignments only. A program name read from a file/DB or indexed purely numerically stays unresolvable -> remains visible as a CALLNAT_DYNAMIC placeholder (step 34), it is not swallowed.
  • Bound/fix (item 75): module scope via sourceFile equality instead of (caller)-[:CONTAINS*0..]-> — otherwise cubic path explosion over cyclic copycode (step 17 hung for >2 h with it). A CALLNAT that sits physically in shared copycode (in upms exactly 1 case without a module of its own) is no longer attributed to every including program.
  • 16 runs in every mode (cheap), 17 only with resolveFields.

Phase D — field level: resolving placeholder fields (steps 18-26)

Only with resolveFields (deep). Re-points qualified/bare field references (READS/WRITES onto sourceFile="" fields) at the real field of the included data structure.

18. resolve-field-placeholder READS / 19. resolve-field-placeholder WRITES — qualified field reference ( CDPDA-M.SORT-KEY): find the real DATA_STRUCTURE of the same name INCLUDEd by the module, and within it the field of the same name.

  • Bound: the field lookup (real)-[:CONTAINS*1..10]->(realv) is bounded to depth 10 — groups nested more deeply (>10 levels) are not found. The field tree per structure is a pure tree (not a DAG), so it is cheap per real.
  • Constraint (item 76 — fixed 2026-07-18): since the #76 fix, field placeholders have a per-module identity (ownerModule in the merge key), so they are no longer shared project-wide. Before that, step 19 was by far the most expensive (~55 min on upms), because the shared node produced a ph × writer cartesian product (~28.7 M rows); now each phv_M has only the writers of its module -> step 19 in ~55 s (60x), field block 18-26 in ~2 min in total instead of ~57 min.
  • Constraint: resolves only on an unambiguous INCLUDES hit; ambiguous/missing includes stay unresolved.
  • Ordering (item 159 — 2026-09-05): the query first resolves (ph, phv) → real → realv and attaches the reference edges afterwards; real comes from the index (project, type, name) instead of a name scan over the module's INCLUDES list. The field subtree is thereby expanded once per (phv, real) pair (520 on upms) instead of once per row (~220 000). The INCLUDES still decides which data area applies — it just no longer finds real, it checks it. Two clauses carry this, both verified with EXPLAIN: the WITH DISTINCT as a planner barrier (without it Neo4j plans reference-driven again) and the EXISTS { (m)-[:INCLUDES]->(real) } as a predicate rather than a MATCH (as a MATCH the planner draws m from the includers of real and builds an m × src cross product). The _SCOPED variant deliberately keeps the module-driven ordering: it starts from $names and processes few modules, so there is nothing to hoist there.
  • Precondition of the ordering: the placeholder name has to stay selective enough that the index seek does not fan out wider than the INCLUDES list it replaced — on upms at most 4 real structures share a placeholder name (mean 1.45). The result stays correct in any case, since the INCLUDES check still filters; only the cost tips back.

20. resolve-field-placeholder-by-name READS / 21. …WRITES — fallback: the same resolution, but real globally by name instead of via INCLUDES (for GLOBAL USING/copycode without an INCLUDES edge).

  • Constraint: purely name-based and global — a weaker guarantee than the INCLUDES route; it can grab the wrong thing when structure names are duplicated project-wide. Runs after 18/19, so it sees only the remainder.

22. delete-resolved-field-contains — removes the now unused CONTAINS edges to resolved field placeholders.

23. resolve-bare-included READS / 24. …WRITES — unqualified (bare) field reference onto a field of a USING include, only if exactly one field of that name is reachable through the module's includes (item 18).

  • Bound (item 77): evaluated per module (m stays in the aggregation, src via (m)-[:CONTAINS*0..1]->), otherwise misattribution across module boundaries (previously: BMTABBP0 wrongly wrote ##MSG-NR of CDPDA-M, which it does not include). INCLUDE_FIELD_DEPTH = 10 is a correctness bound here (not tuning), because CONTAINS is cyclic.
  • Constraint: ambiguity (several fields of the same name) -> stays unresolved (conservative). Known remainder (item 76): 132 shared copycode FUNCTIONs in upms are one node for several modules — per-module binding cannot separate them.
  • Driving side (item 163, 2026-09-06): the query is driven by the placeholder name (realv via the (project, name) index), no longer by the include side. A realv.sourceFile IN includedFiles prefilter keeps the subtree walk small; it is semantics-bearing, not tuning, and presupposes that a field lies in the file of its data area (checked on upms: no cross-file CONTAINS edge out of a DATA_STRUCTURE anywhere in the project). The WITH DISTINCT in front of it is a planner barrier — without it the filter migrates behind the SemiApply and the step no longer completes. 135.8 s -> 90.2 s.
  • Driving seek (item 167, 2026-09-06): realv is looked up per include file via (project, sourceFile, type, name) instead of by the name alone. The name seek returned 9 363 762 rows in order to keep 6 973 — around 8.4 M of them placeholders of other modules (sourceFile = ""), which a name index cannot exclude. 9 363 756 db-hits become 6 973, and the step drops from 102.0 s to 53.1 s. Caveat: the type list now determines the number of seeks (placeholders × include files × types), no longer just a filter predicate.
  • Resolve once instead of six times (item 169, 2026-09-06): after item 167 the entire read time lay in the number of seeks — full read side 8 045 ms, prefix alone 7 997 ms, so EXISTS and aggregation together 48 ms. The seeks were redundant by a factor of 6.05: 330 408 (module, file, name) triples over only 54 630 distinct (file, name) pairs. The query now resolves once per pair and joins the modules back afterwards: 53.1 s -> 40.1 s. Caveat: collect materialises the hit list, so phase 1 has to run to completion before phase 2 begins.

25. delete-bare-placeholder-contains — removes CONTAINS to resolved bare placeholders (per-module condition, so that the still unresolved edges of another module do not keep the placeholder alive).

26. delete-resolved-placeholders — reaps the now fully resolved placeholder structures.

  • Constraint: safe only after the field placeholders are resolved — otherwise structures would be removed that a later scoped deep ingest still needs.

Phase E — dataflow argument->parameter (steps 27-29)

Only with dataflow (deep). Expensive, hence not in the fast whole-root pass.

27. link-args-to-params — links every positional CALLNAT argument of the caller with the parameter at the same position in the callee via ARG_TO_PARAM.

  • Bound/fix (item 75): with three unbounded CONTAINS* anchors this was the slowest finalize step (~ 25 min). Switched to sourceFile owner (call site) + sourceFile/INCLUDES scope (argument variable, callee parameter). paramPosition is unique within the callee scope.
  • Constraint: purely positional — no type/name matching; a wrong argument order in the source would be wired up wrongly. Natural only.
  • Index + caller module (item 164, 2026-09-06): the callee parameter is looked up via the index (project, sourceFile, paramPosition). Without it the seek stops at (project, sourceFile) and reads every candidate file in full — 62 048 947 rows for 29 484 hits. A/B shows: 66.6 s without, 25.2 s with the index; not measurable at persist, because a composite index only holds nodes with all the properties (1 988 of 938 746). The caller module now comes via (cm)-[:CONTAINS*0..1]->(src) instead of via src.sourceFile. That also captures call sites inside copycodes, which were previously skipped silently (a .cpy has no MODULE node of its own): ARG_TO_PARAM on upms 17 813 -> 17 871. A size(cms) = 1 guard keeps the step conservative in case a copycode FUNCTION is ever contained by several modules (item 76).

28. link-args-to-params-java / 29. link-args-to-params-java-cross — the same for Java, intra- and cross-class respectively.

Phase F — cross-module dynamic CALLNAT resolution (step 30)

30. resolve-dynamic-callnat-cross — CALLNAT <var> whose program name flows in from another module via a parameter: follows ARG_TO_PARAM (from 27) into the callee and takes the literals written into that parameter there as targets.

  • Dependency: needs the ARG_TO_PARAM edges from 27 -> dataflow only.
  • Bound: ARG_TO_PARAM*1..5 — flow over at most 5 hops; longer chains are unresolvable.
  • Bound/fix (item 75): call site via sourceFile, dispatch variable via sourceFile/INCLUDES scope (the name alone would catch ~20 958 foreign variables of the same name). Over-approximation as with 16/17.

Phase G — polymorphism, cleanup, projection (steps 31-36)

31. link-calls-to-implementations — fans class-level CALLS on an interface/base type out to the implementations/subclasses (CHA).

  • Constraint: class hierarchy analysis — over-approximates (all implementations count as possible targets), no precise target determination. Runs last, so that it sees resolved CALLS+IMPLEMENTS/EXTENDS.

32. delete-dynamic-callnat-placeholders — removes the CALLNAT_DYNAMIC marker edges (the unresolved dispatch placeholders); resolved edges remain.

33. delete-data-literal-call-placeholders — reaps placeholder MODULEs that are really data literals (browse keys) and were wrongly restored as call targets (item 62).

  • Constraint: a name-based test that only removes edges which are wrong everywhere.

34. stamp-unresolved-placeholders — marks leftover placeholders as (un)resolved, depending on whether a real definition now exists. Runs last, sees the fully re-pointed/reaped graph.

35. delete-calls-module / 36. build-calls-module — projects the finished call graph onto CALLS_MODULE module->module edges (so that traversals can be bounded in module hops rather than raw CALLS hops).

  • Constraint: a pure projection — only as correct as the graph at the moment of the run. Must run after all steps that add CALLS (placeholder/dynamic resolution, 31) or remove it (33). The preceding delete (35) makes it correct across refreshes (a MERGE alone would never remove stale edges).

Summary of the most important constraints

Topic Steps affected Constraint
Mode gating 17, 18-26 (resolveFields); 27-30 (dataflow) Without deep ingest, field and dataflow edges are missing; queries on them return empty/guarded.
CONTAINS cycles (item 75) 16, 17, 27, 30 (fixed); query endpoints (open) Unbounded CONTAINS* explodes over cyclic copycode; worked around via sourceFile owner.
Shared placeholders (item 76) 19 (cardinality), 23/24 (discrimination) One node per (name, project) -> step 19 very expensive; 132 shared copycode FUNCTIONs not separable per module.
Over-approximation 16, 17, 30 (dynamic dispatch), 31 (CHA) All possible targets are wired up, more than are really hit.
Fixed depth bounds 18-24 (CONTAINS*1..10), 30 (ARG_TO_PARAM*1..5) Fields nested more deeply / longer flow chains are not captured.
Purely name-based 1-7, 20/21, 33 Targets sharing a name but differing semantically cannot be distinguished.
Positional dataflow 27-29 No type/name matching of the arguments.
Java only / Natural only 8-15 (Java), 16-17/27/30 (Natural) No-ops for the other language.
Partial enrichment on abort all An aborted deep refresh leaves the graph inconsistent; do not stop halfway.