19 KiB
Finalize pipeline: the 36 enrichment steps
Numbering caveat (2026-09-06). This document describes a 36-step pipeline. The pipeline as it runs today has 51 steps, so the step numbers below no longer line up one-to-one with the numbers in the
Finalize <project> [n/51]log lines. The descriptions and the recorded constraints are still accurate; only the numbering is stale. Content is otherwise unchanged from the German original.
What each step of a
finalizeProjectrun does — with the emphasis on its limitations and constraints. Source of the ordering:GraphRepository.enrichmentSteps(dataflow, resolveFields, scoped). The numbering below corresponds to a full deep run over the whole root (recreate?deep=true/refresh?deep=true):dataflow = true,resolveFields = true,scoped = false. In other modes steps are skipped (see "mode gating").
Modes and what they switch on/off
The step list is not fixed — three flags control which steps run:
| Flag | When true |
Switches on |
|---|---|---|
| (always) | every run | placeholder resolution at call-graph level (1-16), Java DB, inheritance, link-calls-to-implementations, cleanup/projection (31-36) |
resolveFields |
deep ingest only | field level: 17 (indirect), 18-26 (field/bare-field resolution + the associated deletes) |
dataflow |
deep ingest only | argument->parameter dataflow: 27-29, and building on it 30 (cross-module dynamic CALLNAT) |
scoped |
per-program deep ingest | replaces project-wide variants with …_SCOPED (only modules in $names), same ordering |
Fundamental constraints of the pipeline as a whole:
- Every step runs in its own transaction and is idempotent (all
MERGE/SET, deletes are conditional). A step may safely be repeated — but a half-run deep refresh leaves the graph partially enriched: earlier steps are committed, later ones are not. Do not abort halfway. - The ordering is binding. Many steps assume that an earlier one has already resolved/built edges ( dependencies noted per step below). Reordering breaks correctness.
CONTAINSis not acyclic (item 75: self-loops on","-named parser-artefactDATA_STRUCTUREs + shared copycodeCONTROL_FLOWnodes). Every step with an unboundedCONTAINS*over module subtrees is therefore at risk of exploding. Where that was critical, it has been switched tosourceFileowner joins (16, 17, 27, 30) or bounded to a fixed depth (18-24).- Shared placeholder nodes (item 76: one node per
(name, project),sourceFile="") are referenced by hundreds of modules. That makes cardinality-heavy steps (above all 19) expensive and limits the discriminating power of resolution.
Phase A — placeholders at call-graph level (steps 1-7)
For every edge type in RESOLVABLE_EDGE_TYPES the cross-file placeholder
(sourceFile="") is re-pointed at the real target node.
1. resolve-placeholder CALLS — PERFORM/CALLNAT '…' targets onto the real module.
2. resolve-placeholder INCLUDES — USING copybooks/PDAs onto the real DATA_STRUCTURE.
3. resolve-placeholder USES_TYPE — type references.
4. resolve-placeholder EXTENDS, 5. IMPLEMENTS, 6. INJECTS, 7. REFERENCES —
Java inheritance/DI edges.
- Constraint: resolves only if a real definition with a matching name exists in the project. If the target is missing ( external module, copybook not ingested), the placeholder stays and is only marked "unresolved" in step 34 — not deleted (so that a later scoped deep ingest can still resolve it).
- Constraint: purely name-based. Targets that share a name but differ semantically are not distinguished.
- Always project-wide (no scoped variant) — including in call-graph-only mode, because it is cheap.
Phase B — Java-specific resolution (steps 8-15)
8. resolve-panache-inherited-entity — sets repositoryEntity on concrete Panache repositories whose entity
is inherited via a project base class.
- Dependency: needs resolved
EXTENDS(after step 4). Constraint: project base classes only; inheritance external to the JAR is invisible.
9. resolve-java-db-access — JPA/Panache DB_ACCESS candidates onto the entity's DB_TABLE.
- Constraint: needs persisted
MAPS_TOedges (entity->table); without a mapping there is no resolution.
10. resolve-java-query-jpql / 11. resolve-java-query-native-sql — @Query JPQL and native SQL respectively
onto
DB_TABLE.
- JPQL constraint: needs
MAPS_TO. Native SQL: matches only a literal table name — dynamically assembled SQL is not recognised.
12. build-implemented-by / 13. build-overridden-by — materialises interface->impl and
base-method->override edges respectively.
- Dependency: needs resolved
IMPLEMENTS/EXTENDS.
14. link-references-to-subclasses / 15. link-injects-to-subclasses — wires REFERENCES/INJECTS
declared in ancestors onto concrete subclasses.
- Constraint across 8-15: purely Java. For a pure Natural project these are no-ops (0 edges), but they run anyway (cheap).
Phase C — dynamic CALLNAT resolution, intra-module (steps 16-17)
16. resolve-dynamic-callnat-intra — CALLNAT <var> where <var> is assigned a string literal in the same
module becomes a real CALLS edge (tagged CALLNAT_DYNAMIC).
17. resolve-dynamic-callnat-intra-indirect — like 16, but the program name is copied indirectly from another
variable/lookup table (#PROG := #TBL(#I)).
- Constraint (over-approximation): every literal the dispatch variable can take becomes a target. Correct for genuine dispatchers, but more edges arise than a single call really hits.
- Constraint: literal assignments only. A program name read from a file/DB or indexed purely numerically
stays unresolvable -> remains visible as a
CALLNAT_DYNAMICplaceholder (step 34), it is not swallowed. - Bound/fix (item 75): module scope via
sourceFileequality instead of(caller)-[:CONTAINS*0..]->— otherwise cubic path explosion over cyclic copycode (step 17 hung for >2 h with it). ACALLNATthat sits physically in shared copycode (inupmsexactly 1 case without a module of its own) is no longer attributed to every including program. - 16 runs in every mode (cheap), 17 only with
resolveFields.
Phase D — field level: resolving placeholder fields (steps 18-26)
Only with resolveFields (deep). Re-points qualified/bare field references (READS/WRITES onto sourceFile=""
fields)
at the real field of the included data structure.
18. resolve-field-placeholder READS / 19. resolve-field-placeholder WRITES — qualified field reference (
CDPDA-M.SORT-KEY): find the real DATA_STRUCTURE of the same name INCLUDEd by the module, and within it the field
of
the same name.
- Bound: the field lookup
(real)-[:CONTAINS*1..10]->(realv)is bounded to depth 10 — groups nested more deeply (>10 levels) are not found. The field tree per structure is a pure tree (not a DAG), so it is cheap perreal. - Constraint (item 76 — fixed 2026-07-18): since the #76 fix, field placeholders have a
per-module identity (
ownerModulein the merge key), so they are no longer shared project-wide. Before that, step 19 was by far the most expensive (~55 min onupms), because the shared node produced aph × writercartesian product (~28.7 M rows); now eachphv_Mhas only the writers of its module -> step 19 in ~55 s (60x), field block 18-26 in ~2 min in total instead of ~57 min. - Constraint: resolves only on an unambiguous
INCLUDEShit; ambiguous/missing includes stay unresolved. - Ordering (item 159 — 2026-09-05): the query first resolves
(ph, phv) → real → realvand attaches the reference edges afterwards;realcomes from the index(project, type, name)instead of a name scan over the module's INCLUDES list. The field subtree is thereby expanded once per(phv, real)pair (520 onupms) instead of once per row (~220 000). TheINCLUDESstill decides which data area applies — it just no longer findsreal, it checks it. Two clauses carry this, both verified withEXPLAIN: theWITH DISTINCTas a planner barrier (without it Neo4j plans reference-driven again) and theEXISTS { (m)-[:INCLUDES]->(real) }as a predicate rather than aMATCH(as aMATCHthe planner drawsmfrom the includers ofrealand builds anm × srccross product). The_SCOPEDvariant deliberately keeps the module-driven ordering: it starts from$namesand processes few modules, so there is nothing to hoist there. - Precondition of the ordering: the placeholder name has to stay selective enough that the
index seek does not fan out wider than the INCLUDES list it replaced — on
upmsat most 4 real structures share a placeholder name (mean 1.45). The result stays correct in any case, since theINCLUDEScheck still filters; only the cost tips back.
20. resolve-field-placeholder-by-name READS / 21. …WRITES — fallback: the same resolution, but real
globally by name instead of via INCLUDES (for GLOBAL USING/copycode without an INCLUDES edge).
- Constraint: purely name-based and global — a weaker guarantee than the INCLUDES route; it can grab the wrong thing when structure names are duplicated project-wide. Runs after 18/19, so it sees only the remainder.
22. delete-resolved-field-contains — removes the now unused CONTAINS edges to resolved
field placeholders.
23. resolve-bare-included READS / 24. …WRITES — unqualified (bare) field reference onto a field of a
USING include, only if exactly one field of that name is reachable through the module's includes (item 18).
- Bound (item 77): evaluated per module (
mstays in the aggregation,srcvia(m)-[:CONTAINS*0..1]->), otherwise misattribution across module boundaries (previously:BMTABBP0wrongly wrote##MSG-NRofCDPDA-M, which it does not include).INCLUDE_FIELD_DEPTH = 10is a correctness bound here (not tuning), becauseCONTAINSis cyclic. - Constraint: ambiguity (several fields of the same name) -> stays unresolved (conservative). Known
remainder (item 76): 132 shared copycode
FUNCTIONs inupmsare one node for several modules — per-module binding cannot separate them. - Driving side (item 163, 2026-09-06): the query is driven by the placeholder name (
realvvia the(project, name)index), no longer by the include side. Arealv.sourceFile IN includedFilesprefilter keeps the subtree walk small; it is semantics-bearing, not tuning, and presupposes that a field lies in the file of its data area (checked onupms: no cross-fileCONTAINSedge out of aDATA_STRUCTUREanywhere in the project). TheWITH DISTINCTin front of it is a planner barrier — without it the filter migrates behind theSemiApplyand the step no longer completes. 135.8 s -> 90.2 s. - Driving seek (item 167, 2026-09-06):
realvis looked up per include file via(project, sourceFile, type, name)instead of by the name alone. The name seek returned 9 363 762 rows in order to keep 6 973 — around 8.4 M of them placeholders of other modules (sourceFile = ""), which a name index cannot exclude. 9 363 756 db-hits become 6 973, and the step drops from 102.0 s to 53.1 s. Caveat: the type list now determines the number of seeks (placeholders × include files × types), no longer just a filter predicate. - Resolve once instead of six times (item 169, 2026-09-06): after item 167 the entire read time lay in
the number of seeks — full read side 8 045 ms, prefix alone 7 997 ms, so
EXISTSand aggregation together 48 ms. The seeks were redundant by a factor of 6.05: 330 408(module, file, name)triples over only 54 630 distinct(file, name)pairs. The query now resolves once per pair and joins the modules back afterwards: 53.1 s -> 40.1 s. Caveat:collectmaterialises the hit list, so phase 1 has to run to completion before phase 2 begins.
25. delete-bare-placeholder-contains — removes CONTAINS to resolved bare placeholders (per-module condition, so
that the still unresolved edges of another module do not keep the placeholder alive).
26. delete-resolved-placeholders — reaps the now fully resolved placeholder structures.
- Constraint: safe only after the field placeholders are resolved — otherwise structures would be removed that a later scoped deep ingest still needs.
Phase E — dataflow argument->parameter (steps 27-29)
Only with dataflow (deep). Expensive, hence not in the fast whole-root pass.
27. link-args-to-params — links every positional CALLNAT argument of the caller with the parameter at the same
position in the callee via ARG_TO_PARAM.
- Bound/fix (item 75): with three unbounded
CONTAINS*anchors this was the slowest finalize step (~ 25 min). Switched tosourceFileowner (call site) +sourceFile/INCLUDESscope (argument variable, callee parameter).paramPositionis unique within the callee scope. - Constraint: purely positional — no type/name matching; a wrong argument order in the source would be wired up wrongly. Natural only.
- Index + caller module (item 164, 2026-09-06): the callee parameter is looked up via the index
(project, sourceFile, paramPosition). Without it the seek stops at(project, sourceFile)and reads every candidate file in full — 62 048 947 rows for 29 484 hits. A/B shows: 66.6 s without, 25.2 s with the index; not measurable at persist, because a composite index only holds nodes with all the properties (1 988 of 938 746). The caller module now comes via(cm)-[:CONTAINS*0..1]->(src)instead of viasrc.sourceFile. That also captures call sites inside copycodes, which were previously skipped silently (a.cpyhas noMODULEnode of its own):ARG_TO_PARAMonupms17 813 -> 17 871. Asize(cms) = 1guard keeps the step conservative in case a copycodeFUNCTIONis ever contained by several modules (item 76).
28. link-args-to-params-java / 29. link-args-to-params-java-cross — the same for Java, intra- and
cross-class respectively.
Phase F — cross-module dynamic CALLNAT resolution (step 30)
30. resolve-dynamic-callnat-cross — CALLNAT <var> whose program name flows in from another module via a
parameter: follows ARG_TO_PARAM (from 27) into the callee and takes the literals written into that parameter there
as targets.
- Dependency: needs the
ARG_TO_PARAMedges from 27 ->dataflowonly. - Bound:
ARG_TO_PARAM*1..5— flow over at most 5 hops; longer chains are unresolvable. - Bound/fix (item 75): call site via
sourceFile, dispatch variable viasourceFile/INCLUDESscope (the name alone would catch ~20 958 foreign variables of the same name). Over-approximation as with 16/17.
Phase G — polymorphism, cleanup, projection (steps 31-36)
31. link-calls-to-implementations — fans class-level CALLS on an interface/base type out to the
implementations/subclasses (CHA).
- Constraint: class hierarchy analysis — over-approximates (all implementations count as possible
targets), no precise target determination. Runs last, so that it sees resolved
CALLS+IMPLEMENTS/EXTENDS.
32. delete-dynamic-callnat-placeholders — removes the CALLNAT_DYNAMIC marker edges (the unresolved
dispatch placeholders); resolved edges remain.
33. delete-data-literal-call-placeholders — reaps placeholder MODULEs that are really data literals (browse
keys)
and were wrongly restored as call targets (item 62).
- Constraint: a name-based test that only removes edges which are wrong everywhere.
34. stamp-unresolved-placeholders — marks leftover placeholders as (un)resolved, depending on whether a real
definition now exists. Runs last, sees the fully re-pointed/reaped graph.
35. delete-calls-module / 36. build-calls-module — projects the finished call graph onto CALLS_MODULE
module->module edges (so that traversals can be bounded in module hops rather than raw CALLS hops).
- Constraint: a pure projection — only as correct as the graph at the moment of the run. Must run after all
steps that add
CALLS(placeholder/dynamic resolution, 31) or remove it (33). The precedingdelete(35) makes it correct across refreshes (aMERGEalone would never remove stale edges).
Summary of the most important constraints
| Topic | Steps affected | Constraint |
|---|---|---|
| Mode gating | 17, 18-26 (resolveFields); 27-30 (dataflow) |
Without deep ingest, field and dataflow edges are missing; queries on them return empty/guarded. |
CONTAINS cycles (item 75) |
16, 17, 27, 30 (fixed); query endpoints (open) | Unbounded CONTAINS* explodes over cyclic copycode; worked around via sourceFile owner. |
| Shared placeholders (item 76) | 19 (cardinality), 23/24 (discrimination) | One node per (name, project) -> step 19 very expensive; 132 shared copycode FUNCTIONs not separable per module. |
| Over-approximation | 16, 17, 30 (dynamic dispatch), 31 (CHA) | All possible targets are wired up, more than are really hit. |
| Fixed depth bounds | 18-24 (CONTAINS*1..10), 30 (ARG_TO_PARAM*1..5) |
Fields nested more deeply / longer flow chains are not captured. |
| Purely name-based | 1-7, 20/21, 33 | Targets sharing a name but differing semantically cannot be distinguished. |
| Positional dataflow | 27-29 | No type/name matching of the arguments. |
| Java only / Natural only | 8-15 (Java), 16-17/27/30 (Natural) | No-ops for the other language. |
| Partial enrichment on abort | all | An aborted deep refresh leaves the graph inconsistent; do not stop halfway. |