Files
agenticCode/x-docs/neo4j-introduction.md
Ingo Schnabel da9308e98a Add project scoping, project CRUD, CLI config file, and Neo4j docs
Scope all AST nodes/queries by project, add /api/projects CRUD endpoints
and matching ac-cli commands (use, project create/update/delete/list,
-p/--project), persist server URL and selected project to
~/.agenticcode/config.properties, and add a Neo4j/Cypher introduction
for the team. Also clarify CLAUDE.md's commit policy.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-14 09:20:59 +02:00

22 KiB

Introduction to Neo4j (for AgenticCode)

This document is a practical introduction to Neo4j aimed at developers working on AgenticCode who have no prior graph-database experience. It covers how Neo4j works, how this project models its data, how the existing CypherQueries work, and how graph queries map onto program-analysis questions.


1. How Neo4j works in general

Neo4j is a graph database. Instead of tables and rows (relational) or documents (NoSQL document stores), everything is stored as:

  • Nodes — entities, roughly like "rows" but with no fixed schema. Each node has:
    • zero or more labels (a node's "type", e.g. :AstNode, :Project) — a node can have multiple labels
    • a set of properties (key/value pairs, e.g. name: "YADDRBN0", startLine: 42)
  • Relationships — directed, typed edges between two nodes. Each relationship has:
    • exactly one type (e.g. :CALLS, :CONTAINS, :READS)
    • a direction ((a)-[:CALLS]->(b))
    • optionally its own properties

Everything is stored "pre-joined": a relationship is a physical pointer between two nodes, stored once at creation time. This means traversing a relationship is a cheap pointer lookup, not a join computed at query time. The more your queries are about connections (who calls whom, what depends on what, how does data flow), the better a graph database performs compared to a relational database with many JOINs.

Key concepts

Concept Relational analogy Example in AgenticCode
Node Row One AstNode (a function, variable, module, ...)
Label Table name AstNode, Project
Property Column value name, type, sourceFile, startLine
Relationship Foreign key / join (:AstNode)-[:CALLS]->(:AstNode)
Relationship type — (no real equivalent) CALLS, CONTAINS, READS, WRITES

Cypher — the query language

Cypher is Neo4j's query language. It's declarative and visual: patterns in a query look like the graph itself, using ASCII-art for nodes () and relationships -->.

MATCH (caller:AstNode)-[:CALLS]->(callee:AstNode {name: 'FOO'})
RETURN caller.name

This reads almost like English: "find all AstNodes that have a CALLS relationship pointing to an AstNode named FOO, and return the caller's name."

Schema-optional, but indexed in practice

Neo4j doesn't require you to declare a schema up front — any node can have any properties. In practice, for performance, you create indexes/constraints on the properties you frequently search by (e.g. an index on AstNode.name and AstNode.project), so lookups like MATCH (n:AstNode {name: $name}) don't scan every node.

How AgenticCode talks to Neo4j

  • ac-neo4j-store uses the official neo4j-java-driver.
  • GraphRepository opens a Session, runs Cypher via session.executeWrite(...) / session.executeRead(...), and maps result Records to Java records (CallReference, DbAccess, IdentifierMatch, ProjectInfo, ...).
  • All Cypher query strings live in CypherQueries (per the project's "no inline Cypher in services" convention) — GraphRepository only supplies parameters.

1a. Writing Cypher queries — syntax basics

Cypher reads like ASCII-art of the graph pattern you're looking for, followed by what to do with what you found. Most queries are built from a small set of clauses, used in roughly this order:

MATCH ...        -- find a pattern in the graph
WHERE ...        -- filter the matches
WITH ...         -- reshape/aggregate before continuing
RETURN ...       -- produce results
ORDER BY ...
SKIP / LIMIT ...

for writes:

CREATE ...       -- always create new node(s)/relationship(s)
MERGE ...        -- find-or-create (upsert)
SET ...          -- add/overwrite properties
DELETE / DETACH DELETE ...

Patterns: nodes and relationships

(n)                              -- any node, bound to variable n
(n:AstNode)                      -- node with label AstNode
(n:AstNode {name: 'FOO'})        -- label + property filter (inline)
(a)-[:CALLS]->(b)                -- directed relationship of type CALLS
(a)-[r:CALLS]->(b)               -- relationship bound to variable r
(a)-[:CALLS|CONTAINS]->(b)       -- either relationship type
(a)-[:CALLS*1..3]->(b)           -- variable-length path, 1 to 3 hops
(a)-[:CONTAINS*0..]->(b)         -- 0 or more hops (a itself, or any descendant)
(a)--(b)                         -- relationship, direction/type don't matter
  • () = node, [] = relationship, -->/<--/-- = direction.
  • Anything not given a variable name ((), [:CALLS]) is just a pattern shape — you can't reference it later, but it still constrains the match.
  • Property filters in {...} are an AND of equality checks; for anything more complex (ranges, IN, regex, IS NULL, ...), match without the filter and use WHERE instead.

MATCH + WHERE — finding things

MATCH (m:AstNode {type: 'MODULE'})
WHERE m.project = $project AND m.name STARTS WITH 'YADDR'
RETURN m.name, m.sourceFile
  • {type: 'MODULE'} and WHERE m.project = $project are equivalent ways to filter — inline filters are slightly more index-friendly, WHERE is more flexible (operators: =, <>, <, >, IN, STARTS WITH/CONTAINS/ENDS WITH, IS NULL, AND/OR/NOT, regex with =~).
  • Multiple MATCH clauses (or comma-separated patterns) are joined like an inner join on shared variables — see CALLERS in section 3.3, which uses two MATCH clauses connected via m/target.

RETURN — shaping output

RETURN m.name AS name, m.type AS type     -- alias columns (becomes the JSON key)
RETURN DISTINCT caller.name               -- de-duplicate rows
RETURN count(*) AS total                  -- aggregation

Aggregations (count, collect, sum, avg, min, max) implicitly group by every other non-aggregated expression in the RETURN — there's no separate GROUP BY.

Parameters — never inline user input

MATCH (n:AstNode {project: $project, name: $name}) RETURN n

$name-style parameters are supplied separately as a map (tx.run(query, Map.of("project", project, "name", name)) in GraphRepository). Always use parameters instead of string-concatenating values into the query — it avoids Cypher injection (the graph equivalent of SQL injection) and lets Neo4j cache/reuse the query plan.

CREATE vs MERGE — the most important distinction for writes

CREATE (p:Project {name: $name})          -- always inserts a new node, even if one
                                           -- with the same properties exists

MERGE (n:AstNode {type: $type, name: $name, sourceFile: $sourceFile, project: $project})
ON CREATE SET n.id = $id                  -- only runs if a new node was created
ON MATCH SET n.lastSeen = timestamp()     -- only runs if an existing node matched
SET n.language = $language                -- runs either way
  • MERGE treats everything inside {...} as the identity to match-or-create — put only the stable "key" properties there (as MERGE_NODE does), and use a separate SET for everything else, otherwise a single differing property (e.g. a changed line number) would cause MERGE to create a duplicate node instead of updating the existing one.
  • ON CREATE SET / ON MATCH SET let you distinguish "first time" vs "update" — not currently used in this project, but useful for e.g. tracking createdAt vs updatedAt.

SET / REMOVE / DELETE

SET n.dataType = $dataType        -- set/overwrite a property
SET n += $propsMap                -- merge a map of properties into a node
REMOVE n.dataType                 -- remove a property entirely
DELETE n                          -- delete a node (fails if it has relationships)
DETACH DELETE n                   -- delete a node and all its relationships

WITH — chaining query stages

WITH passes variables (optionally aggregated/filtered/reshaped) from one part of a query to the next — it's how you build multi-step queries:

MATCH (m:AstNode {type: 'MODULE', project: $project})-[:CONTAINS]->(f:AstNode {type: 'FUNCTION'})
WITH m, count(f) AS functionCount
WHERE functionCount > 10
RETURN m.name, functionCount
ORDER BY functionCount DESC

A minimal mental checklist when writing a new query

  1. Sketch the pattern: which node(s) am I anchoring on, and what path connects them to what I want to return? (Draw it as ()-->() on paper first.)
  2. Anchor on the most selective filter first (usually project + name + type, ideally backed by an index — see section 4.4).
  3. Use MATCH/WHERE for reads, MERGE (with a minimal identity key) for idempotent writes, CREATE only when duplicates are impossible/acceptable.
  4. Always pass values as $parameters, never string-concatenate.
  5. Try it in Neo4j Browser with literal values before wiring it into CypherQueries/GraphRepository.

For anything beyond this, the official Cypher manual and the interactive Cypher cheat sheet are the best references.


2. Structure of the database

2.1 Node labels and properties

AgenticCode currently uses two node labels:

  • :AstNode — every parsed element (module, function, variable, data structure, DB table, constant, field). The type property (a NodeType enum value) distinguishes what kind of AST element it is — it is not a separate label.
  • :Project — a logical grouping/namespace for ingested source code.
classDiagram
    class AstNode {
        +String id
        +String type
        +String name
        +String sourceFile
        +String project
        +String language
        +int startLine
        +int endLine
        +String dataType
        +String value
    }
    class Project {
        +String name
        +String description
    }

AstNode.type is one of (from NodeType):

MODULE | FUNCTION | VARIABLE | DATA_STRUCTURE | DB_TABLE | CONSTANT | FIELD

2.2 Relationship types

From EdgeType:

CONTAINS | CALLS | READS | WRITES | USES_TYPE | EXTENDS | IMPLEMENTS | INCLUDES

2.3 Conceptual schema (entity/relationship view)

erDiagram
    MODULE ||--o{ FUNCTION : CONTAINS
    FUNCTION ||--o{ FUNCTION : CALLS
    FUNCTION ||--o{ VARIABLE : READS
    FUNCTION ||--o{ VARIABLE : WRITES
    FUNCTION ||--o{ DB_TABLE : READS
    FUNCTION ||--o{ DB_TABLE : WRITES
    FUNCTION ||--o{ DATA_STRUCTURE : USES_TYPE
    MODULE ||--o{ MODULE : EXTENDS
    MODULE ||--o{ MODULE : IMPLEMENTS
    MODULE ||--o{ MODULE : INCLUDES
    PROJECT ||--o{ MODULE : "scopes (via project property)"

Note: MODULE, FUNCTION, VARIABLE, DATA_STRUCTURE, DB_TABLE are not separate labels in the real database — they are all :AstNode nodes whose type property has that value. The diagram shows the conceptual model; section 2.1 shows the physical model.

2.4 Example graph instance

A small Java class SampleRequest extends BaseRequest with a method getPartnerId() that reads field partnerId, ingested into project demo, looks like this:

graph TD
    M["AstNode<br/>type=MODULE<br/>name=SampleRequest<br/>project=demo"]
    BASE["AstNode<br/>type=MODULE<br/>name=BaseRequest<br/>project=demo"]
    F["AstNode<br/>type=FUNCTION<br/>name=getPartnerId<br/>project=demo"]
    V["AstNode<br/>type=FIELD<br/>name=partnerId<br/>project=demo"]

    M -- CONTAINS --> F
    M -- EXTENDS --> BASE
    F -- READS --> V

2.5 Project isolation

Every :AstNode carries a project property, and is merged (deduplicated) on (type, name, sourceFile, project). This means:

  • The same module name can exist independently in two different projects.
  • A :Project node is separate metadata (name, description) used only by the project CRUD API — it is not itself connected to :AstNodes via relationships; the link is the shared project property value.
  • Deleting a project (DELETE_PROJECT) removes the :Project node and every :AstNode (and its relationships, via DETACH DELETE) with that project value.
graph LR
    subgraph "project = orderapp"
        A1[AstNode MODULE Order]
        A2[AstNode FUNCTION calculateTotal]
        A1 -- CONTAINS --> A2
    end
    subgraph "project = billingapp"
        B1[AstNode MODULE Order]
        B2[AstNode FUNCTION calculateTotal]
        B1 -- CONTAINS --> B2
    end
    P1[Project name=orderapp] -. "same project value, no edge" .-> A1
    P2[Project name=billingapp] -. "same project value, no edge" .-> B1

Two modules named Order in different projects are distinct nodes — there is no relationship between them and no way to accidentally cross-reference data between projects.


3. Explaining CypherQueries

All queries live in ac-neo4j-store/src/main/java/com/agenticcode/neo4jstore/graph/CypherQueries.java. Parameters ($name, $project, ...) are supplied by GraphRepository as a Map<String, Object>.

3.1 MERGE_NODE — upsert a parsed AST element

MERGE (n:AstNode {type: $type, name: $name, sourceFile: $sourceFile, project: $project})
SET n.id = $id, n.language = $language, n.startLine = $startLine,
    n.endLine = $endLine, n.dataType = $dataType, n.value = $value
  • MERGE = "find a node matching this pattern, or create it if it doesn't exist." The properties inside {...} form the identity key. Two ingests of the same (type, name, sourceFile, project) combination update the same node rather than creating duplicates — this makes re-ingesting a changed file idempotent.
  • SET then (re-)writes the remaining properties, including a fresh id (UUID) and line numbers, so re-parsing a file refreshes its data.
  • Placeholder nodes for external references (e.g. a CALLNAT target that hasn't been ingested yet) use sourceFile = "". Once the real file is ingested with the same (type, name, project), the MERGE matches the same placeholder node and enriches it — this is how forward references resolve.

3.2 Edge merges — mergeEdge(EdgeType)

MATCH (a:AstNode {id: $sourceId}), (b:AstNode {id: $targetId})
MERGE (a)-[:CALLS]->(b)

(generated once per EdgeType value, e.g. CALLS, CONTAINS, READS, ...)

  • Nodes are matched by their stable id (UUID) — not by name — so edges always point at the exact node instances created/merged by MERGE_NODE.
  • MERGE on the relationship avoids duplicate edges if the same call/read/write is ingested twice.
  • Because id is globally unique, edges can never accidentally cross between projects, even without an explicit project check.

3.3 CALLERS — who calls this module?

MATCH (m:AstNode {type: 'MODULE', name: $name, project: $project})
MATCH (caller:AstNode)-[:CALLS]->(target:AstNode)
WHERE target = m OR (m)-[:CONTAINS]->(target)
RETURN DISTINCT caller.name AS name, caller.type AS type, caller.sourceFile AS sourceFile
  • First finds the anchor MODULE node m (scoped by project).
  • Then finds any CALLS edge whose target is either m itself, or something m CONTAINS (i.e. a function defined inside that module) — so calling a function inside SampleRequest counts as "calling SampleRequest".
  • target = m is a node-identity comparison (same internal node).

3.4 CALLEES — what does this module call?

MATCH (m:AstNode {type: 'MODULE', name: $name, project: $project})
MATCH (m)-[:CONTAINS*0..]->(source:AstNode)-[:CALLS|EXTENDS|IMPLEMENTS]->(callee:AstNode)
RETURN DISTINCT callee.name AS name, callee.type AS type, callee.sourceFile AS sourceFile
  • [:CONTAINS*0..] is a variable-length path: "zero or more CONTAINS hops." *0 means source can be m itself; more hops would reach functions-within- functions etc.
  • [:CALLS|EXTENDS|IMPLEMENTS] is relationship-type alternation — match any one of these three edge types in a single step.
  • Result: every module/function/class that m (or anything it contains) calls, extends, or implements.

3.5 DB_ACCESSES — which DB tables, and how?

MATCH (m:AstNode {type: 'MODULE', name: $name, project: $project})
MATCH (m)-[:CONTAINS*0..]->(f:AstNode)-[r:READS|WRITES]->(t:AstNode {type: 'DB_TABLE'})
RETURN DISTINCT t.name AS name, type(r) AS mode
  • Same "module or anything it contains" pattern as CALLEES.
  • r is bound to the relationship itself (not just a node), so type(r) returns the relationship's type as a string ("READS" or "WRITES") — this becomes the mode in the response.
  • Filters the target node by type: 'DB_TABLE'.

3.6 SEARCH_IDENTIFIER — find a name anywhere in the project

MATCH (n:AstNode {project: $project})
WHERE n.name = $name
RETURN n.type AS type, n.name AS name, n.sourceFile AS sourceFile,
       n.startLine AS startLine, n.endLine AS endLine,
       n.dataType AS dataType, n.value AS value
  • A flat scan over all :AstNodes in the project filtered by exact name. With an index on (project, name) this is fast even for large graphs.
  • Returns every kind of node (function, variable, constant, ...) that has this name — useful for "where is partnerId used/defined?"

3.7 Project CRUD queries

-- PROJECT_EXISTS
MATCH (p:Project {name: $name}) RETURN p

-- CREATE_PROJECT
CREATE (p:Project {name: $name, description: $description})

-- UPDATE_PROJECT
MATCH (p:Project {name: $name}) SET p.description = $description

-- DELETE_PROJECT  (cascades to all AstNodes in the project)
MATCH (p:Project {name: $name})
OPTIONAL MATCH (n:AstNode {project: $name})
DETACH DELETE p, n

-- LIST_PROJECTS
MATCH (p:Project) RETURN p.name AS name, p.description AS description ORDER BY p.name
  • CREATE (unlike MERGE) always creates a new node — GraphRepository checks PROJECT_EXISTS first to return 409 Conflict instead of creating duplicates.
  • DETACH DELETE removes a node and all its relationships in one step — required because Neo4j refuses to delete a node that still has relationships attached. OPTIONAL MATCH ensures the query doesn't fail if the project has no AstNodes yet (an empty/never-ingested project).

4. Using Neo4j for program analysis (and beyond)

4.1 Why a graph model fits program analysis

Most interesting questions about source code are graph questions:

Question Graph operation
"Who calls this function?" incoming CALLS edges (1 hop)
"What's the full call tree under this module?" variable-length traversal (*0..N or *0..)
"Is there a cycle in the call graph?" path-finding / cycle detection
"What tables does this module touch, transitively?" traversal + filter by node type
"What's the shortest call path from A to B?" shortestPath()
"Which modules are most central / most depended-upon?" graph algorithms (degree, PageRank via GDS)
"Show me everything that would break if I delete this function" reverse traversal of CALLS/USES_TYPE/READS/WRITES

A relational database can answer these with recursive CTEs, but each additional hop means another self-join; performance degrades quickly with depth. In Neo4j, traversal cost depends on the size of the result, not the size of the whole table — a 5-hop traversal over a 10-million-node graph is fast if only a handful of nodes match.

4.2 Patterns already used in AgenticCode

  • Call graph navigation (CALLERS/CALLEES): direct application of the table above — "who calls / is called by this module."
  • Data/DB lineage (DB_ACCESSES): traversal from a module down to the DB tables it reads/writes, classified by access mode.
  • Cross-module search (SEARCH_IDENTIFIER): flat lookup by property, independent of graph structure — Neo4j is also a perfectly good "property store" for this.
  • Multi-tenancy (project property): a lightweight way to partition the graph without separate databases, while still allowing (in principle) cross-project queries later if ever needed (e.g. "does any project call this shared library?").

4.3 Ideas for future enrichment / queries

These map directly to the EnrichmentPipelines design and the agent-facing API:

  • Impact analysis: MATCH (n {name:$name})<-[:CALLS|READS|WRITES|USES_TYPE*1..3]-(dep) RETURN dep — "what depends on this, up to 3 hops away?"
  • Dead code detection: MATCH (f:AstNode {type:'FUNCTION', project:$project}) WHERE NOT ()-[:CALLS]->(f) RETURN f — functions with no incoming CALLS edge.
  • Full call chains to a DB table: MATCH path = (m:AstNode {type:'MODULE'})-[:CONTAINS*0..]->()-[:READS|WRITES]->(t:AstNode {type:'DB_TABLE', name:$table}) RETURN path — every module that ultimately touches a given table, with the full path for explanation.
  • Cycle detection in the call graph: MATCH (f:AstNode)-[:CALLS*1..10]->(f) (with a hop limit to bound cost) — useful for spotting recursive or circular dependencies in legacy Natural code.

4.4 Practical tips while developing

  • Neo4j Browser (http://localhost:7474) is the fastest way to explore: run any CypherQueries string directly with literal values to see the actual graph and tune the query before wiring it into GraphRepository.
  • EXPLAIN/PROFILE a query (PROFILE MATCH ...) to see whether it's using an index or doing a full label scan — important once a project has many thousands of AstNodes.
  • Add indexes for properties used in WHERE/MATCH patterns, e.g.:
    CREATE INDEX ast_node_project_name IF NOT EXISTS
    FOR (n:AstNode) ON (n.project, n.name);
    
    CREATE INDEX ast_node_project_type_name IF NOT EXISTS
    FOR (n:AstNode) ON (n.project, n.type, n.name);
    
    (Not yet present in the codebase — worth adding once ingest volumes grow; see the open TODO "Deploy unified AST schema to Neo4j (constraints + indexes)".)
  • Resetting data during development: MATCH (n) DETACH DELETE n wipes the entire database (all projects) — use the project-scoped DELETE_PROJECT query instead when you only want to clear one project.