Scope all AST nodes/queries by project, add /api/projects CRUD endpoints and matching ac-cli commands (use, project create/update/delete/list, -p/--project), persist server URL and selected project to ~/.agenticcode/config.properties, and add a Neo4j/Cypher introduction for the team. Also clarify CLAUDE.md's commit policy. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
22 KiB
Introduction to Neo4j (for AgenticCode)
This document is a practical introduction to Neo4j aimed at developers working on
AgenticCode who have no prior graph-database experience. It covers how Neo4j works,
how this project models its data, how the existing CypherQueries work, and how
graph queries map onto program-analysis questions.
1. How Neo4j works in general
Neo4j is a graph database. Instead of tables and rows (relational) or documents (NoSQL document stores), everything is stored as:
- Nodes — entities, roughly like "rows" but with no fixed schema. Each node has:
- zero or more labels (a node's "type", e.g.
:AstNode,:Project) — a node can have multiple labels - a set of properties (key/value pairs, e.g.
name: "YADDRBN0",startLine: 42)
- zero or more labels (a node's "type", e.g.
- Relationships — directed, typed edges between two nodes. Each relationship has:
- exactly one type (e.g.
:CALLS,:CONTAINS,:READS) - a direction (
(a)-[:CALLS]->(b)) - optionally its own properties
- exactly one type (e.g.
Everything is stored "pre-joined": a relationship is a physical pointer between two nodes, stored once at creation time. This means traversing a relationship is a cheap pointer lookup, not a join computed at query time. The more your queries are about connections (who calls whom, what depends on what, how does data flow), the better a graph database performs compared to a relational database with many JOINs.
Key concepts
| Concept | Relational analogy | Example in AgenticCode |
|---|---|---|
| Node | Row | One AstNode (a function, variable, module, ...) |
| Label | Table name | AstNode, Project |
| Property | Column value | name, type, sourceFile, startLine |
| Relationship | Foreign key / join | (:AstNode)-[:CALLS]->(:AstNode) |
| Relationship type | — (no real equivalent) | CALLS, CONTAINS, READS, WRITES |
Cypher — the query language
Cypher is Neo4j's query language. It's declarative and visual: patterns in a query
look like the graph itself, using ASCII-art for nodes () and relationships -->.
MATCH (caller:AstNode)-[:CALLS]->(callee:AstNode {name: 'FOO'})
RETURN caller.name
This reads almost like English: "find all AstNodes that have a CALLS
relationship pointing to an AstNode named FOO, and return the caller's name."
Schema-optional, but indexed in practice
Neo4j doesn't require you to declare a schema up front — any node can have any
properties. In practice, for performance, you create indexes/constraints on the
properties you frequently search by (e.g. an index on AstNode.name and
AstNode.project), so lookups like MATCH (n:AstNode {name: $name}) don't scan
every node.
How AgenticCode talks to Neo4j
ac-neo4j-storeuses the officialneo4j-java-driver.GraphRepositoryopens aSession, runs Cypher viasession.executeWrite(...)/session.executeRead(...), and maps resultRecords to Java records (CallReference,DbAccess,IdentifierMatch,ProjectInfo, ...).- All Cypher query strings live in
CypherQueries(per the project's "no inline Cypher in services" convention) —GraphRepositoryonly supplies parameters.
1a. Writing Cypher queries — syntax basics
Cypher reads like ASCII-art of the graph pattern you're looking for, followed by what to do with what you found. Most queries are built from a small set of clauses, used in roughly this order:
MATCH ... -- find a pattern in the graph
WHERE ... -- filter the matches
WITH ... -- reshape/aggregate before continuing
RETURN ... -- produce results
ORDER BY ...
SKIP / LIMIT ...
for writes:
CREATE ... -- always create new node(s)/relationship(s)
MERGE ... -- find-or-create (upsert)
SET ... -- add/overwrite properties
DELETE / DETACH DELETE ...
Patterns: nodes and relationships
(n) -- any node, bound to variable n
(n:AstNode) -- node with label AstNode
(n:AstNode {name: 'FOO'}) -- label + property filter (inline)
(a)-[:CALLS]->(b) -- directed relationship of type CALLS
(a)-[r:CALLS]->(b) -- relationship bound to variable r
(a)-[:CALLS|CONTAINS]->(b) -- either relationship type
(a)-[:CALLS*1..3]->(b) -- variable-length path, 1 to 3 hops
(a)-[:CONTAINS*0..]->(b) -- 0 or more hops (a itself, or any descendant)
(a)--(b) -- relationship, direction/type don't matter
()= node,[]= relationship,-->/<--/--= direction.- Anything not given a variable name (
(),[:CALLS]) is just a pattern shape — you can't reference it later, but it still constrains the match. - Property filters in
{...}are an AND of equality checks; for anything more complex (ranges,IN, regex,IS NULL, ...), match without the filter and useWHEREinstead.
MATCH + WHERE — finding things
MATCH (m:AstNode {type: 'MODULE'})
WHERE m.project = $project AND m.name STARTS WITH 'YADDR'
RETURN m.name, m.sourceFile
{type: 'MODULE'}andWHERE m.project = $projectare equivalent ways to filter — inline filters are slightly more index-friendly,WHEREis more flexible (operators:=,<>,<,>,IN,STARTS WITH/CONTAINS/ENDS WITH,IS NULL,AND/OR/NOT, regex with=~).- Multiple
MATCHclauses (or comma-separated patterns) are joined like an inner join on shared variables — seeCALLERSin section 3.3, which uses twoMATCHclauses connected viam/target.
RETURN — shaping output
RETURN m.name AS name, m.type AS type -- alias columns (becomes the JSON key)
RETURN DISTINCT caller.name -- de-duplicate rows
RETURN count(*) AS total -- aggregation
Aggregations (count, collect, sum, avg, min, max) implicitly group by
every other non-aggregated expression in the RETURN — there's no separate
GROUP BY.
Parameters — never inline user input
MATCH (n:AstNode {project: $project, name: $name}) RETURN n
$name-style parameters are supplied separately as a map
(tx.run(query, Map.of("project", project, "name", name)) in GraphRepository).
Always use parameters instead of string-concatenating values into the query —
it avoids Cypher injection (the graph equivalent of SQL injection) and lets Neo4j
cache/reuse the query plan.
CREATE vs MERGE — the most important distinction for writes
CREATE (p:Project {name: $name}) -- always inserts a new node, even if one
-- with the same properties exists
MERGE (n:AstNode {type: $type, name: $name, sourceFile: $sourceFile, project: $project})
ON CREATE SET n.id = $id -- only runs if a new node was created
ON MATCH SET n.lastSeen = timestamp() -- only runs if an existing node matched
SET n.language = $language -- runs either way
MERGEtreats everything inside{...}as the identity to match-or-create — put only the stable "key" properties there (asMERGE_NODEdoes), and use a separateSETfor everything else, otherwise a single differing property (e.g. a changed line number) would causeMERGEto create a duplicate node instead of updating the existing one.ON CREATE SET/ON MATCH SETlet you distinguish "first time" vs "update" — not currently used in this project, but useful for e.g. trackingcreatedAtvsupdatedAt.
SET / REMOVE / DELETE
SET n.dataType = $dataType -- set/overwrite a property
SET n += $propsMap -- merge a map of properties into a node
REMOVE n.dataType -- remove a property entirely
DELETE n -- delete a node (fails if it has relationships)
DETACH DELETE n -- delete a node and all its relationships
WITH — chaining query stages
WITH passes variables (optionally aggregated/filtered/reshaped) from one part of a
query to the next — it's how you build multi-step queries:
MATCH (m:AstNode {type: 'MODULE', project: $project})-[:CONTAINS]->(f:AstNode {type: 'FUNCTION'})
WITH m, count(f) AS functionCount
WHERE functionCount > 10
RETURN m.name, functionCount
ORDER BY functionCount DESC
A minimal mental checklist when writing a new query
- Sketch the pattern: which node(s) am I anchoring on, and what path connects them
to what I want to return? (Draw it as
()-->()on paper first.) - Anchor on the most selective filter first (usually
project+name+type, ideally backed by an index — see section 4.4). - Use
MATCH/WHEREfor reads,MERGE(with a minimal identity key) for idempotent writes,CREATEonly when duplicates are impossible/acceptable. - Always pass values as
$parameters, never string-concatenate. - Try it in Neo4j Browser with literal values before wiring it into
CypherQueries/GraphRepository.
For anything beyond this, the official Cypher manual and the interactive Cypher cheat sheet are the best references.
2. Structure of the database
2.1 Node labels and properties
AgenticCode currently uses two node labels:
:AstNode— every parsed element (module, function, variable, data structure, DB table, constant, field). Thetypeproperty (aNodeTypeenum value) distinguishes what kind of AST element it is — it is not a separate label.:Project— a logical grouping/namespace for ingested source code.
classDiagram
class AstNode {
+String id
+String type
+String name
+String sourceFile
+String project
+String language
+int startLine
+int endLine
+String dataType
+String value
}
class Project {
+String name
+String description
}
AstNode.type is one of (from NodeType):
MODULE | FUNCTION | VARIABLE | DATA_STRUCTURE | DB_TABLE | CONSTANT | FIELD
2.2 Relationship types
From EdgeType:
CONTAINS | CALLS | READS | WRITES | USES_TYPE | EXTENDS | IMPLEMENTS | INCLUDES
2.3 Conceptual schema (entity/relationship view)
erDiagram
MODULE ||--o{ FUNCTION : CONTAINS
FUNCTION ||--o{ FUNCTION : CALLS
FUNCTION ||--o{ VARIABLE : READS
FUNCTION ||--o{ VARIABLE : WRITES
FUNCTION ||--o{ DB_TABLE : READS
FUNCTION ||--o{ DB_TABLE : WRITES
FUNCTION ||--o{ DATA_STRUCTURE : USES_TYPE
MODULE ||--o{ MODULE : EXTENDS
MODULE ||--o{ MODULE : IMPLEMENTS
MODULE ||--o{ MODULE : INCLUDES
PROJECT ||--o{ MODULE : "scopes (via project property)"
Note:
MODULE,FUNCTION,VARIABLE,DATA_STRUCTURE,DB_TABLEare not separate labels in the real database — they are all:AstNodenodes whosetypeproperty has that value. The diagram shows the conceptual model; section 2.1 shows the physical model.
2.4 Example graph instance
A small Java class SampleRequest extends BaseRequest with a method
getPartnerId() that reads field partnerId, ingested into project demo, looks
like this:
graph TD
M["AstNode<br/>type=MODULE<br/>name=SampleRequest<br/>project=demo"]
BASE["AstNode<br/>type=MODULE<br/>name=BaseRequest<br/>project=demo"]
F["AstNode<br/>type=FUNCTION<br/>name=getPartnerId<br/>project=demo"]
V["AstNode<br/>type=FIELD<br/>name=partnerId<br/>project=demo"]
M -- CONTAINS --> F
M -- EXTENDS --> BASE
F -- READS --> V
2.5 Project isolation
Every :AstNode carries a project property, and is merged (deduplicated) on
(type, name, sourceFile, project). This means:
- The same module name can exist independently in two different projects.
- A
:Projectnode is separate metadata (name,description) used only by the project CRUD API — it is not itself connected to:AstNodes via relationships; the link is the sharedprojectproperty value. - Deleting a project (
DELETE_PROJECT) removes the:Projectnode and every:AstNode(and its relationships, viaDETACH DELETE) with thatprojectvalue.
graph LR
subgraph "project = orderapp"
A1[AstNode MODULE Order]
A2[AstNode FUNCTION calculateTotal]
A1 -- CONTAINS --> A2
end
subgraph "project = billingapp"
B1[AstNode MODULE Order]
B2[AstNode FUNCTION calculateTotal]
B1 -- CONTAINS --> B2
end
P1[Project name=orderapp] -. "same project value, no edge" .-> A1
P2[Project name=billingapp] -. "same project value, no edge" .-> B1
Two modules named Order in different projects are distinct nodes — there is no
relationship between them and no way to accidentally cross-reference data between
projects.
3. Explaining CypherQueries
All queries live in
ac-neo4j-store/src/main/java/com/agenticcode/neo4jstore/graph/CypherQueries.java.
Parameters ($name, $project, ...) are supplied by GraphRepository as a
Map<String, Object>.
3.1 MERGE_NODE — upsert a parsed AST element
MERGE (n:AstNode {type: $type, name: $name, sourceFile: $sourceFile, project: $project})
SET n.id = $id, n.language = $language, n.startLine = $startLine,
n.endLine = $endLine, n.dataType = $dataType, n.value = $value
MERGE= "find a node matching this pattern, or create it if it doesn't exist." The properties inside{...}form the identity key. Two ingests of the same(type, name, sourceFile, project)combination update the same node rather than creating duplicates — this makes re-ingesting a changed file idempotent.SETthen (re-)writes the remaining properties, including a freshid(UUID) and line numbers, so re-parsing a file refreshes its data.- Placeholder nodes for external references (e.g. a
CALLNATtarget that hasn't been ingested yet) usesourceFile = "". Once the real file is ingested with the same(type, name, project), theMERGEmatches the same placeholder node and enriches it — this is how forward references resolve.
3.2 Edge merges — mergeEdge(EdgeType)
MATCH (a:AstNode {id: $sourceId}), (b:AstNode {id: $targetId})
MERGE (a)-[:CALLS]->(b)
(generated once per EdgeType value, e.g. CALLS, CONTAINS, READS, ...)
- Nodes are matched by their stable
id(UUID) — not by name — so edges always point at the exact node instances created/merged byMERGE_NODE. MERGEon the relationship avoids duplicate edges if the same call/read/write is ingested twice.- Because
idis globally unique, edges can never accidentally cross between projects, even without an explicitprojectcheck.
3.3 CALLERS — who calls this module?
MATCH (m:AstNode {type: 'MODULE', name: $name, project: $project})
MATCH (caller:AstNode)-[:CALLS]->(target:AstNode)
WHERE target = m OR (m)-[:CONTAINS]->(target)
RETURN DISTINCT caller.name AS name, caller.type AS type, caller.sourceFile AS sourceFile
- First finds the anchor
MODULEnodem(scoped byproject). - Then finds any
CALLSedge whose target is eithermitself, or somethingmCONTAINS(i.e. a function defined inside that module) — so calling a function insideSampleRequestcounts as "callingSampleRequest". target = mis a node-identity comparison (same internal node).
3.4 CALLEES — what does this module call?
MATCH (m:AstNode {type: 'MODULE', name: $name, project: $project})
MATCH (m)-[:CONTAINS*0..]->(source:AstNode)-[:CALLS|EXTENDS|IMPLEMENTS]->(callee:AstNode)
RETURN DISTINCT callee.name AS name, callee.type AS type, callee.sourceFile AS sourceFile
[:CONTAINS*0..]is a variable-length path: "zero or moreCONTAINShops."*0meanssourcecan bemitself; more hops would reach functions-within- functions etc.[:CALLS|EXTENDS|IMPLEMENTS]is relationship-type alternation — match any one of these three edge types in a single step.- Result: every module/function/class that
m(or anything it contains) calls, extends, or implements.
3.5 DB_ACCESSES — which DB tables, and how?
MATCH (m:AstNode {type: 'MODULE', name: $name, project: $project})
MATCH (m)-[:CONTAINS*0..]->(f:AstNode)-[r:READS|WRITES]->(t:AstNode {type: 'DB_TABLE'})
RETURN DISTINCT t.name AS name, type(r) AS mode
- Same "module or anything it contains" pattern as
CALLEES. ris bound to the relationship itself (not just a node), sotype(r)returns the relationship's type as a string ("READS"or"WRITES") — this becomes themodein the response.- Filters the target node by
type: 'DB_TABLE'.
3.6 SEARCH_IDENTIFIER — find a name anywhere in the project
MATCH (n:AstNode {project: $project})
WHERE n.name = $name
RETURN n.type AS type, n.name AS name, n.sourceFile AS sourceFile,
n.startLine AS startLine, n.endLine AS endLine,
n.dataType AS dataType, n.value AS value
- A flat scan over all
:AstNodes in the project filtered by exactname. With an index on(project, name)this is fast even for large graphs. - Returns every kind of node (function, variable, constant, ...) that has this name
— useful for "where is
partnerIdused/defined?"
3.7 Project CRUD queries
-- PROJECT_EXISTS
MATCH (p:Project {name: $name}) RETURN p
-- CREATE_PROJECT
CREATE (p:Project {name: $name, description: $description})
-- UPDATE_PROJECT
MATCH (p:Project {name: $name}) SET p.description = $description
-- DELETE_PROJECT (cascades to all AstNodes in the project)
MATCH (p:Project {name: $name})
OPTIONAL MATCH (n:AstNode {project: $name})
DETACH DELETE p, n
-- LIST_PROJECTS
MATCH (p:Project) RETURN p.name AS name, p.description AS description ORDER BY p.name
CREATE(unlikeMERGE) always creates a new node —GraphRepositorychecksPROJECT_EXISTSfirst to return409 Conflictinstead of creating duplicates.DETACH DELETEremoves a node and all its relationships in one step — required because Neo4j refuses to delete a node that still has relationships attached.OPTIONAL MATCHensures the query doesn't fail if the project has noAstNodes yet (an empty/never-ingested project).
4. Using Neo4j for program analysis (and beyond)
4.1 Why a graph model fits program analysis
Most interesting questions about source code are graph questions:
| Question | Graph operation |
|---|---|
| "Who calls this function?" | incoming CALLS edges (1 hop) |
| "What's the full call tree under this module?" | variable-length traversal (*0..N or *0..) |
| "Is there a cycle in the call graph?" | path-finding / cycle detection |
| "What tables does this module touch, transitively?" | traversal + filter by node type |
| "What's the shortest call path from A to B?" | shortestPath() |
| "Which modules are most central / most depended-upon?" | graph algorithms (degree, PageRank via GDS) |
| "Show me everything that would break if I delete this function" | reverse traversal of CALLS/USES_TYPE/READS/WRITES |
A relational database can answer these with recursive CTEs, but each additional hop means another self-join; performance degrades quickly with depth. In Neo4j, traversal cost depends on the size of the result, not the size of the whole table — a 5-hop traversal over a 10-million-node graph is fast if only a handful of nodes match.
4.2 Patterns already used in AgenticCode
- Call graph navigation (
CALLERS/CALLEES): direct application of the table above — "who calls / is called by this module." - Data/DB lineage (
DB_ACCESSES): traversal from a module down to the DB tables it reads/writes, classified by access mode. - Cross-module search (
SEARCH_IDENTIFIER): flat lookup by property, independent of graph structure — Neo4j is also a perfectly good "property store" for this. - Multi-tenancy (
projectproperty): a lightweight way to partition the graph without separate databases, while still allowing (in principle) cross-project queries later if ever needed (e.g. "does any project call this shared library?").
4.3 Ideas for future enrichment / queries
These map directly to the EnrichmentPipelines design and the agent-facing API:
- Impact analysis:
MATCH (n {name:$name})<-[:CALLS|READS|WRITES|USES_TYPE*1..3]-(dep) RETURN dep— "what depends on this, up to 3 hops away?" - Dead code detection:
MATCH (f:AstNode {type:'FUNCTION', project:$project}) WHERE NOT ()-[:CALLS]->(f) RETURN f— functions with no incomingCALLSedge. - Full call chains to a DB table:
MATCH path = (m:AstNode {type:'MODULE'})-[:CONTAINS*0..]->()-[:READS|WRITES]->(t:AstNode {type:'DB_TABLE', name:$table}) RETURN path— every module that ultimately touches a given table, with the full path for explanation. - Cycle detection in the call graph:
MATCH (f:AstNode)-[:CALLS*1..10]->(f)(with a hop limit to bound cost) — useful for spotting recursive or circular dependencies in legacy Natural code.
4.4 Practical tips while developing
- Neo4j Browser (
http://localhost:7474) is the fastest way to explore: run anyCypherQueriesstring directly with literal values to see the actual graph and tune the query before wiring it intoGraphRepository. EXPLAIN/PROFILEa query (PROFILE MATCH ...) to see whether it's using an index or doing a full label scan — important once a project has many thousands ofAstNodes.- Add indexes for properties used in
WHERE/MATCHpatterns, e.g.:(Not yet present in the codebase — worth adding once ingest volumes grow; see the open TODO "Deploy unified AST schema to Neo4j (constraints + indexes)".)CREATE INDEX ast_node_project_name IF NOT EXISTS FOR (n:AstNode) ON (n.project, n.name); CREATE INDEX ast_node_project_type_name IF NOT EXISTS FOR (n:AstNode) ON (n.project, n.type, n.name); - Resetting data during development:
MATCH (n) DETACH DELETE nwipes the entire database (all projects) — use the project-scopedDELETE_PROJECTquery instead when you only want to clear one project.