Files
agenticCode/README.md
Ingo Schnabel 7e0d75cc5e Roadmap
2026-09-23 12:39:15 +02:00

521 lines
31 KiB
Markdown

# AgenticCode
> AI-agent-optimized code analysis platform — parse, store, enrich, and query source code as a graph.
AgenticCode ingests source code (Natural/Software AG, Java and TypeScript/React), parses it into a unified AST, persists
it in Neo4j,
enriches it with semantic information (call graphs, DB accesses, data structures, dynamic-dispatch resolution), and
exposes everything through an agent-ready **REST** API, a **CLI**, and a **web UI**.
---
## Architecture
```mermaid
flowchart TD
subgraph Sources["Source Files"]
N([Natural\nSoftware AG])
J([Java])
PY([Python\n_planned_])
CB([COBOL\n_planned_])
end
subgraph Parsers["Parser Layer"]
NP["NaturalParser\n<i>custom implementation</i>"]
JP["JavaParser\n<i>com.github.javaparser</i>"]
PI["LanguageParser SPI\n<i>shared interface + unified AST schema</i>"]
end
subgraph Store["AST Store"]
NEO[("Neo4j 5.x\nGraph Database\nModules · Functions · Variables\nData Structures · DB Tables")]
end
subgraph Enrichment["Enrichment Layer"]
CG["CallGraphEnricher\nPERFORM / CALLNAT → CALLS edges"]
DB["DbAccessEnricher\nREAD / FIND / STORE → READS / WRITES edges"]
ID["IdentifierIndexEnricher\ncross-module variable index"]
DS["DataStructureEnricher\nDEFINE DATA → USES_TYPE edges"]
end
subgraph API["Agentic API — Quarkus"]
REST["REST API\nJAX-RS · RESTEasy Reactive · tool-use optimized"]
end
subgraph Clients["Clients"]
AG(["Claude / LLM Agent\nvia REST"])
CLI(["ac CLI"])
UI(["Web UI\nReact + Vite"])
end
N --> NP
J --> JP
PY --> PI
CB --> PI
NP --> PI
JP --> PI
PI --> NEO
NEO --> CG & DB & ID & DS
CG & DB & ID & DS --> REST
REST --> AG & CLI & UI
```
---
## Quick Start
**Prerequisites:** Java 21+, Maven 3.9+, Docker (with the Compose plugin). Node 20+ only if you want to run the web UI.
The repo ships a `docker-compose.yml` (Neo4j 5 + `ac-code-server` + `ac-ui`) and a `manage-ac.sh` wrapper.
```bash
git clone https://github.com/your-org/agenticcode.git
cd agenticcode
# Build all modules, start Neo4j + ac-code-server + ac-ui, and install the `ac` CLI launcher
./manage-ac.sh deploy
```
`./manage-ac.sh deploy` builds the project, brings the Compose stack up (leaving an already-running Neo4j untouched),
serves the web UI at `http://localhost:5174`, and installs the `ac` launcher to `~/.local/bin/ac`. See
[`manage-ac.sh`](#manage-acsh--the-stack-manager) below for all commands.
Once up:
- **Web UI** — `http://localhost:5174`
- **REST API** — `http://localhost:8787/api`
- **OpenAPI / health** — `http://localhost:8787/q/openapi`, `http://localhost:8787/q/health`
- **Neo4j browser** — `http://localhost:7474` (credentials `neo4j` / `agenticcode`)
### Ingest a codebase
Ingestion is **project-root based**: you register a project pointing at a server-side source root, then trigger a scan.
There is no per-file upload step.
```bash
# 1. Create a project pointing at a source root, with its language
ac project create upms /path/to/natural-sources -l natural -d "UPMS legacy"
# 2. Scan the root into the graph
# default: fast whole-root pass (call graph + identifier index; field/dataflow detail is
# resolved lazily per module on demand). Add --deep for a full field-level ingest up front.
ac refresh -p upms
ac refresh -p upms --deep # full field-level pass
# 3. Query it
ac callees WGEAGB0S -p upms
```
Re-run `ac refresh` after the sources change; it reconciles per file (unchanged files are left as-is).
`ac refresh <MODULE> -p upms` deep-ingests one module plus its transitive `CALLNAT`/`PERFORM` dependency tree (lazy
Tier-2).
### `manage-ac.sh` — the stack manager
`manage-ac.sh` builds and runs the whole stack (Neo4j + `ac-code-server` + `ac-ui`) via docker-compose and installs the
`ac` CLI. Run it with no argument (or `help`) to print the command list — a bare invocation deliberately does **not**
deploy, since a full deploy bumps the version and rebuilds everything.
```bash
./manage-ac.sh <command>
```
| Command | What it does |
|-------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `deploy` | Full deploy: bump version, `mvn clean install`, rebuild + restart `ac-code-server`, rebuild + start `ac-ui`, install/refresh the `ac` CLI. Neo4j is left running if already up. |
| `restart` | Restart `ac-code-server` only — **no build**. Also the way to abort a long server-side job (a deep refresh keeps running after its HTTP client is killed). |
| `stop` | Stop `ac-code-server` only; Neo4j and `ac-ui` keep running. |
| `down` | Stop the whole stack, Neo4j included. **The graph volume is kept** (never `down -v`). |
| `status` | Show containers, the answering server version, and the ingested projects — works even when the stack is down. |
| `cli` | Build and (re)install `ac` only — no Docker involved. |
| `ui` | Rebuild + restart `ac-ui` only (npm build runs inside Docker; no local Node needed). |
| `logs [-f]` | Last 200 lines of server logs; `-f` to follow. |
The server answers on `http://localhost:8787`, and the **Dockerized UI on `http://localhost:5174`** (distinct from the
local Vite dev server on 5173). The version bump lives in the Maven build, so every `deploy` (a full `install`) bumps
`agenticcode.version` and re-stamps the CLI; `mvn test`/`compile`/`quarkus:dev` do not.
### `rebuild-and-refresh.sh` — redeploy then deep-refresh
A one-shot convenience script that redeploys the server and re-ingests the given project(s) from scratch — use it after
code changes that affect parsing or enrichment, so the graph reflects the new build. **One or more project names are
required** (there is no default; running it with no argument prints usage and exits).
```bash
./rebuild-and-refresh.sh upms # one project
./rebuild-and-refresh.sh upms pur # several, refreshed in order
```
It runs the full sequence, blocking until done: **stop** the server → **`manage-ac.sh deploy`** (version bump +
server/UI rebuild) → **wait** for `http://localhost:8787/api/projects` to answer (timeout `READY_TIMEOUT`, default 300
s) → **deep-refresh** each named project synchronously → print per-project timings and ring the terminal bell (and
`notify-send` if available).
Overridable via env: `AC` (CLI launcher, default `ac`), `AC_SERVER_URL` (default `http://localhost:8787`),
`READY_TIMEOUT`.
> A deep refresh is long and mutates the graph — **don't interrupt it once running**; the earlier enrichment steps are
> already committed, so an aborted refresh leaves the graph half-updated.
### Dev mode (hot reload)
```bash
mvn clean install
docker compose up -d neo4j # just the Neo4j service
cd ac-code-server && mvn quarkus:dev # server with hot reload
```
---
## REST API
All analysis endpoints are scoped to a project: `/api/projects/{project}/...`. Responses are concise machine-readable
JSON with `limit`/`offset` pagination where lists can be large.
### Projects & ingestion
| Endpoint | Description |
|----------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------|
| `GET /api/projects` | List all projects |
| `POST /api/projects/{project}` | Create a project (body: `root`, `language`, optional `description`, `excludeDirs`, `generatedDir`, `userExitDir`, `counterparts`) |
| `PUT /api/projects/{project}` | Update a project (unset fields left unchanged) |
| `DELETE /api/projects/{project}` | Delete a project and all its data |
| `POST /api/projects/{project}/refresh[?deep=true]` | (Re)scan the whole root — call-graph pass, or full field-level with `deep=true` |
| `POST /api/projects/{project}/refresh/{name}` | Deep-ingest one module + its dependency tree (`maxDepth`, `maxNodes`, `neighborhood`) |
| `POST /api/projects/{project}/recreate` | Drop and rebuild the project's graph from disk |
### Call graph & structure
| Endpoint | Description |
|----------------------------------------------------------------------------------|------------------------------------------------------------|
| `GET .../modules` | List modules (filter by name/kind) |
| `GET .../modules/{name}/callers` | Who calls this module (`scope=external\|internal`) |
| `GET .../modules/{name}/callees` | What this module calls (`scope=external\|internal`) |
| `GET .../modules/{name}/call-tree` | Recursive callee tree |
| `GET .../modules/{name}/context` | Compact module summary for agents |
| `GET .../modules/{name}/digest` | One-line module digest |
| `GET .../modules/{name}/functions` | Subroutines/methods in the module |
| `GET .../modules/{name}/functions/{fn}/callers` | Callers of a specific function |
| `GET .../modules/{name}/functions/{fn}/overrides` | Polymorphic overrides of a method |
| `GET .../modules/{name}/graph` | Ego graph (bounded neighbourhood; `dir`, `depth`, `limit`) |
| `GET .../modules/{name}/source` · `GET .../nodes/{id}/source` · `GET .../source` | Source text of a module / node / file |
### Data, DB access & payload
| Endpoint | Description |
|-----------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------|
| `GET .../modules/{name}/db-accesses` | DB tables accessed and mode (READ/WRITE) |
| `GET .../store?slice=` · `GET .../store/{slice}/accesses?field=&mode=` | Frontend Redux store: slices, state keys, who reads/writes them (item 194) |
| `GET .../bindings?dto=&field=&mode=&module=` | DTO field bindings: which component reads/writes which backend field, with its Java counterpart (item 195) |
| `GET .../theme?unused=` · `GET .../theme/{token}/usages` · `GET .../styles` | MUI theme tokens with use counts, where a token is read, and the sx/style/styled/CSS inventory with hard-coded literals (item 196) |
| `GET .../modules/{name}/sql-statements` | SQL/ADABAS statements |
| `GET .../modules/{name}/workfile-accesses` | Natural work-file reads/writes |
| `GET .../modules/{name}/data-structures` | Data structures used by the module |
| `GET .../modules/{name}/payload` | Parameter-data-area I/O contract |
| `GET .../modules/{name}/columns` · `.../modules/{name}/dispatch-table` | Entity columns · DECIDE dispatch table |
| `GET .../data-structures/{name}/fields` · `.../db-tables/{name}/columns` | Fields of a structure · columns of a table |
### Search & dataflow
| Endpoint | Description |
|----------------------------------------------------------------------------------|---------------------------------------------------------------------------------------------|
| `GET .../search/identifier?name=[&priorityModule=]` | Find an identifier across modules (`priorityModule` pins one module's matches to the front) |
| `GET .../search/value` · `.../search/source` · `.../search/annotation` | Search literals · source text · annotations |
| `GET .../variables/{name}/reads` · `.../writes` | Where a variable is read / written (optional `module=`) |
| `GET .../variables/{name}/flow-forward` · `.../flow-backward` · `.../field-flow` | Argument→parameter dataflow, and shared-field producer→consumer |
| `GET .../nodes/{id}` | Inspect a raw graph node |
| `GET .../loc` · `GET /api/version` | Project LOC metrics · server version |
### Dynamic-dispatch (Natural `CALLNAT <var>`)
| Endpoint | Description |
|--------------------------------------|--------------------------------------------------------------------|
| `GET .../dynamic-calls/unresolved` | Dynamic call sites no resolver recovered — candidates to pin |
| `GET .../dynamic-calls/overrides` | Manual overrides (flagged `obsolete` if auto-resolution caught up) |
| `POST .../dynamic-calls/overrides` | Pin a dynamic call site to target module(s) |
| `DELETE .../dynamic-calls/overrides` | Reset one override (`originFile`+`lineNo`) or all |
Dynamic `CALLNAT` targets are recovered automatically where possible: direct string literals, indirect lookup-array
copies, cross-module parameter flow, and **constant-folded string assembly** (`MOVE 'YABALKEY' TO #M` +
`MOVE 'GN0' TO SUBSTR(#M,6,3)` → `YABALGN0`). A manual override always wins over an auto-fold. What can't be recovered
stays visible as an unresolved site.
---
## Agent usage (Claude Code & other agents)
The server exposes every query capability as REST endpoints under `http://localhost:8787/api` — e.g.
`/projects`, `/projects/{p}/modules`, `.../callers`, `.../callees`, `.../call-tree`, `.../db-accesses`,
`/search/identifier`, `.../context`, `.../payload`, `.../dispatch-table`, `/variables/{n}/flow-forward` and
`/flow-backward`, `/variables/{n}/field-flow`, `/variables/{n}/reads` and `/writes`,
`/dynamic-calls/unresolved`, `/dynamic-calls/overrides`, `/refresh`, and more. The OpenAPI spec is served at
`/q/openapi`. A full usage guide with response fields and semantics lives in [
`x-docs/agent-api-usage-ac-implementation.md`](x-docs/agent-api-usage-ac-implementation.md).
---
## Web UI
A React + Vite + Tailwind front-end (`ac-ui/`) for browsing projects and modules interactively — built for reading
legacy Natural/Java and planning migrations.
### Run it
If you ran `./manage-ac.sh deploy` (or `./manage-ac.sh ui`), the UI is **already built and served in Docker
at `http://localhost:5174`** — no local Node needed. For front-end development, run the Vite dev server instead:
```bash
cd ac-ui
npm install
npm run dev # Vite dev server on http://localhost:5173 (auto-bumps to 5174 etc. if taken)
```
It talks to the REST API on `http://localhost:8787` (the same server `./manage-ac.sh deploy` starts). Other scripts:
`npm run build` (type-check + production build), `npm run preview` (serve the build), `npm run gen:api` (regenerate the
typed client `src/api/schema.ts` from the live server's `/q/openapi`), `npm run e2e` (Playwright tests).
### How to use it
**1. Pick a project.** The landing page lists every project (name, language, root). Click one — or use the project
dropdown in the top bar to switch at any time.
**2. Find a module.** The left pane lists all modules. Filter with the search box (a **regex** on name or file — toggle
**names / source** to search identifiers *inside* the code instead), and narrow by kind with the **all kinds** dropdown.
Each row shows the module kind, sloc, and ingest status (`INGESTED · FULL`, call-graph-only, …). Click a module to open
it. `refresh project` re-scans the root; the per-module `refresh` / `ingest +callers/callees` buttons deepen just that
module on demand.
**3. Read the module.** The detail pane has one tab per view:
| Tab | What it shows |
|---------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **Source** | The full source with a subroutine **outline** (click to jump), **find-in-file** (regex, `n/m` match counter), and **Ctrl/⌘-click any identifier to "identify"** it (see below). |
| **Dossier** | The I/O contract — every parameter-data-area field with direction (REQUEST/RESPONSE) and source; line numbers link back into Source. |
| **Data-flow** | Trace one field: **comes-from (backward)**, **flows-to (forward)**, and **field-flow (producer → consumer)** across calls. Seed it from a field's **data-flow →** link in the identify popover. |
| **Impact** | The blast radius — transitive **callers** grouped by hop distance (direct / 2 hops / 3 hops …). |
| **Overview** | One-screen summary: description, kind, loc/sloc, functions, all callees (internal + external), callers, DB tables, plus a per-browser **notes** field for migration annotations. |
| **Calls** | **Called-by** (callers) and **calls** (callees) side by side, each tagged with the edge kind (`CALLNAT`, `PERFORM`, `CALLNAT_DYNAMIC`, `EXTENDS`, …). |
| **Call-tree** | The transitive callee tree; nodes **expand lazily** on click. |
| **Graph** | An interactive **ego graph** around the module, with `dir` (in/out/both), `depth`, and `limit` controls. Dynamic/unresolved edges are colour-coded (orange); a legend explains the node/edge types. |
**4. Identify an identifier (Ctrl/⌘-click in Source).** A popover shows every match for that name across the project (
e.g. *"25 matches"*), with the **current module's** declaration pinned to the top and the rest grouped as *"N in other
modules"*. Expand an entry for its type, scope, and its **reads/writes** occurrence list, plus **open definition →** and
**data-flow →** links that jump to the definition or seed the Data-flow tab.
**5. Save a view.** `save view` / `saved (N)` in the top bar bookmarks the current module+tab so you can jump back while
working through a codebase.
---
## CLI
`ac` (installed by `./manage-ac.sh deploy` to `~/.local/bin/ac`) is a client for ingesting and querying the graph —
handy for scripting and bulk work without going through HTTP directly.
```bash
ac --help
ac version
```
Every command prints the API's JSON response, pretty-printed, and exits non-zero on error. The general form is:
```
ac <command> [args] [-p <project>] [-s <server>] [--limit N] [--offset N]
```
**Where the server and project come from** (each row overrides the ones below it):
| Setting | Command flag | Environment | Config file (`~/.agenticcode/config.properties`) | Shell command |
|------------|--------------------|-----------------|--------------------------------------------------|-----------------|
| Server URL | `-s` / `--server` | `AC_SERVER_URL` | `server.url` | `connect <url>` |
| Project | `-p` / `--project` | `AC_PROJECT` | `project` | `use <project>` |
Default server is `http://localhost:8787`. `connect` and `use` write their values to the config file, so once set they
persist across invocations and you can drop `-s`/`-p`. Commands that need a project fail with a clear error if none is
selected.
> Not deployed via `manage-ac.sh`? Build the uber-jar with `mvn -pl ac-cli -am package` and run
> `java -jar ac-cli/target/ac-cli-*.jar ...` — same commands.
### Interactive shell
Running `ac` with no arguments starts a `psql`-like shell with line editing and history (`~/.agenticcode_history`):
```
AgenticCode interactive shell. Type 'help' for commands, 'exit' to quit.
agenticcode> connect http://my-server:8787
agenticcode> use upms
agenticcode> callers WGEAGB0S
agenticcode> exit
```
`connect <url>` and `use <project>` set (and persist) the server/project for the session.
### Common commands
```bash
# Projects
ac project create upms /path/to/sources -l natural -d "UPMS legacy"
ac project list
ac project update upms -d "New description"
ac project rename upms upms_alt # every node, override and counterpart reference follows
ac project delete upms
# Ingest / refresh
ac refresh -p upms # fast whole-root scan
ac refresh -p upms --deep # full field-level
ac refresh WGEAGB0S -p upms # one module + its dependency tree
# Query the call graph
ac callers WGEAGB0S -p upms
ac callees WGEAGB0S -p upms
ac call-tree WGEAGB0S -p upms
ac context WGEAGB0S -p upms
ac functions WGEAGB0S -p upms
# Data & DB
ac db-accesses WGEAGB0S -p upms
ac sql-statements WGEAGB0S -p upms
ac payload WGEAGB0S -p upms
ac data-structure-fields SOME-PDA -p upms
# Search & dataflow
ac search-identifier I-LINE-LEV -p upms
ac search-value 'YABAL' -p upms
ac variable-reads '#I-LINE-LEV' -p upms
ac flow-forward '#SOME-FIELD' -p upms
# Dynamic dispatch
ac dynamic-calls unresolved -p upms
ac dynamic-calls set --file X.nat --line 403 --target YABALGN0 -p upms
ac dynamic-calls reset --file X.nat --line 403 -p upms
```
### Command reference
`ac --help` lists everything; `ac <command> --help` shows a command's options. The main commands:
| Group | Commands |
|-----------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **Project & ingest** | `project create\|update\|list\|delete`, `refresh [MODULE] [--deep] [--max-depth N] [--max-nodes N] [--neighborhood]`, `recreate` |
| **Call graph** | `callers`, `callees`, `call-tree`, `context`, `digest`, `functions`, `function-callers`, `function-overrides`, `ego-graph`, `modules`, `loc` |
| **Data & DB** | `db-accesses`, `sql-statements`, `workfile-accesses`, `payload`, `dispatch-table`, `module-data-structures`, `data-structure-fields`, `db-table-columns`, `entity-columns`, `store`, `store-accesses`, `bindings`, `theme`, `theme-usages`, `styles` |
| **Search & dataflow** | `search-identifier`, `search-value`, `search-source`, `search-annotation`, `variable-reads`, `variable-writes`, `flow-forward`, `flow-backward`, `field-flow` |
| **Source & nodes** | `module-source`, `file-source`, `node-source`, `inspect-node` |
| **Dynamic dispatch** | `dynamic-calls unresolved\|overrides\|set\|reset` |
| **Session / misc** | `connect`, `use`, `version` |
**Useful options.** Most list commands accept `--limit` / `--offset` for pagination. `callers` and `callees` take
`--scope external` (module-to-module `CALLNAT`/inheritance — the default for `callers`) or `--scope internal` (
own-subroutine `PERFORM` wiring). `search-identifier` takes `--type`, `--module`, `--source-file`, and
`--priority-module` (pin one module's matches to the front so its local declaration survives the limit). For Natural
names, a leading sigil (`#`, `&`, `+`) is optional — `ac search-identifier I-LINE-LEV` and `'#I-LINE-LEV'` are
equivalent.
Everything the CLI does is also available as a REST endpoint — the two are kept in lockstep.
---
## Tech Stack
| Component | Technology |
|----------------|---------------------------------------------------------|
| Server | [Quarkus](https://quarkus.io) (latest stable) |
| Language | Java 21 — Records, Sealed Classes, Virtual Threads |
| Graph DB | Neo4j 5.x via `neo4j-java-driver` |
| REST | RESTEasy Reactive (JAX-RS) |
| Natural Parser | Custom implementation |
| Java Parser | [JavaParser](https://javaparser.org) |
| CLI | Picocli uber-jar |
| Web UI | React + Vite + Tailwind, typed via `openapi-typescript` |
| Tests | JUnit 5 + RestAssured + Testcontainers |
| Build | Maven (multi-module) |
---
## Neo4j Graph Schema
```mermaid
erDiagram
MODULE ||--o{ FUNCTION : CONTAINS
FUNCTION ||--o{ FUNCTION : CALLS
FUNCTION ||--o{ VARIABLE : READS
FUNCTION ||--o{ VARIABLE : WRITES
FUNCTION ||--o{ DB_TABLE : READS
FUNCTION ||--o{ DB_TABLE : WRITES
FUNCTION ||--o{ DATA_STRUCTURE : USES_TYPE
MODULE {
string id
string name
string language
string sourceFile
}
FUNCTION {
string id
string name
int startLine
int endLine
}
VARIABLE {
string id
string name
string dataType
}
DATA_STRUCTURE {
string id
string name
string scope
}
DB_TABLE {
string id
string name
string dbType
}
```
Full schema with Cypher examples: [`x-docs/ast-graph-schema.md`](x-docs/ast-graph-schema.md).
---
## Project Structure
```
agenticcode/
├── ac-code-server/ # Quarkus application (REST API)
├── ac-cli/ # Picocli command-line client (`ac`)
├── ac-ui/ # React + Vite web UI
├── ac-parser-core/ # Shared AST model + LanguageParser SPI
├── ac-parser-natural/ # Custom Natural/Software AG parser
├── ac-parser-java/ # Java parser (JavaParser-based)
├── ac-neo4j-store/ # Neo4j persistence + Cypher queries
├── docker-compose.yml # Neo4j + ac-code-server
├── manage-ac.sh # build / run / install-CLI wrapper
└── x-docs/ # architecture, schema, API-usage, roadmap, features
```
---
## Testing
```bash
mvn test # all tests (integration tests use Testcontainers Neo4j)
mvn test -Dtest="*Test" # unit tests only
mvn test -Dtest="*IT" # integration tests only (need Docker)
```
Integration tests spin up their own Neo4j via Testcontainers — independent of the Compose stack.
---
## Supported Languages
| Language | Status | Parser |
|-----------------------|-------------|------------|
| Natural (Software AG) | ✅ Supported | Custom |
| Java | ✅ Supported | JavaParser |
| Python | 📋 Planned | — |
| COBOL | 📋 Planned | — |
```