2. System architecture
If you only remember 3 things 1. Seven nodes, one shared state:
intake → {retrieval ∥ rag} → rca → recommend → approval → (router) → code_fix → END(backend/app/graph/builder.py). 2. Parallel branches merge safely becauseerrorsandeventsare append-only (operator.addreducers inbackend/app/graph/state.py) and the two branches write different fields. 3. There are two SQLite databases with separate jobs:checkpoints.sqlite(LangGraph resume state) andapp.sqlite3(queryable incident/RCA/fix records) (backend/app/graph/checkpointer.py,backend/app/db/repository.py).
Terms, defined once (plain English)
- LangGraph: a library for building LLM workflows as a graph of nodes. Each node is a Python function that reads a shared state and returns a partial update. Edges define execution order. The framework handles parallelism, merging, and persistence.
- Node: one step in the graph — a plain function
def node(state) -> dict. - State: a single Pydantic object (
IncidentState) shared by all nodes in a run. - Reducer: a rule for merging two writes to the same state field. Default is
"last write wins";
operator.addmeans "append" — used forerrorsandevents. - Checkpointer: storage that saves the state after every node so a run can pause (at the human gate) and resume later, even after a process restart.
- Interrupt: LangGraph's pause mechanism — the graph stops, saves state, and waits
for a
Command(resume=...)call to continue.
Components
| Component | File(s) | Job |
|---|---|---|
| Splunk poller | backend/app/services/splunk_poller.py |
Poll Splunk every 2s, dedupe, group, start/resume graph runs |
| Event normalizer | backend/app/services/splunk_normalizer.py |
Raw Splunk row → ErrorEvent (stable IDs, severity mapping) |
| Correlation service | backend/app/services/incident_correlation.py |
Group errors by request/flow; stable INC-FLOW-* IDs |
| LangGraph workflow | backend/app/graph/builder.py + nodes/*.py |
The 7-node RCA pipeline |
| LLM client | backend/app/llm/client.py |
Cached ChatOpenAI (temperature 0.0), tenacity retries |
| RAG store | backend/app/rag/store.py, rag/ingest.py |
In-memory ChromaDB over data/knowledge/error_docs.txt |
| Splunk client | backend/app/tools/splunk_query.py |
REST Search API: poll errors, fetch log windows |
| Code search | backend/app/tools/code_search.py |
Keyword search over the sample app (no execution) |
| Fix proposal | backend/app/tools/issue_fix.py |
LLM proposes a fix, restricted to candidate files |
| GitHub tool | backend/app/tools/github_tool.py |
Branch/commit/push the fix (sample app only) |
| App DB | backend/app/db/repository.py |
Queryable incidents, RCAs, fixes, dedupe, groups |
| Checkpointer | backend/app/graph/checkpointer.py |
LangGraph resume state (separate SQLite file) |
| API | backend/app/api/routes/*.py |
POS app, incidents dashboard API, metrics, health |
| Frontend | frontend/src/ |
React POS (App.jsx) + RCA dashboard (Incidents.jsx) |
| Monitoring | backend/app/monitoring/metrics.py, logging.py |
Thread-safe counters; structlog JSON logs |
| Guardrails | backend/app/services/incident_presentation.py |
Demo-flavor scrubbing, confidence cap 0.3 |
The graph, as it exists in code today
Verified against backend/app/graph/builder.py (NODE_NAMES and the add_edge calls):
graph TD
S([START]) --> intake
intake -->|fan-out| retrieval
intake -->|fan-out| rag
retrieval -->|fan-in| rca
rag -->|fan-in| rca
rca --> recommend
recommend --> approval
approval -->|interrupt: human gate| G{route_after_approval}
G -->|approved AND fix_eligible| code_fix
G -->|rejected, or approved but not eligible| E([END])
code_fix --> E
- Fan-out:
intakehas edges to bothretrievalandrag. LangGraph runs them concurrently (verified by a test that spies on both executing —backend/tests/unit/graph/test_builder.py). - Fan-in: both
retrievalandraghave edges torca, which waits for both. - Router:
route_after_approvalis a conditional edge afterapproval(builder.py). It callsfix_eligibility(rca): - no RCA → not eligible
confidence < CODE_FIX_MIN_CONFIDENCE(default 0.5,backend/app/config.py) → not eligibleseverity == "low"→ not eligible- otherwise eligible →
code_fix - Interrupt:
approvalcallsinterrupt()before the router runs, so the human gate always fires; the router only decides what happens after the decision (backend/app/graph/nodes/approval.py).
Shared state: IncidentState
backend/app/graph/state.py — a Pydantic model. Nodes return partial updates
(plain dicts); LangGraph merges them.
| Field | Type | Written by | Read by |
|---|---|---|---|
incident_id |
str |
intake | checkpointer thread key, all nodes |
error_event |
ErrorEvent \| None |
intake | retrieval, rag, rca, code_fix |
retrieved_logs |
list[LogBundle] |
retrieval | rca |
similar_error_docs |
list[ErrorDoc] |
rag | rca, approval payload |
rca |
RCAResult \| None |
rca, then recommend (re-save) | approval, code_fix, DB, API |
approval |
ApprovalDecision \| None |
approval | router, code_fix |
fix_result |
FixResult \| None |
code_fix | API, UI, DB |
errors |
list[str] — append |
any node | API, UI |
events |
list[NodeEvent] — append |
any node | UI run timeline |
Reducers and why parallel branches merge safely
In state.py:
errors: Annotated[list[str], operator.add] = []
events: Annotated[list[NodeEvent], operator.add] = []
operator.add tells LangGraph: when two parallel nodes both write this field, append
both instead of one overwriting the other. That is the only field the two parallel
branches could race on — retrieval writes retrieved_logs, rag writes
similar_error_docs, and they never touch each other's fields. So the fan-in at rca
is deterministic: it always sees both the log window and the RAG docs, plus the union of
any events both branches appended. (Tested in backend/tests/unit/graph/test_builder.py
— a parallel-execution test proves both branches run and merge.)
Why this design? The alternative — one sequential node doing Splunk retrieval then
RAG — would work but doubles latency for no benefit: the two lookups are independent
(neither needs the other's output; both only need error_event). A Send()-based
map-reduce was considered and rejected in the git history: "No Send() needed since it's
a static fan-out, not per-item map-reduce" (commit 9f65393).
Runtime lifecycle (what happens when the server starts)
backend/app/main.py lifespan, in order:
init_db()— create app-DB tables (backend/app/db/repository.py).ingest_error_docs()— chunk + embeddata/knowledge/error_docs.txtinto the in-memory Chroma collection. RaisesRuntimeErrorif 0 chunks — the app refuses to start with an empty knowledge base.get_checkpointer()— open SQLite checkpointer (backend/checkpoints.sqlite).build_graph(checkpointer, ...)— wire the 7 nodes.- Start the Splunk poller as an asyncio task (
run_splunk_poller). - On shutdown: set the stop event, wait up to 2s, cancel, close the checkpointer connection.
GET /api/health reports database/graph/rag_documents/splunk_poller status
(backend/app/main.py).
Data flow, end to end
flowchart LR
POS[Keele POS app] -->|HEC| Splunk[(Splunk index incidentiq)]
Splunk -->|REST Search API, poll every 2s| Poller[splunk_poller]
Poller -->|normalize + correlate| Graph[LangGraph run, thread_id = incident_id]
Graph -->|fetch_log_window| Splunk
Graph -->|similarity query| Chroma[(in-memory ChromaDB)]
Graph -->|invoke| LLM[(OpenAI-compatible LLM, gpt-5-codex)]
Graph -->|save_rca / save_fix_result| AppDB[(app.sqlite3)]
Graph -->|checkpoint every node| Ckpt[(checkpoints.sqlite)]
Dashboard[Incidents.jsx /incidents] -->|GET /api/incidents| API[FastAPI]
Dashboard -->|POST /api/incidents/id/fix = Command resume| Graph
Key point for the demo: the FIX button is not a separate code path — it is a graph
resume. POST /api/incidents/{id}/fix checks the snapshot's next node is
approval, then calls graph.invoke(Command(resume=ApprovalDecision(...)))
(backend/app/api/routes/incidents.py).
The two SQLite databases (do not confuse them)
| DB | File | Created/used by | Job |
|---|---|---|---|
| Checkpointer | backend/checkpoints.sqlite (CHECKPOINT_DB_PATH) |
backend/app/graph/checkpointer.py (LangGraph SqliteSaver) |
Resume state for paused runs — the workflow's memory |
| App DB | backend/app.sqlite3 (APP_DB_PATH) |
backend/app/db/repository.py (plain sqlite3) |
Queryable business records: incidents, RCAs, fixes, processed-event dedupe, incident groups, analysis runs |
Why separate? The checkpointer's schema belongs to LangGraph (opaque serialized state);
the app DB is ours to query and render. Mixing them would couple dashboard queries to
framework internals. The checkpointer uses JsonPlusSerializer with an explicit
allowlist of our state classes (state_serde() in backend/app/graph/checkpointer.py).
Known weaknesses (honest)
- The RAG collection is in-memory (
chromadb.EphemeralClient,backend/app/rag/store.py) — rebuilt on every start, one collection per process. - Metrics are in-process counters — reset on restart, no latency histograms
(
backend/app/monitoring/metrics.py). - The corpus is 10 curated entries (
data/knowledge/error_docs.txt— verified count). See 05-rag-pipeline.md. - No LLM output-quality evals exist (see 10-testing.md).