2. System architecture

If you only remember 3 things 1. Seven nodes, one shared state: intake → {retrieval ∥ rag} → rca → recommend → approval → (router) → code_fix → END (backend/app/graph/builder.py). 2. Parallel branches merge safely because errors and events are append-only (operator.add reducers in backend/app/graph/state.py) and the two branches write different fields. 3. There are two SQLite databases with separate jobs: checkpoints.sqlite (LangGraph resume state) and app.sqlite3 (queryable incident/RCA/fix records) (backend/app/graph/checkpointer.py, backend/app/db/repository.py).

Terms, defined once (plain English)

Components

Component File(s) Job
Splunk poller backend/app/services/splunk_poller.py Poll Splunk every 2s, dedupe, group, start/resume graph runs
Event normalizer backend/app/services/splunk_normalizer.py Raw Splunk row → ErrorEvent (stable IDs, severity mapping)
Correlation service backend/app/services/incident_correlation.py Group errors by request/flow; stable INC-FLOW-* IDs
LangGraph workflow backend/app/graph/builder.py + nodes/*.py The 7-node RCA pipeline
LLM client backend/app/llm/client.py Cached ChatOpenAI (temperature 0.0), tenacity retries
RAG store backend/app/rag/store.py, rag/ingest.py In-memory ChromaDB over data/knowledge/error_docs.txt
Splunk client backend/app/tools/splunk_query.py REST Search API: poll errors, fetch log windows
Code search backend/app/tools/code_search.py Keyword search over the sample app (no execution)
Fix proposal backend/app/tools/issue_fix.py LLM proposes a fix, restricted to candidate files
GitHub tool backend/app/tools/github_tool.py Branch/commit/push the fix (sample app only)
App DB backend/app/db/repository.py Queryable incidents, RCAs, fixes, dedupe, groups
Checkpointer backend/app/graph/checkpointer.py LangGraph resume state (separate SQLite file)
API backend/app/api/routes/*.py POS app, incidents dashboard API, metrics, health
Frontend frontend/src/ React POS (App.jsx) + RCA dashboard (Incidents.jsx)
Monitoring backend/app/monitoring/metrics.py, logging.py Thread-safe counters; structlog JSON logs
Guardrails backend/app/services/incident_presentation.py Demo-flavor scrubbing, confidence cap 0.3

The graph, as it exists in code today

Verified against backend/app/graph/builder.py (NODE_NAMES and the add_edge calls):

graph TD
    S([START]) --> intake
    intake -->|fan-out| retrieval
    intake -->|fan-out| rag
    retrieval -->|fan-in| rca
    rag -->|fan-in| rca
    rca --> recommend
    recommend --> approval
    approval -->|interrupt: human gate| G{route_after_approval}
    G -->|approved AND fix_eligible| code_fix
    G -->|rejected, or approved but not eligible| E([END])
    code_fix --> E

Shared state: IncidentState

backend/app/graph/state.py — a Pydantic model. Nodes return partial updates (plain dicts); LangGraph merges them.

Field Type Written by Read by
incident_id str intake checkpointer thread key, all nodes
error_event ErrorEvent \| None intake retrieval, rag, rca, code_fix
retrieved_logs list[LogBundle] retrieval rca
similar_error_docs list[ErrorDoc] rag rca, approval payload
rca RCAResult \| None rca, then recommend (re-save) approval, code_fix, DB, API
approval ApprovalDecision \| None approval router, code_fix
fix_result FixResult \| None code_fix API, UI, DB
errors list[str] — append any node API, UI
events list[NodeEvent] — append any node UI run timeline

Reducers and why parallel branches merge safely

In state.py:

errors: Annotated[list[str], operator.add] = []
events: Annotated[list[NodeEvent], operator.add] = []

operator.add tells LangGraph: when two parallel nodes both write this field, append both instead of one overwriting the other. That is the only field the two parallel branches could race on — retrieval writes retrieved_logs, rag writes similar_error_docs, and they never touch each other's fields. So the fan-in at rca is deterministic: it always sees both the log window and the RAG docs, plus the union of any events both branches appended. (Tested in backend/tests/unit/graph/test_builder.py — a parallel-execution test proves both branches run and merge.)

Why this design? The alternative — one sequential node doing Splunk retrieval then RAG — would work but doubles latency for no benefit: the two lookups are independent (neither needs the other's output; both only need error_event). A Send()-based map-reduce was considered and rejected in the git history: "No Send() needed since it's a static fan-out, not per-item map-reduce" (commit 9f65393).

Runtime lifecycle (what happens when the server starts)

backend/app/main.py lifespan, in order:

  1. init_db() — create app-DB tables (backend/app/db/repository.py).
  2. ingest_error_docs() — chunk + embed data/knowledge/error_docs.txt into the in-memory Chroma collection. Raises RuntimeError if 0 chunks — the app refuses to start with an empty knowledge base.
  3. get_checkpointer() — open SQLite checkpointer (backend/checkpoints.sqlite).
  4. build_graph(checkpointer, ...) — wire the 7 nodes.
  5. Start the Splunk poller as an asyncio task (run_splunk_poller).
  6. On shutdown: set the stop event, wait up to 2s, cancel, close the checkpointer connection.

GET /api/health reports database/graph/rag_documents/splunk_poller status (backend/app/main.py).

Data flow, end to end

flowchart LR
    POS[Keele POS app] -->|HEC| Splunk[(Splunk index incidentiq)]
    Splunk -->|REST Search API, poll every 2s| Poller[splunk_poller]
    Poller -->|normalize + correlate| Graph[LangGraph run, thread_id = incident_id]
    Graph -->|fetch_log_window| Splunk
    Graph -->|similarity query| Chroma[(in-memory ChromaDB)]
    Graph -->|invoke| LLM[(OpenAI-compatible LLM, gpt-5-codex)]
    Graph -->|save_rca / save_fix_result| AppDB[(app.sqlite3)]
    Graph -->|checkpoint every node| Ckpt[(checkpoints.sqlite)]
    Dashboard[Incidents.jsx /incidents] -->|GET /api/incidents| API[FastAPI]
    Dashboard -->|POST /api/incidents/id/fix = Command resume| Graph

Key point for the demo: the FIX button is not a separate code path — it is a graph resume. POST /api/incidents/{id}/fix checks the snapshot's next node is approval, then calls graph.invoke(Command(resume=ApprovalDecision(...))) (backend/app/api/routes/incidents.py).

The two SQLite databases (do not confuse them)

DB File Created/used by Job
Checkpointer backend/checkpoints.sqlite (CHECKPOINT_DB_PATH) backend/app/graph/checkpointer.py (LangGraph SqliteSaver) Resume state for paused runs — the workflow's memory
App DB backend/app.sqlite3 (APP_DB_PATH) backend/app/db/repository.py (plain sqlite3) Queryable business records: incidents, RCAs, fixes, processed-event dedupe, incident groups, analysis runs

Why separate? The checkpointer's schema belongs to LangGraph (opaque serialized state); the app DB is ours to query and render. Mixing them would couple dashboard queries to framework internals. The checkpointer uses JsonPlusSerializer with an explicit allowlist of our state classes (state_serde() in backend/app/graph/checkpointer.py).

Known weaknesses (honest)