7. Human-in-the-loop, memory, and durability
If you only remember 3 things 1. HITL =
interrupt()at the approval node +Command(resume=ApprovalDecision)from the FIX API — the graph pauses, persists, and resumes; it never polls or blocks a thread (backend/app/graph/nodes/approval.py,backend/app/api/routes/incidents.py). 2. Durability = a SqliteSaver checkpointer withthread_id = incident_id; a paused run survives process restarts. The poller resumes failed runs and merges new evidence, but a run paused at approval waits for the human's FIX click (backend/app/graph/checkpointer.py,backend/app/services/splunk_poller.py). 3. "Memory" here = durable per-incident workflow state (checkpointed graph state), not chatbot conversation memory. There is no cross-incident conversational memory — and that's a deliberate scope choice.
Terms
- Interrupt: LangGraph's
interrupt()— a node calls it, the graph halts at that node, persists state, and returns control to the caller. Later,Command(resume=...)injects a value and continues from exactly that point. - Checkpointer: a persistence layer that saves the graph state after every
super-step. Ours is
SqliteSaver(LangGraph's SQLite implementation) writing tobackend/checkpoints.sqlite(backend/app/graph/checkpointer.py). - Thread: a named checkpoint stream. We set
thread_id = incident_id, so each incident's run history is isolated and resumable.
The HITL mechanism, end to end
approval.pybuilds a review payload (incident_id,root_cause,confidence,recommended_action,similar_error_docs) and callsinterrupt(payload).- The graph stops.
graph.invoke(...)returns with the state paused atapproval. The poller detects this (snapshot.next == ("approval",)) and marks the incidentrca_ready(backend/app/services/splunk_poller.py). - The human opens the dashboard (
frontend/src/Incidents.jsx), reads the RCA, and clicks FIX. POST /api/incidents/{id}/fixloads the thread and callsgraph.invoke(Command(resume=ApprovalDecision(status="approved", reviewer=...)))(backend/app/api/routes/incidents.py).- The graph resumes at
approval, writes the decision into state, and the router takes over (backend/app/graph/builder.py).
Why interrupt/resume and not a callback or a blocking wait: the process can restart
between step 2 and 4 without losing the run; no thread sits blocked waiting for a
human; and the approval decision becomes part of the durable audit trail (it's in the
checkpointed state, and approval_approved/approval_rejected counters increment —
backend/app/monitoring/metrics.py).
Tested: backend/tests/integration/test_stub_flow.py asserts the interrupt payload
shape and both resume paths (approve → code_fix, reject → END);
backend/tests/unit/services/test_splunk_poller.py::test_recovers_checkpoint_already_waiting_at_approval
proves the poller finds a run paused at approval, marks the event processed, and does
not re-invoke (graph.invocations == []) — resumption is the human's FIX click.
The checkpointer (backend/app/graph/checkpointer.py)
SqliteSaverfromlanggraph-checkpoint-sqlite, pointed atCHECKPOINT_DB_PATH(defaultbackend/checkpoints.sqlite,backend/app/config.py).JsonPlusSerializerwith a customstate_serde()that allowlists our state classes (IncidentState,RCAResult,ErrorDoc,ApprovalDecision,FixResult,NodeEvent,SplunkErrorEvent,LogEntry— the exact list is in the file). This keeps Pydantic objects round-tripping as their real types, not dicts.thread_id = incident_id— one thread per incident. A second Splunk event for the same incident correlates to the same thread (see below).
The two SQLite databases and their separate jobs
| DB | File | Written by | Job |
|---|---|---|---|
| App DB | backend/app.sqlite3 |
backend/app/db/repository.py (plain sqlite3, no ORM) |
Domain records: incidents, analysis_runs, incident_groups, incident_events, processed_splunk_events — what the dashboard and APIs read |
| Checkpointer DB | backend/checkpoints.sqlite |
LangGraph SqliteSaver |
Graph mechanics: every super-step's state, so a paused run can be resumed after a restart |
Why two databases: different owners, different lifecycles. The app DB is the business record (queryable, paginated, deletable via the API). The checkpointer DB is engine state (opaque, versioned blobs, owned by LangGraph). Merging them would couple business queries to LangGraph's internal format and make either harder to evolve. (Honest note: both are single-file SQLite — fine for a POC, a real deployment would use Postgres for both; see 15-production-roadmap.md.)
What "memory" means here (be precise — the rubric scores this)
| Memory type | Do we have it? | Where |
|---|---|---|
| Short-term / workflow memory — the evolving state of one incident's run (logs, docs, RCA, approval, fix) | Yes | Checkpointed IncidentState, backend/app/graph/state.py |
| Cross-restart durability — a paused run survives a process restart | Yes | SqliteSaver, backend/app/graph/checkpointer.py |
| Long-term / semantic memory — "what did we learn from past incidents?" | Partially, by design choice | The RAG corpus (data/knowledge/error_docs.txt) is curated institutional memory, but it is hand-written, not auto-accumulated from past RCAs |
| Chatbot conversation memory — multi-turn dialogue history with a user | No, deliberately | Not our use case; the pipeline is event-driven, not conversational |
The precise sentence to say in the demo: "Our memory model is durable workflow state: every super-step of every incident is checkpointed to SQLite keyed by incident ID, so an interrupted run resumes exactly where it stopped, even after a restart. We deliberately do not have conversational memory — this is a pipeline, not a chatbot — and long-term knowledge lives in a curated, version-controlled RAG corpus rather than an auto-growing vector store."
New evidence while paused (the subtle case)
If a second Splunk event arrives for an incident while its run is paused at
approval, the poller does not re-run the graph. It uses
graph.update_state(..., as_node="intake") to merge the new evidence into the
checkpointed state, so the human reviews with the fuller picture
(backend/app/services/splunk_poller.py). This is a genuinely non-trivial LangGraph
pattern — worth mentioning in the demo.
Honest limits
- No multi-actor memory: two humans reviewing the same incident share the thread, but there's no per-user session concept.
- Checkpoint DB is append-mostly: no cleanup/TTL policy for old threads.
Command(resume)is only wired to "approved" in the API — rejection exists in the graph, tests, and smoke script but not the dashboard (see 06-workflow-design.md).