7. Human-in-the-loop, memory, and durability

If you only remember 3 things 1. HITL = interrupt() at the approval node + Command(resume=ApprovalDecision) from the FIX API — the graph pauses, persists, and resumes; it never polls or blocks a thread (backend/app/graph/nodes/approval.py, backend/app/api/routes/incidents.py). 2. Durability = a SqliteSaver checkpointer with thread_id = incident_id; a paused run survives process restarts. The poller resumes failed runs and merges new evidence, but a run paused at approval waits for the human's FIX click (backend/app/graph/checkpointer.py, backend/app/services/splunk_poller.py). 3. "Memory" here = durable per-incident workflow state (checkpointed graph state), not chatbot conversation memory. There is no cross-incident conversational memory — and that's a deliberate scope choice.

Terms

The HITL mechanism, end to end

  1. approval.py builds a review payload (incident_id, root_cause, confidence, recommended_action, similar_error_docs) and calls interrupt(payload).
  2. The graph stops. graph.invoke(...) returns with the state paused at approval. The poller detects this (snapshot.next == ("approval",)) and marks the incident rca_ready (backend/app/services/splunk_poller.py).
  3. The human opens the dashboard (frontend/src/Incidents.jsx), reads the RCA, and clicks FIX.
  4. POST /api/incidents/{id}/fix loads the thread and calls graph.invoke(Command(resume=ApprovalDecision(status="approved", reviewer=...))) (backend/app/api/routes/incidents.py).
  5. The graph resumes at approval, writes the decision into state, and the router takes over (backend/app/graph/builder.py).

Why interrupt/resume and not a callback or a blocking wait: the process can restart between step 2 and 4 without losing the run; no thread sits blocked waiting for a human; and the approval decision becomes part of the durable audit trail (it's in the checkpointed state, and approval_approved/approval_rejected counters increment — backend/app/monitoring/metrics.py).

Tested: backend/tests/integration/test_stub_flow.py asserts the interrupt payload shape and both resume paths (approve → code_fix, reject → END); backend/tests/unit/services/test_splunk_poller.py::test_recovers_checkpoint_already_waiting_at_approval proves the poller finds a run paused at approval, marks the event processed, and does not re-invoke (graph.invocations == []) — resumption is the human's FIX click.

The checkpointer (backend/app/graph/checkpointer.py)

The two SQLite databases and their separate jobs

DB File Written by Job
App DB backend/app.sqlite3 backend/app/db/repository.py (plain sqlite3, no ORM) Domain records: incidents, analysis_runs, incident_groups, incident_events, processed_splunk_events — what the dashboard and APIs read
Checkpointer DB backend/checkpoints.sqlite LangGraph SqliteSaver Graph mechanics: every super-step's state, so a paused run can be resumed after a restart

Why two databases: different owners, different lifecycles. The app DB is the business record (queryable, paginated, deletable via the API). The checkpointer DB is engine state (opaque, versioned blobs, owned by LangGraph). Merging them would couple business queries to LangGraph's internal format and make either harder to evolve. (Honest note: both are single-file SQLite — fine for a POC, a real deployment would use Postgres for both; see 15-production-roadmap.md.)

What "memory" means here (be precise — the rubric scores this)

Memory type Do we have it? Where
Short-term / workflow memory — the evolving state of one incident's run (logs, docs, RCA, approval, fix) Yes Checkpointed IncidentState, backend/app/graph/state.py
Cross-restart durability — a paused run survives a process restart Yes SqliteSaver, backend/app/graph/checkpointer.py
Long-term / semantic memory — "what did we learn from past incidents?" Partially, by design choice The RAG corpus (data/knowledge/error_docs.txt) is curated institutional memory, but it is hand-written, not auto-accumulated from past RCAs
Chatbot conversation memory — multi-turn dialogue history with a user No, deliberately Not our use case; the pipeline is event-driven, not conversational

The precise sentence to say in the demo: "Our memory model is durable workflow state: every super-step of every incident is checkpointed to SQLite keyed by incident ID, so an interrupted run resumes exactly where it stopped, even after a restart. We deliberately do not have conversational memory — this is a pipeline, not a chatbot — and long-term knowledge lives in a curated, version-controlled RAG corpus rather than an auto-growing vector store."

New evidence while paused (the subtle case)

If a second Splunk event arrives for an incident while its run is paused at approval, the poller does not re-run the graph. It uses graph.update_state(..., as_node="intake") to merge the new evidence into the checkpointed state, so the human reviews with the fuller picture (backend/app/services/splunk_poller.py). This is a genuinely non-trivial LangGraph pattern — worth mentioning in the demo.

Honest limits