3. Agent-by-agent deep dive

If you only remember 3 things 1. Only 3 of 7 nodes are LLM-driven (rca, recommend, code_fix); the rest are deterministic Python. Say this proudly — it's good engineering, not a weakness. 2. Every node has the same contract: read IncidentState, return a partial update dict, append a NodeEvent for the audit trail (backend/app/graph/state.py). 3. The honest answer to "is this an agent?" is: rca, recommend, and code_fix are LLM agents with tools and structured output; intake/retrieval/rag/approval are workflow steps. The graph is the agent system; not every node needs a brain.

The table below is verified against each node file in backend/app/graph/nodes/.

Node LLM? Tools called Metrics emitted Failure behavior
intake No — — placeholder event if none supplied
retrieval No splunk_query.fetch_log_window — fallback "no log window" entry
rag No Chroma query — empty list if no docs
rca Yes LLM invoke rca_success/rca_failed raises (poller marks run failed, retries)
recommend Yes LLM invoke recommend_success/recommend_failed skips without LLM call if no RCA
approval No interrupt() approval_approved/approval_rejected pauses forever until resumed
code_fix Yes (via tools) code_search, issue_fix.propose_fix, github_tool.create_fix_branch fix_created/fix_failed catches everything → FixResult(status="failed")

intake — normalize the event

File: backend/app/graph/nodes/intake.py

retrieval — Splunk log window

File: backend/app/graph/nodes/retrieval.py

rag — similar historical errors

File: backend/app/graph/nodes/rag.py

rca — the core LLM agent

File: backend/app/graph/nodes/rca.py

recommend — the second LLM agent

File: backend/app/graph/nodes/recommend.py

approval — the human gate

File: backend/app/graph/nodes/approval.py

code_fix — the acting agent

File: backend/app/graph/nodes/code_fix.py

The honest framing for "only 3 of 7 nodes are LLM agents"

Say: "We use LLMs only where judgment is required — diagnosis, recommendation, and fix synthesis. Everything else is deterministic Python, because determinism is a feature: you want the same evidence collected the same way every time. This is the standard agentic-pipeline pattern — orchestration around a few LLM reasoning steps — not a chatbot with seven prompts."