3. Agent-by-agent deep dive
If you only remember 3 things 1. Only 3 of 7 nodes are LLM-driven (
rca,recommend,code_fix); the rest are deterministic Python. Say this proudly — it's good engineering, not a weakness. 2. Every node has the same contract: readIncidentState, return a partial update dict, append aNodeEventfor the audit trail (backend/app/graph/state.py). 3. The honest answer to "is this an agent?" is: rca, recommend, and code_fix are LLM agents with tools and structured output; intake/retrieval/rag/approval are workflow steps. The graph is the agent system; not every node needs a brain.
The table below is verified against each node file in backend/app/graph/nodes/.
| Node | LLM? | Tools called | Metrics emitted | Failure behavior |
|---|---|---|---|---|
| intake | No | — | — | placeholder event if none supplied |
| retrieval | No | splunk_query.fetch_log_window |
— | fallback "no log window" entry |
| rag | No | Chroma query | — | empty list if no docs |
| rca | Yes | LLM invoke | rca_success/rca_failed |
raises (poller marks run failed, retries) |
| recommend | Yes | LLM invoke | recommend_success/recommend_failed |
skips without LLM call if no RCA |
| approval | No | interrupt() |
approval_approved/approval_rejected |
pauses forever until resumed |
| code_fix | Yes (via tools) | code_search, issue_fix.propose_fix, github_tool.create_fix_branch |
fix_created/fix_failed |
catches everything → FixResult(status="failed") |
intake — normalize the event
File: backend/app/graph/nodes/intake.py
- Purpose: turn the raw polled event into a normalized
ErrorEventand fix the incident ID. - Inputs:
state.error_event(pre-seeded by the poller) orstate.incident_id. - Outputs:
error_event,incident_id(=INC-{error_id}if not pre-seeded),eventsaudit entry. - No LLM. If no event is supplied at all, it creates a placeholder
(
ERR-STUB-001) so the graph can still run in demos/tests. - Tests:
backend/tests/integration/test_stub_flow.py(intake-first ordering),backend/tests/unit/services/test_splunk_normalizer.py(the normalizer it relies on). - "Is this an agent?" — No, and it shouldn't be. Normalization is deterministic
rule-based work (severity mapping, stable IDs —
backend/app/services/splunk_normalizer.py). An LLM here would add latency, cost, and hallucination risk for zero benefit.
retrieval — Splunk log window
File: backend/app/graph/nodes/retrieval.py
- Purpose: fetch the logs around the failure (±5 minutes) so the RCA agent sees what actually happened, not just the error line.
- Inputs:
error_event(txn_id, system, occurred_at, session_id). - Outputs:
retrieved_logs(aLogBundlesorted by timestamp),events. - How:
get_splunk_client().fetch_log_window(...)— a correlated SPL search:(txn_id OR orderNumber OR requestId) OR (system OR application), restricted to the same session,| head 200(backend/app/tools/splunk_query.py). It also merges any incident-event rows already recorded in the app DB for this incident (dedupe by Splunk_cd). - Failure: if the search fails or returns nothing, it writes a single fallback
LogEntry("no log window available") — the graph continues; RCA then works with less evidence and should lower its confidence (the prompt tells it to). - No LLM. The client is injected via
get_splunk_client()/set_splunk_client()(aProtocol), so tests substitute a fake without network access (backend/tests/unit/test_splunk_connection.py,backend/tests/integration/test_rca_dashboard.py). - "Is this an agent?" — No: it's a tool call inside a workflow step. The decision of what to fetch is fixed (correlation keys), which is right — you want determinism about what evidence is collected.
rag — similar historical errors
File: backend/app/graph/nodes/rag.py
- Purpose: retrieve documented error families that resemble this error, so the RCA can say "this looks like CO-500, which historically was a divide-by-zero in checkout."
- Inputs:
error_event(error_code, event, message). - Query:
f"{error_code} {event} {message}"— topmin(RAG_TOP_K=5, collection size)results from Chroma (backend/app/config.py,backend/app/rag/store.py). - Outputs:
similar_error_docs(list ofErrorDocwithdoc_id,content,similarity),events. - Similarity score:
clamp(1.0 - chroma_distance, 0, 1)— explicitly an approximation for display, not a calibrated probability (comment inrag.py). - No LLM. Embedding happens inside Chroma (default
ONNXMiniLM_L6_V2—backend/app/rag/store.py). - Tests:
backend/tests/integration/test_rca_dashboard.py(real ingestion + retrieval with a deterministic offline embedding function).
rca — the core LLM agent
File: backend/app/graph/nodes/rca.py
- Purpose: produce the structured, evidence-cited
RCAResult— root cause, contributing factors, evidence IDs, impacted component, severity, confidence, summary, sequence of events. - Inputs:
error_event,retrieved_logs,similar_error_docs(all three branches of the fan-in). - LLM: yes —
get_chat_model()(temperature 0.0) wrapped inwith_retries(backend/app/llm/client.py). - Flow (in code order):
1.
save_analysis_context(...)— persist what the agent was shown (auditability). 2. Build the user prompt: JSON witherror_event,failed_request(redacted evidence metadata),log_window(each log gets an IDlog-{i}),similar_error_docs(id + content). 3.with_retries(model.invoke)— 3 attempts, exponential backoff. 4.RCAResult.model_validate_json(strip_code_fences(...))— strict Pydantic parse. 5.operational_rca(...)— demo-flavor guardrail (caps confidence at 0.3 if the "cause" is just the injected-error setup —backend/app/services/incident_presentation.py). 6. Evidence-ID whitelist filter: keep only IDs from the supplied docs, logs, and the error itself. 7.save_rca(...)— persist immediately so the dashboard shows it while the run is paused at approval. 8.metrics.increment("rca_success")or"rca_failed"(then re-raise). - Failure: any exception re-raises after counting
rca_failed. The poller marks the analysis failed and does not mark the event processed, so the next poll retries it (backend/app/services/splunk_poller.py). - Tests:
backend/tests/integration/test_rca_dashboard.py(real parsing + persistence with a stubbed LLM),backend/tests/unit/graph/test_builder.py(topology),backend/tests/integration/test_stub_flow.py(override path). - "Is this an agent?" — Yes. It's an LLM that reasons over retrieved evidence, decides what the root cause is, and produces a structured judgment (severity, confidence) that changes the control flow of the system (the router uses its confidence). It has no tool-calling loop, but agentic ≠ "calls tools in a loop" — the pipeline gives it its evidence deterministically, which is a deliberate reliability choice.
recommend — the second LLM agent
File: backend/app/graph/nodes/recommend.py
- Purpose: fill
rca.recommended_actionwith a separate LLM call that reasons only over the completed RCA — never the raw logs or RAG docs. - Why separate (this is a strong rubric point): it shrinks the context to just the
RCA fields, which (a) reduces hallucination surface, (b) enforces the role split
"diagnose" vs "prescribe", and (c) means the recommendation can't cite evidence the
RCA never established. Commit
bda6c87introduced it as "a genuinely separate LLM call that reasons only over the completed RCA". - Inputs:
state.rca(root_cause, contributing_factors, evidence, impacted_component, severity, confidence, summary). - Outputs: updated
rca.recommended_action(re-saved viasave_rcaso the dashboard reflects it),events. - Prompt rule: "If the RCA states the underlying cause could not be determined,
recommend an investigation step, not a fix" (
recommend.py). - Guardrail:
operational_recommendation()scrubs demo-flavored recommendations (backend/app/services/incident_presentation.py). - Failure: counts
recommend_failed; if there is no RCA at all it skips without an LLM call (tested inbackend/tests/unit/graph/nodes/test_recommend.py). - Tests:
backend/tests/unit/graph/nodes/test_recommend.py— skip path, happy path + persistence, guardrail scrub, failure metrics. - "Is this an agent?" — Yes, same argument as rca; and the split is the point: two small specialists beat one generalist here because each has a smaller, cleaner job (see 04-prompt-engineering.md).
approval — the human gate
File: backend/app/graph/nodes/approval.py
- Purpose: pause the run for a human decision. This is the HITL (human-in-the-loop) centerpiece.
- Mechanics: calls
interrupt()with a review payload:{incident_id, root_cause, confidence, recommended_action, similar_error_docs}. The graph saves state and stops. The FIX button resumes it withCommand(resume=ApprovalDecision(status="approved", reviewer=...))(backend/app/api/routes/incidents.py). - Pre-seeded pass-through: if
state.approvalis already set (tests, replays), the node validates it and skips the interrupt — but the router still applies fix-eligibility, so auto-approve can't bypass the confidence/severity policy. - No LLM. The human is the decision-maker here.
- Metrics:
approval_approved/approval_rejected(backend/tests/unit/graph/nodes/test_approval_metrics.py). - "Is this an agent?" — No: it's a control-flow gate. That's exactly what it should be — a deterministic guarantee that no LLM output reaches code without a human.
code_fix — the acting agent
File: backend/app/graph/nodes/code_fix.py
- Purpose: turn an approved, fix-eligible RCA into a reviewable Git branch.
- Only reached when: approval status is
approvedandfix_eligibility(rca)is true (confidence ≥ 0.5, severity ≠ low) —route_after_approvalinbackend/app/graph/builder.py. - Flow:
1. Re-apply
operational_rca(defense in depth — the RCA text is re-scrubbed before it is shown to the fix LLM). 2.search_code(root_cause, impacted_component, evidence)— keyword search overdata/sample_app/, returns top-5 candidate files (backend/app/tools/code_search.py). 3.propose_fix(rca, event, candidates)— LLM must return strict JSON{file_path, fixed_content, explanation}; the path must be one of the candidates or it raises; a no-op fix (empty/unchanged content) raises (backend/app/tools/issue_fix.py). 4.create_fix_branch([FixWrite(path, after_code)])— branchfix/{incident_id}-{timestamp}, commit (--onlyso unrelated staged work can't leak), push if a token is configured, restore the original checkout (backend/app/tools/github_tool.py). 5. BuildFixResult(status="created", branch_name, commit_sha, before_code, after_code, diff, push_status, ...)andsave_fix_result. - Failure: the whole body is one
try/except— any failure becomesFixResult(status="failed")with the error message, and the run still ends cleanly (metricsfix_failed). The API can later re-run justcode_fixon the saved state (backend/app/api/routes/incidents.py— the "fix failed, retry" path). - Tests:
backend/tests/unit/graph/nodes/test_code_fix_metrics.py,backend/tests/integration/test_stub_flow.py(override path),backend/tests/unit/graph/test_builder.py(router gating). - "Is this an agent?" — Yes, the most agentic node: an LLM that acts on the world (writes a file, creates a branch, pushes) through tools, with safety rails (candidate-file restriction, no-op rejection, human gate upstream). If an evaluator wants one example of "an agent", use this one.
The honest framing for "only 3 of 7 nodes are LLM agents"
Say: "We use LLMs only where judgment is required — diagnosis, recommendation, and fix synthesis. Everything else is deterministic Python, because determinism is a feature: you want the same evidence collected the same way every time. This is the standard agentic-pipeline pattern — orchestration around a few LLM reasoning steps — not a chatbot with seven prompts."