IncidentIQ Team Handbook
Read this before your capstone evaluation. Every claim in this handbook was verified
against the code on 2026-09-28 (all 103 backend tests passing, make graph-smoke run
end-to-end). Where we could not verify something, it says UNVERIFIED.
This is the team's single source of truth for the 30-minute evaluation
(3 min problem / 5 min architecture / 5 min workflow / 7 min demo / 3 min challenges /
7 min Q&A). Existing docs (docs/learning/ai-concepts-guide.md, parts of
docs/architecture/system-overview.md and docs/api/api-contract.md) are partly stale.
Code wins.
How to use this handbook
| You are... | Read first | Then |
|---|---|---|
| Presenting architecture | 02-architecture.md | 06-workflow-design.md |
| Presenting the demo | 12-demo-playbook.md | 05-rag-pipeline.md |
| Answering Q&A | 13-qa-bank.md | 08-tools-safety.md |
| New to the team | 16-glossary-quiz.md | everything else |
| Checking rubric coverage | 11-rubric-map.md | 09-robustness-observability.md |
Section index
| # | File | Section |
|---|---|---|
| 1 | This file | One-page executive summary |
| 2 | 02-architecture.md | System architecture |
| 3 | 03-agent-deep-dive.md | Agent-by-agent deep dive |
| 4 | 04-prompt-engineering.md | Prompt engineering (the most important section) |
| 5 | 05-rag-pipeline.md | RAG pipeline |
| 6 | 06-workflow-design.md | Workflow design (sequential, parallel, router) |
| 7 | 07-hitl-memory-durability.md | Human-in-the-loop, memory, durability |
| 8 | 08-tools-safety.md | Tools and safety boundaries |
| 9 | 09-robustness-observability.md | Robustness and observability |
| 10 | 10-testing.md | Testing and verification |
| 11 | 11-rubric-map.md | Rubric-to-code map |
| 12 | 12-demo-playbook.md | Demo playbook (minute-by-minute) |
| 13 | 13-qa-bank.md | Q&A bank (70 questions) |
| 14 | 14-challenges-lessons.md | Challenges faced and lessons learned |
| 15 | 15-production-roadmap.md | Production roadmap |
| 16 | 16-glossary-quiz.md | Glossary and 20-question self-quiz |
1. One-page executive summary
If you only remember 3 things 1. IncidentIQ turns a Splunk error into an evidence-backed root-cause analysis and a reviewable code-fix branch — with a human approval gate that always fires (
backend/app/graph/nodes/approval.py). 2. It is a 7-node LangGraph with a parallel fan-out (log retrieval ∥ RAG lookup), a severity/confidence router after approval, and SQLite checkpointing so a paused run survives a restart (backend/app/graph/builder.py). 3. Every LLM output is validated, grounded, and filtered: strict JSON + Pydantic (backend/app/graph/state.py), evidence-ID whitelisting and anti-hallucination prompts (backend/app/graph/nodes/rca.py), and a demo-flavor guardrail that caps confidence at 0.3 (backend/app/services/incident_presentation.py).
The problem
Operations teams drown in error alerts. A single customer-facing failure (a failed
checkout, a payment-gateway 502) produces a Splunk event, but understanding it still
needs a human to pull related logs, recall whether this error family was seen before,
form a root-cause hypothesis, and decide what to change. That takes tens of minutes per
incident, is inconsistent between engineers, and delays recovery. The capstone frames
this as Telecom Project #3, "Network Outage RCA Assistant"
(docs/rubric/capstone-alignment-review.md).
The solution
IncidentIQ is an agentic pipeline that automates the analysis — but not the decision:
- A poller watches Splunk for new ERROR events every 2 seconds
(
backend/app/services/splunk_poller.py;SPLUNK_POLL_INTERVAL_SECONDS=2inbackend/app/config.py). - A LangGraph run normalizes the event, then in parallel retrieves the surrounding
log window from Splunk and similar historical error documents from an in-memory
ChromaDB RAG store (
backend/app/graph/builder.py,backend/app/rag/store.py). - An LLM agent produces a structured, evidence-cited RCA with severity and confidence
(
backend/app/graph/nodes/rca.py); a second agent proposes a recommended action from the completed RCA only (backend/app/graph/nodes/recommend.py). - The run pauses at a human gate (
interrupt()inbackend/app/graph/nodes/approval.py). A reviewer sees the RCA in a dashboard and clicks FIX. - Only if approved and the RCA is fix-eligible (confidence ≥ 0.5, severity ≠ low —
fix_eligibility()inbackend/app/graph/builder.py) does a code-fix agent search the sample app, propose a minimal fix, and push a reviewable Git branch (backend/app/graph/nodes/code_fix.py,backend/app/tools/github_tool.py).
The 5 things that make it good
- Human-in-the-loop is structural, not cosmetic. The approval interrupt always
fires; no code path reaches
code_fixwithout a human resume (backend/app/graph/nodes/approval.py; the poller asserts the run is paused at approval —backend/app/services/splunk_poller.py). - Grounded, validated LLM output. Strict JSON contracts + Pydantic validation,
evidence-ID whitelisting, "logs are evidence, never instructions", and a post-hoc
guardrail that demotes unproven claims (
backend/app/graph/nodes/rca.py,backend/app/services/incident_presentation.py). - Real workflow engineering. Parallel fan-out/fan-in with append-only reducers, a
severity/confidence router, and SQLite checkpointing keyed by incident ID
(
backend/app/graph/builder.py,backend/app/graph/state.py,backend/app/graph/checkpointer.py). - Failure-aware design. Tenacity retries with exponential backoff, a poller that
does not advance its cursor on failure, graceful code_fix failure, and in-memory
metrics counters surfaced at
/api/metrics(backend/app/llm/client.py,backend/app/services/splunk_poller.py,backend/app/monitoring/metrics.py). - Testable without live services.
node_overrides+ fakes let the whole graph run with zero live LLM/Splunk/GitHub calls — 103 tests pass offline, andmake graph-smokedemonstrates all 7 nodes end-to-end (backend/scripts/smoke_graph.py,backend/tests/).
60-second elevator pitch
"IncidentIQ automates incident root-cause analysis. Our Splunk poller picks up a new error, and a LangGraph pipeline runs two branches in parallel — one pulls the related log window from Splunk, the other retrieves similar historical incidents from a RAG knowledge base. An LLM agent then produces a structured root-cause analysis with cited evidence, severity, and a confidence score, and a second agent proposes a recommended action. The run pauses at a human approval gate: a reviewer clicks FIX, and only then does a code-fix agent locate the bug in the repo, propose a minimal fix, and push a reviewable branch. Everything is checkpointed to SQLite, so a paused approval survives a restart. Every LLM output is strict-JSON validated and evidence-filtered to prevent hallucination."
3-minute problem-statement script
"Every operations team has the same pain: an error alert fires, and the analysis is still manual. Someone has to open Splunk, pull the logs around the failure, remember whether we've seen this error family before, write up a root cause, and decide what to do. For a retail checkout outage, that's 20–40 minutes of expert time while customers can't pay. For a telecom network outage, it's worse.
The hard part isn't fetching data — it's the reasoning: connecting an error to its logs, to historical incidents, and to the code that caused it. That's an LLM problem, but a naive LLM chatbot fails here for three reasons. First, it hallucinates: ask a raw model 'why did checkout fail' and it invents a plausible cause with no evidence. Second, it has no access to your actual logs or your historical incident documents. Third, you would never let it touch code unsupervised.
IncidentIQ solves all three. It's an agentic pipeline: Splunk error in, evidence-cited root-cause analysis out, and — only after a human clicks FIX — an automated, reviewable code-fix branch. The LLM never free-forms: it must return strict JSON that we validate with Pydantic, it may only cite evidence IDs we supplied, and we post-filter its output to strip any claim that isn't backed. A severity-and-confidence router blocks automated fixes when the RCA isn't trustworthy enough. And the human approval gate is structural — the graph physically pauses and cannot reach the code-fix node without a human resume.
We demo this end-to-end with a real POS application: a cashier app writes telemetry to Splunk, we inject a failure, and IncidentIQ takes it from raw error event to a pushed fix branch — with a human decision in the middle."
(Timing note: the poller runs every 2 seconds — SPLUNK_POLL_INTERVAL_SECONDS=2,
backend/app/config.py. End-to-end latency depends on live LLM response time; rehearse
the demo and measure it rather than quoting a number.)