16. Glossary and 20-question self-quiz
If you only remember 3 things 1. If a teammate can pass the 20-question quiz below cold, they can survive Q&A. 2. Every term is defined the way we use it in this codebase — not the textbook way. 3. Quiz answers are at the bottom; don't peek until you've answered out loud.
Glossary
| Term | Plain-English meaning (as used in IncidentIQ) |
|---|---|
| Agent | An LLM step that decides something unscripted from unstructured input (diagnose, prescribe, choose a fix). Our control flow around it is deterministic on purpose. |
| At-least-once processing | An event may be processed more than once in failure cases, but never silently dropped. Made safe by dedupe (processed_splunk_events). |
| Checkpoint | A saved snapshot of graph state after every super-step, so a paused run survives restarts (SqliteSaver). |
| Chunking | Splitting the knowledge file into pieces — here, one per ----separated entry. |
| Command(resume) | The value injected into a paused graph to continue it — our FIX button sends an ApprovalDecision. |
| Confidence | The RCA model's self-reported 0–1 certainty. Not calibrated; gated at 0.5 for automation, capped at 0.3 by the guardrail. |
| Correlation key | The grouping identity for related events (request:{id} or a legacy-UI message hash) → one incident per outage. |
| Embedding | Text → vector of numbers, so "similar meaning" becomes "nearby vectors". |
| Evidence ID | A citable identifier: the error event ID, log-{i} window entries, or error-doc-* RAG doc IDs. Whitelist-filtered after generation. |
| Fan-out / fan-in | One node branching to several parallel nodes / several branches converging on one node. Ours: intake → {retrieval ∥ rag} → rca. |
| fix_eligibility | The policy gate: no RCA, confidence < 0.5, or severity "low" → not eligible for automated fix. |
| Guardrail | Code that rewrites or demotes LLM output that violates policy (operational_rca, operational_recommendation). |
| interrupt() | LangGraph primitive: halt the graph, persist, return control; resumed later with Command(resume). |
| LangGraph | A framework for building LLM workflows as explicit state machines with checkpointing and interrupts. |
| Node | One step of the graph. Seven total: intake, retrieval, rag, rca, recommend, approval, code_fix. |
| NodeEvent | Audit record appended per node (state.events): which node, detail, timestamp. |
| operator.add reducer | The rule that concurrent writes to errors/events append instead of overwriting — why parallel branches merge safely. |
| Prompt injection | Attacker-controlled text (e.g., in logs) trying to steer the LLM. Mitigated by "untrusted evidence" prompting + validation + human gate. |
| Pydantic | Library that validates JSON against a typed schema (RCAResult.model_validate_json). Malformed output fails loudly. |
| RAG | Retrieval-Augmented Generation: fetch relevant docs, put them in the prompt, answer from them. |
| RCA | Root-Cause Analysis: the structured diagnosis (cause, factors, evidence, severity, confidence, sequence). |
| Reducer | The merge function for a state field when multiple branches write it. |
| Router | A conditional edge choosing the next node from state (route_after_approval). |
| Super-step | One execution round of a LangGraph graph; checkpointing happens per super-step. |
| Temperature | LLM randomness dial; 0.0 = most deterministic. Our choice for reproducible analysis. |
| Thread | A named checkpoint stream; ours is keyed thread_id = incident_id. |
| Tool | A narrow, validated capability the pipeline can use (Splunk query, code search, fix validation, GitHub). |
| Vector store | Database for nearest-vector search. Ours: ChromaDB, in-memory, rebuilt on start. |
20-question self-quiz
- Name the seven graph nodes in order, and mark which three call an LLM.
- What are the two parallel branches, what does each write into state, and why can't they clobber each other?
- What three conditions make an incident eligible for automated code fix?
- What does the FIX button actually send, and what does the graph do with it?
- Where does the graph pause, and what guarantees the pause survives a restart?
- What are the two SQLite databases, and which component owns each?
- Why is
max_retries=0set on the LLM client when we retry three times anyway? - What happens to an incident when its RCA comes back with confidence 0.3?
- What does
operational_rcado to a demo-flavored analysis, and why does that matter for the router? - How many documents are in the RAG corpus, how are chunks split, and what embedding model is used?
- An evidence ID appears in the RCA that was never supplied. What happens to it?
- The fix agent proposes editing a file that wasn't in the search results. What happens?
- What stops the LLM from "fixing" the incident by disabling the demo switch?
- What happens to a Splunk event whose processing fails — is it marked processed? What does the cursor do?
- How does the system avoid processing the same Splunk event twice given the 5-minute overlap window?
- What are the 10 metric counters, where are they exposed, and what happens to them on restart?
- What does the reviewer see in the approval interrupt payload, and why does that matter for trust?
- How do tests run the full graph without any live LLM, Splunk, or GitHub call?
- What is
make graph-smoke, and why is it also the demo fallback? - What is the single biggest honest weakness of this system, and what is the proposed fix?
Answers (check yourself)
- intake → {retrieval ∥ rag} → rca → recommend → approval → code_fix; LLM: rca, recommend, code_fix (
builder.py). - retrieval (
retrieved_logs) and rag (similar_error_docs); shared lists useoperator.addappend reducers (state.py). - Human approved AND confidence ≥ 0.5 AND severity ≠ "low" (
builder.py,config.py). Command(resume=ApprovalDecision(status="approved", ...))viaPOST /api/incidents/{id}/fix; the router then picks code_fix or END (api/routes/incidents.py).- At
interrupt()in approval; the SqliteSaver checkpointer persists every super-step (approval.py,checkpointer.py). - App DB
backend/app.sqlite3(domain records,db/repository.py); checkpointerbackend/checkpoints.sqlite(graph state,graph/checkpointer.py). - So one layer owns the retry budget — the tenacity wrapper (3 attempts, 1–10s backoff); nested retries would multiply attempts (
llm/client.py). - It's saved, but the router blocks automation (0.3 < 0.5); human-reviewed recommendation only (
builder.py). - Rewrites summary/root cause to "observed failure…", caps confidence at 0.3, filters factors/sequence — which makes
fix_eligibilityblock the fix (incident_presentation.py). - 10 entries; split on
\n---\n; ONNXMiniLM_L6_V2 (error_docs.txt,rag/ingest.py,rag/store.py). - The whitelist filter drops it — allowed IDs are only the supplied doc IDs,
log-{i}, and the error ID (rca.py). ValueError("The fix agent selected a file outside the searched candidates.")(issue_fix.py).- Two layers: the prompt forbids it, and
code_searchexcludes demo/config paths viaIGNORED_TERMS(issue_fix.py,code_search.py). - Not marked processed; the cursor doesn't advance, so it's retried next cycle; dedupe makes eventual reprocessing safe (
splunk_poller.py). - Stable content-derived event IDs (
SPLUNK-+ sha256) checked againstprocessed_splunk_events(splunk_normalizer.py,db/repository.py). - incidents_processed/failed, rca_success/failed, recommend_success/failed, approval_approved/rejected, fix_created/failed;
GET /api/metrics; they reset on restart (in-memory) (metrics.py). - The root cause, confidence, recommendation, and the similar docs the model saw — informed review, not rubber-stamping (
approval.py). node_overridesswaps LLM nodes at build time;SplunkClientProtocol takes fakes;MemorySaver/offline embeddings for the rest (builder.py,test_stub_flow.py,test_rca_dashboard.py).- A deterministic end-to-end run with zero external dependencies (mock Splunk, stubbed LLM, MemorySaver) — it always works, so it's the fallback when live services fail (
smoke_graph.py). - No LLM output-quality evals — correctness is guardrails + human, not measured; the fix is a golden-set eval harness with scoring (10-testing.md).