16. Glossary and 20-question self-quiz

If you only remember 3 things 1. If a teammate can pass the 20-question quiz below cold, they can survive Q&A. 2. Every term is defined the way we use it in this codebase — not the textbook way. 3. Quiz answers are at the bottom; don't peek until you've answered out loud.

Glossary

Term Plain-English meaning (as used in IncidentIQ)
Agent An LLM step that decides something unscripted from unstructured input (diagnose, prescribe, choose a fix). Our control flow around it is deterministic on purpose.
At-least-once processing An event may be processed more than once in failure cases, but never silently dropped. Made safe by dedupe (processed_splunk_events).
Checkpoint A saved snapshot of graph state after every super-step, so a paused run survives restarts (SqliteSaver).
Chunking Splitting the knowledge file into pieces — here, one per ----separated entry.
Command(resume) The value injected into a paused graph to continue it — our FIX button sends an ApprovalDecision.
Confidence The RCA model's self-reported 0–1 certainty. Not calibrated; gated at 0.5 for automation, capped at 0.3 by the guardrail.
Correlation key The grouping identity for related events (request:{id} or a legacy-UI message hash) → one incident per outage.
Embedding Text → vector of numbers, so "similar meaning" becomes "nearby vectors".
Evidence ID A citable identifier: the error event ID, log-{i} window entries, or error-doc-* RAG doc IDs. Whitelist-filtered after generation.
Fan-out / fan-in One node branching to several parallel nodes / several branches converging on one node. Ours: intake → {retrieval ∥ rag} → rca.
fix_eligibility The policy gate: no RCA, confidence < 0.5, or severity "low" → not eligible for automated fix.
Guardrail Code that rewrites or demotes LLM output that violates policy (operational_rca, operational_recommendation).
interrupt() LangGraph primitive: halt the graph, persist, return control; resumed later with Command(resume).
LangGraph A framework for building LLM workflows as explicit state machines with checkpointing and interrupts.
Node One step of the graph. Seven total: intake, retrieval, rag, rca, recommend, approval, code_fix.
NodeEvent Audit record appended per node (state.events): which node, detail, timestamp.
operator.add reducer The rule that concurrent writes to errors/events append instead of overwriting — why parallel branches merge safely.
Prompt injection Attacker-controlled text (e.g., in logs) trying to steer the LLM. Mitigated by "untrusted evidence" prompting + validation + human gate.
Pydantic Library that validates JSON against a typed schema (RCAResult.model_validate_json). Malformed output fails loudly.
RAG Retrieval-Augmented Generation: fetch relevant docs, put them in the prompt, answer from them.
RCA Root-Cause Analysis: the structured diagnosis (cause, factors, evidence, severity, confidence, sequence).
Reducer The merge function for a state field when multiple branches write it.
Router A conditional edge choosing the next node from state (route_after_approval).
Super-step One execution round of a LangGraph graph; checkpointing happens per super-step.
Temperature LLM randomness dial; 0.0 = most deterministic. Our choice for reproducible analysis.
Thread A named checkpoint stream; ours is keyed thread_id = incident_id.
Tool A narrow, validated capability the pipeline can use (Splunk query, code search, fix validation, GitHub).
Vector store Database for nearest-vector search. Ours: ChromaDB, in-memory, rebuilt on start.

20-question self-quiz

  1. Name the seven graph nodes in order, and mark which three call an LLM.
  2. What are the two parallel branches, what does each write into state, and why can't they clobber each other?
  3. What three conditions make an incident eligible for automated code fix?
  4. What does the FIX button actually send, and what does the graph do with it?
  5. Where does the graph pause, and what guarantees the pause survives a restart?
  6. What are the two SQLite databases, and which component owns each?
  7. Why is max_retries=0 set on the LLM client when we retry three times anyway?
  8. What happens to an incident when its RCA comes back with confidence 0.3?
  9. What does operational_rca do to a demo-flavored analysis, and why does that matter for the router?
  10. How many documents are in the RAG corpus, how are chunks split, and what embedding model is used?
  11. An evidence ID appears in the RCA that was never supplied. What happens to it?
  12. The fix agent proposes editing a file that wasn't in the search results. What happens?
  13. What stops the LLM from "fixing" the incident by disabling the demo switch?
  14. What happens to a Splunk event whose processing fails — is it marked processed? What does the cursor do?
  15. How does the system avoid processing the same Splunk event twice given the 5-minute overlap window?
  16. What are the 10 metric counters, where are they exposed, and what happens to them on restart?
  17. What does the reviewer see in the approval interrupt payload, and why does that matter for trust?
  18. How do tests run the full graph without any live LLM, Splunk, or GitHub call?
  19. What is make graph-smoke, and why is it also the demo fallback?
  20. What is the single biggest honest weakness of this system, and what is the proposed fix?

Answers (check yourself)

  1. intake → {retrieval ∥ rag} → rca → recommend → approval → code_fix; LLM: rca, recommend, code_fix (builder.py).
  2. retrieval (retrieved_logs) and rag (similar_error_docs); shared lists use operator.add append reducers (state.py).
  3. Human approved AND confidence ≥ 0.5 AND severity ≠ "low" (builder.py, config.py).
  4. Command(resume=ApprovalDecision(status="approved", ...)) via POST /api/incidents/{id}/fix; the router then picks code_fix or END (api/routes/incidents.py).
  5. At interrupt() in approval; the SqliteSaver checkpointer persists every super-step (approval.py, checkpointer.py).
  6. App DB backend/app.sqlite3 (domain records, db/repository.py); checkpointer backend/checkpoints.sqlite (graph state, graph/checkpointer.py).
  7. So one layer owns the retry budget — the tenacity wrapper (3 attempts, 1–10s backoff); nested retries would multiply attempts (llm/client.py).
  8. It's saved, but the router blocks automation (0.3 < 0.5); human-reviewed recommendation only (builder.py).
  9. Rewrites summary/root cause to "observed failure…", caps confidence at 0.3, filters factors/sequence — which makes fix_eligibility block the fix (incident_presentation.py).
  10. 10 entries; split on \n---\n; ONNXMiniLM_L6_V2 (error_docs.txt, rag/ingest.py, rag/store.py).
  11. The whitelist filter drops it — allowed IDs are only the supplied doc IDs, log-{i}, and the error ID (rca.py).
  12. ValueError("The fix agent selected a file outside the searched candidates.") (issue_fix.py).
  13. Two layers: the prompt forbids it, and code_search excludes demo/config paths via IGNORED_TERMS (issue_fix.py, code_search.py).
  14. Not marked processed; the cursor doesn't advance, so it's retried next cycle; dedupe makes eventual reprocessing safe (splunk_poller.py).
  15. Stable content-derived event IDs (SPLUNK- + sha256) checked against processed_splunk_events (splunk_normalizer.py, db/repository.py).
  16. incidents_processed/failed, rca_success/failed, recommend_success/failed, approval_approved/rejected, fix_created/failed; GET /api/metrics; they reset on restart (in-memory) (metrics.py).
  17. The root cause, confidence, recommendation, and the similar docs the model saw — informed review, not rubber-stamping (approval.py).
  18. node_overrides swaps LLM nodes at build time; SplunkClient Protocol takes fakes; MemorySaver/offline embeddings for the rest (builder.py, test_stub_flow.py, test_rca_dashboard.py).
  19. A deterministic end-to-end run with zero external dependencies (mock Splunk, stubbed LLM, MemorySaver) — it always works, so it's the fallback when live services fail (smoke_graph.py).
  20. No LLM output-quality evals — correctness is guardrails + human, not measured; the fix is a golden-set eval harness with scoring (10-testing.md).