11. Rubric-to-code map

If you only remember 3 things 1. Every rubric line has a file path and a demo moment — this table is your cheat sheet for "show, don't tell". 2. Our honest self-scores: Technical ~20/25, GenAI+Agents ~16/20, Workflow ~13/15, Robustness ~11/15, Presentation = yours to win. The gaps are known and named. 3. The single biggest scoring risk is Presentation (25 marks): the system is stronger than a nervous demo makes it look. Rehearse 12-demo-playbook.md.

Self-scores are calibrated against docs/rubric/capstone-alignment-review.md (which self-estimates Tech 20–21/25, GenAI 15–16/20, Workflow 10–11/15 pre-fix, Robustness 9–10/15 pre-fix) plus the post-fix workflow/robustness work (parallel fan-out commit 9f65393, router commit 098a6df, structlog/metrics commit 2e6d4b3).

Technical Implementation (25)

Line item Implementation Demo evidence Self-score
Working end-to-end pipeline backend/app/main.py lifespan → poller → graph → API Live POS error → dashboard RCA Strong
API design backend/app/api/routes/{incidents,metrics,pos,telemetry}.py Show /api/incidents, /api/metrics Strong
Data persistence backend/app/db/repository.py (5 tables), backend/app/graph/checkpointer.py Show both SQLite files' roles Strong
Integration (Splunk) backend/app/tools/splunk_query.py (Protocol + REST client) Show poller status transitions Strong
Code quality ruff (line-length 100), typed, 103 tests make backend-test Strong
RAG depth 10-doc corpus, default embedding, in-memory Chroma — Weak: ~4/5 lost here

Honest total: ~20/25. The deduction is RAG depth (thin corpus, no hybrid search, no evals) — see 05-rag-pipeline.md.

GenAI + Agents (20)

Line item Implementation Demo evidence Self-score
LLM integration backend/app/llm/client.py (temp 0, retries) Show retry config Strong
Prompt engineering rca.py, recommend.py, issue_fix.py system prompts Quote grounding rules Strong
Structured output + validation Pydantic RCAResult etc., strip_code_fences Show a parsed RCA JSON Strong
Guardrails / anti-hallucination incident_presentation.py (confidence cap 0.3), evidence whitelist Show a demoted demo-flavored RCA Strong
Multi-agent orchestration 7-node LangGraph, 3 LLM nodes, router Show graph diagram Strong
RAG quality as above — Weak

Honest total: ~16/20.

Workflow Design (15)

Line item Implementation Demo evidence Self-score
Sequential flow rca → recommend → approval edges, builder.py Walk the happy path Full
Parallel flow intake → {retrieval ∥ rag} → rca, builder.py Show fan-out/fan-in + parallel test Full
Conditional routing route_after_approval + fix_eligibility, builder.py Show low-confidence → END branch Full
Human-in-the-loop interrupt() + Command(resume), approval.py, incidents.py Click FIX, show resume Full

Honest total: ~13/15 (post-fix; the alignment review's 10–11 predates the router and parallel work). Remaining risk: explaining why parallel is safe (reducers) — rehearse it.

Robustness (15)

Line item Implementation Demo evidence Self-score
Retry/backoff with_retries + max_retries=0 rationale, llm/client.py Show the comment Full
Failure handling graceful code_fix failure, poller retry semantics Kill Splunk mid-demo? (risky — show test instead) Full
Observability structlog JSON, /api/metrics, NodeEvent trail curl /api/metrics Partial
Durable state SqliteSaver, restart-resume test Show checkpoint-recovery test Full
Monitoring depth no latency, in-memory counters — Weak: ~4/5 lost here

Honest total: ~11/15.

Presentation / Demo / Q&A (25)

Line item Risk Mitigation
Problem clarity (3 min) rambling Scripted in 01 exec summary
Architecture explanation (5 min) too deep, too fast 02-architecture.md diagram + 3-things boxes
Live demo (7 min) highest risk — Splunk/LLM/network failure 12-demo-playbook.md fallback ladder
Q&A (7 min) hostile gotchas 13-qa-bank.md — 60+ rehearsed answers
Equal participation (5 marks) one person hogs Speaking split table in the playbook

This category is winnable and losable on preparation alone.

The 5 things to show live (if you show nothing else)

  1. make graph-smoke output — the whole graph, deterministically, in seconds.
  2. POS error → dashboard RCA appearing (the "it works end to end" moment).
  3. The FIX click → branch created (the HITL + agent moment).
  4. curl /api/metrics — counters moved.
  5. The low-confidence test — "the system refuses to automate when unsure" (run uv run pytest backend/tests/unit/graph/test_builder.py -k "low_confidence" -q).