11. Rubric-to-code map
If you only remember 3 things 1. Every rubric line has a file path and a demo moment — this table is your cheat sheet for "show, don't tell". 2. Our honest self-scores: Technical ~20/25, GenAI+Agents ~16/20, Workflow ~13/15, Robustness ~11/15, Presentation = yours to win. The gaps are known and named. 3. The single biggest scoring risk is Presentation (25 marks): the system is stronger than a nervous demo makes it look. Rehearse 12-demo-playbook.md.
Self-scores are calibrated against docs/rubric/capstone-alignment-review.md (which
self-estimates Tech 20–21/25, GenAI 15–16/20, Workflow 10–11/15 pre-fix, Robustness
9–10/15 pre-fix) plus the post-fix workflow/robustness work (parallel fan-out commit
9f65393, router commit 098a6df, structlog/metrics commit 2e6d4b3).
Technical Implementation (25)
| Line item | Implementation | Demo evidence | Self-score |
|---|---|---|---|
| Working end-to-end pipeline | backend/app/main.py lifespan → poller → graph → API |
Live POS error → dashboard RCA | Strong |
| API design | backend/app/api/routes/{incidents,metrics,pos,telemetry}.py |
Show /api/incidents, /api/metrics |
Strong |
| Data persistence | backend/app/db/repository.py (5 tables), backend/app/graph/checkpointer.py |
Show both SQLite files' roles | Strong |
| Integration (Splunk) | backend/app/tools/splunk_query.py (Protocol + REST client) |
Show poller status transitions | Strong |
| Code quality | ruff (line-length 100), typed, 103 tests | make backend-test |
Strong |
| RAG depth | 10-doc corpus, default embedding, in-memory Chroma | — | Weak: ~4/5 lost here |
Honest total: ~20/25. The deduction is RAG depth (thin corpus, no hybrid search, no evals) — see 05-rag-pipeline.md.
GenAI + Agents (20)
| Line item | Implementation | Demo evidence | Self-score |
|---|---|---|---|
| LLM integration | backend/app/llm/client.py (temp 0, retries) |
Show retry config | Strong |
| Prompt engineering | rca.py, recommend.py, issue_fix.py system prompts |
Quote grounding rules | Strong |
| Structured output + validation | Pydantic RCAResult etc., strip_code_fences |
Show a parsed RCA JSON | Strong |
| Guardrails / anti-hallucination | incident_presentation.py (confidence cap 0.3), evidence whitelist |
Show a demoted demo-flavored RCA | Strong |
| Multi-agent orchestration | 7-node LangGraph, 3 LLM nodes, router | Show graph diagram | Strong |
| RAG quality | as above | — | Weak |
Honest total: ~16/20.
Workflow Design (15)
| Line item | Implementation | Demo evidence | Self-score |
|---|---|---|---|
| Sequential flow | rca → recommend → approval edges, builder.py |
Walk the happy path | Full |
| Parallel flow | intake → {retrieval ∥ rag} → rca, builder.py |
Show fan-out/fan-in + parallel test | Full |
| Conditional routing | route_after_approval + fix_eligibility, builder.py |
Show low-confidence → END branch | Full |
| Human-in-the-loop | interrupt() + Command(resume), approval.py, incidents.py |
Click FIX, show resume | Full |
Honest total: ~13/15 (post-fix; the alignment review's 10–11 predates the router and parallel work). Remaining risk: explaining why parallel is safe (reducers) — rehearse it.
Robustness (15)
| Line item | Implementation | Demo evidence | Self-score |
|---|---|---|---|
| Retry/backoff | with_retries + max_retries=0 rationale, llm/client.py |
Show the comment | Full |
| Failure handling | graceful code_fix failure, poller retry semantics |
Kill Splunk mid-demo? (risky — show test instead) | Full |
| Observability | structlog JSON, /api/metrics, NodeEvent trail |
curl /api/metrics |
Partial |
| Durable state | SqliteSaver, restart-resume test | Show checkpoint-recovery test | Full |
| Monitoring depth | no latency, in-memory counters | — | Weak: ~4/5 lost here |
Honest total: ~11/15.
Presentation / Demo / Q&A (25)
| Line item | Risk | Mitigation |
|---|---|---|
| Problem clarity (3 min) | rambling | Scripted in 01 exec summary |
| Architecture explanation (5 min) | too deep, too fast | 02-architecture.md diagram + 3-things boxes |
| Live demo (7 min) | highest risk — Splunk/LLM/network failure | 12-demo-playbook.md fallback ladder |
| Q&A (7 min) | hostile gotchas | 13-qa-bank.md — 60+ rehearsed answers |
| Equal participation (5 marks) | one person hogs | Speaking split table in the playbook |
This category is winnable and losable on preparation alone.
The 5 things to show live (if you show nothing else)
make graph-smokeoutput — the whole graph, deterministically, in seconds.- POS error → dashboard RCA appearing (the "it works end to end" moment).
- The FIX click → branch created (the HITL + agent moment).
curl /api/metrics— counters moved.- The low-confidence test — "the system refuses to automate when unsure" (run
uv run pytest backend/tests/unit/graph/test_builder.py -k "low_confidence" -q).