14. Challenges faced and lessons learned

If you only remember 3 things 1. Every challenge below is traceable to a commit or a code comment — nothing here is invented; that's what makes this section credible in front of an evaluator. 2. The three best stories: the Splunk schema mismatch, the recommend persistence gap, and the guardrail-vs-router test collision. 3. The meta-lesson we actually learned: prompts are suggestions, code is enforcement — every prompt rule in this system eventually got a code backstop.

These are derived from the 24-commit git history (2026-09-22 → 09-27), the alignment review's findings (docs/rubric/capstone-alignment-review.md), and in-code comments. No fabrication; where history is thin, we say so.

Challenge 1: The architecture pivot (commit 4ddbad2)

The project began as a LangGraph skeleton (24104c8) with a POS + Splunk telemetry feed (c6ebd2f), then pivoted to the Splunk-polling RAG→RCA→code-fix pipeline. Lesson: building the demo app first (POS + telemetry) gave the pipeline a real event source to design against — the pipeline's shape (poller, correlation, dedupe) came from the shape of the data, not from a diagram drawn in advance.

Challenge 2: Splunk HEC/Search schema mismatch (commit 3213f66)

Events forwarded via HEC didn't match what search queries returned — a field/schema mismatch between what we sent and what we queried. The patch reconciled the two. Lesson: integration contracts live in both systems; a "working" forwarder and a "working" searcher can still disagree. This is why SplunkClient became a Protocol with fakes — so tests encode the contract we actually verified.

Challenge 3: Correlating logs to an order (commit bd6a03e)

Log lines lacked a shared transaction identifier, so the log window for an incident couldn't be scoped to the failed order. The fix added order-number context to the logged fields, which fetch_log_window now filters on (txn_id OR orderNumber OR requestId, sessionId scoped to exclude other orders — backend/app/tools/splunk_query.py). Lesson: retrieval quality is decided at logging time, not query time. The best SPL in the world can't correlate logs that never carried an ID.

Challenge 4: Parallel fan-out, and when not to use Send() (commit 9f65393)

Adding the intake → {retrieval ∥ rag} → rca fan-out raised the question of LangGraph's Send() API. The commit message records the decision: "No Send() needed since it's a static fan-out, not per-item map-reduce." Static edges express the topology declaratively — visible in builder.py, testable with edge assertions. Lesson: use the least powerful primitive that expresses the design. Send() is for dynamic per-item dispatch; we didn't have that, and pretending we did would have made the graph harder to test and reason about.

Challenge 5: The guardrail-vs-router test collision (commit 098a6df)

When the fix_eligibility router landed, a test that stubbed a demo-flavored RCA started failing — because operational_rca legitimately demoted it to confidence 0.3, which the router then (correctly) blocked. The test had to account for the guardrail's behavior. Lesson: defense layers interact. A guardrail that rewrites confidence is also a policy input to the router — tests must exercise the composed system, not each layer in isolation. (This is also a good demo story: the safety system caught the test's own fake analysis.)

Challenge 6: The recommend persistence gap (commit bda6c87)

After splitting recommendation into its own node, the dashboard stopped showing completed recommendations: recommend updated in-memory state, but the persisted RCA row was written before the node ran. The fix: recommend re-saves the RCA via save_rca after updating recommended_action (the comment in backend/app/graph/nodes/recommend.py documents this). Lesson: graph state and domain persistence are two stores that must be kept in sync explicitly — LangGraph checkpoints the workflow, but the API reads the domain DB. Any node that mutates a field another surface reads must persist the mutation itself.

Challenge 7: Observability was an afterthought — until the review (commit 2e6d4b3)

The alignment review flagged that structlog was declared but unused and there was no metrics endpoint. The response: structlog JSON logging wired at import (backend/app/monitoring/logging.py), a thread-safe counter with 10 named counters, and GET /api/metrics (backend/app/api/routes/metrics.py). Lesson: "robustness" is a rubric category and an engineering reality — you can't argue a system is production-shaped without operational telemetry. But note the honest limit: the counters are in-memory; durability is the next gap (09-robustness-observability.md).

Challenge 8: The LLM provider switch (commit 200c078)

The client was switched to the team's Codex key / OpenAI-compatible endpoint early on. Because all access goes through get_chat_model in backend/app/llm/client.py, the switch was a config change, not a refactor. Lesson: one seam for external dependencies pays for itself the first time the dependency changes.

The meta-lessons (say these in the 3-minute challenges segment)

  1. Prompts are suggestions; code is enforcement. Every prompt rule (cite only supplied IDs, don't touch demo flags, return JSON) eventually got a code backstop (whitelist filter, IGNORED_TERMS, Pydantic validation). The prompt reduces the error rate; the code guarantees the invariant.
  2. Testability seams are design decisions, not test utilities. node_overrides and the SplunkClient Protocol made 103 tests possible without a single live API call — and made the smoke demo deterministic.
  3. Docs drift; code doesn't. We found 7 doc-vs-code discrepancies. The handbook you're reading cites files, not documents, for that reason.
  4. Controlled demos need guardrails against themselves. Because we inject errors, we had to build operational_rca to stop the model from explaining incidents via our own scaffolding — an unusual but real prompt-injection-adjacent problem we solved in code (backend/app/services/incident_presentation.py).