14. Challenges faced and lessons learned
If you only remember 3 things 1. Every challenge below is traceable to a commit or a code comment — nothing here is invented; that's what makes this section credible in front of an evaluator. 2. The three best stories: the Splunk schema mismatch, the recommend persistence gap, and the guardrail-vs-router test collision. 3. The meta-lesson we actually learned: prompts are suggestions, code is enforcement — every prompt rule in this system eventually got a code backstop.
These are derived from the 24-commit git history (2026-09-22 → 09-27), the alignment
review's findings (docs/rubric/capstone-alignment-review.md), and in-code comments.
No fabrication; where history is thin, we say so.
Challenge 1: The architecture pivot (commit 4ddbad2)
The project began as a LangGraph skeleton (24104c8) with a POS + Splunk telemetry
feed (c6ebd2f), then pivoted to the Splunk-polling RAG→RCA→code-fix pipeline.
Lesson: building the demo app first (POS + telemetry) gave the pipeline a real
event source to design against — the pipeline's shape (poller, correlation,
dedupe) came from the shape of the data, not from a diagram drawn in advance.
Challenge 2: Splunk HEC/Search schema mismatch (commit 3213f66)
Events forwarded via HEC didn't match what search queries returned — a field/schema
mismatch between what we sent and what we queried. The patch reconciled the two.
Lesson: integration contracts live in both systems; a "working" forwarder and a
"working" searcher can still disagree. This is why SplunkClient became a Protocol
with fakes — so tests encode the contract we actually verified.
Challenge 3: Correlating logs to an order (commit bd6a03e)
Log lines lacked a shared transaction identifier, so the log window for an incident
couldn't be scoped to the failed order. The fix added order-number context to the
logged fields, which fetch_log_window now filters on (txn_id OR orderNumber OR
requestId, sessionId scoped to exclude other orders —
backend/app/tools/splunk_query.py).
Lesson: retrieval quality is decided at logging time, not query time. The best
SPL in the world can't correlate logs that never carried an ID.
Challenge 4: Parallel fan-out, and when not to use Send() (commit 9f65393)
Adding the intake → {retrieval ∥ rag} → rca fan-out raised the question of LangGraph's
Send() API. The commit message records the decision: "No Send() needed since it's a
static fan-out, not per-item map-reduce." Static edges express the topology
declaratively — visible in builder.py, testable with edge assertions.
Lesson: use the least powerful primitive that expresses the design. Send() is
for dynamic per-item dispatch; we didn't have that, and pretending we did would have
made the graph harder to test and reason about.
Challenge 5: The guardrail-vs-router test collision (commit 098a6df)
When the fix_eligibility router landed, a test that stubbed a demo-flavored RCA
started failing — because operational_rca legitimately demoted it to confidence 0.3,
which the router then (correctly) blocked. The test had to account for the guardrail's
behavior. Lesson: defense layers interact. A guardrail that rewrites confidence
is also a policy input to the router — tests must exercise the composed system, not
each layer in isolation. (This is also a good demo story: the safety system caught the
test's own fake analysis.)
Challenge 6: The recommend persistence gap (commit bda6c87)
After splitting recommendation into its own node, the dashboard stopped showing
completed recommendations: recommend updated in-memory state, but the persisted RCA
row was written before the node ran. The fix: recommend re-saves the RCA via
save_rca after updating recommended_action (the comment in
backend/app/graph/nodes/recommend.py documents this).
Lesson: graph state and domain persistence are two stores that must be kept in
sync explicitly — LangGraph checkpoints the workflow, but the API reads the domain DB.
Any node that mutates a field another surface reads must persist the mutation itself.
Challenge 7: Observability was an afterthought — until the review (commit 2e6d4b3)
The alignment review flagged that structlog was declared but unused and there was no
metrics endpoint. The response: structlog JSON logging wired at import
(backend/app/monitoring/logging.py), a thread-safe counter with 10 named counters,
and GET /api/metrics (backend/app/api/routes/metrics.py).
Lesson: "robustness" is a rubric category and an engineering reality — you
can't argue a system is production-shaped without operational telemetry. But note the
honest limit: the counters are in-memory; durability is the next gap
(09-robustness-observability.md).
Challenge 8: The LLM provider switch (commit 200c078)
The client was switched to the team's Codex key / OpenAI-compatible endpoint early on.
Because all access goes through get_chat_model in backend/app/llm/client.py, the
switch was a config change, not a refactor. Lesson: one seam for external
dependencies pays for itself the first time the dependency changes.
The meta-lessons (say these in the 3-minute challenges segment)
- Prompts are suggestions; code is enforcement. Every prompt rule (cite only
supplied IDs, don't touch demo flags, return JSON) eventually got a code backstop
(whitelist filter,
IGNORED_TERMS, Pydantic validation). The prompt reduces the error rate; the code guarantees the invariant. - Testability seams are design decisions, not test utilities.
node_overridesand theSplunkClientProtocol made 103 tests possible without a single live API call — and made the smoke demo deterministic. - Docs drift; code doesn't. We found 7 doc-vs-code discrepancies. The handbook you're reading cites files, not documents, for that reason.
- Controlled demos need guardrails against themselves. Because we inject
errors, we had to build
operational_rcato stop the model from explaining incidents via our own scaffolding — an unusual but real prompt-injection-adjacent problem we solved in code (backend/app/services/incident_presentation.py).