15. Production roadmap
If you only remember 3 things 1. The POC choices that were features (in-memory Chroma, SQLite, in-process poller, URL tokens) become liabilities at enterprise scale — each has a named replacement below. 2. Priority order: evals and tracing first (you can't improve what you don't measure), then durability/scaling, then hardening (RBAC, secrets, guardrails). 3. Every item maps to an existing seam in the code — this roadmap is evolution, not rewrite. That's the sentence that wins architecture Q&A.
Tier 1 — Measure before you scale
| Item | Why first | Seam today |
|---|---|---|
| LLM output-quality evals | RCA correctness is currently guardrails + human; a golden-set eval harness turns prompt tuning into measurement | node_overrides + smoke_graph.py already run deterministic pipelines; add scoring (10-testing.md proposal) |
| Tracing (LangSmith or OpenTelemetry) | Per-node spans, token counts, latency, and prompt-version regression in one place | structlog JSON is already machine-parseable (monitoring/logging.py); LangGraph natively supports tracing callbacks |
| Latency metrics | No latency data exists at all; add per-node timers/histograms | NodeEvent already timestamps every node (graph/state.py) — compute durations from it |
| Durable metrics | Counters reset on restart; export to Prometheus/VictoriaMetrics | metrics.py counter abstraction isolates the backend swap |
Tier 2 — Durability and scale
| Item | POC choice | Production choice |
|---|---|---|
| Checkpointer | SqliteSaver → backend/checkpoints.sqlite (graph/checkpointer.py) |
Postgres checkpointer; same interface |
| Domain DB | sqlite3 → backend/app.sqlite3 (db/repository.py) |
Postgres + a real migration tool (repository is already plain SQL, no ORM lock-in) |
| Vector store | Chroma EphemeralClient, rebuilt per process (rag/store.py) |
Persistent Chroma or a managed vector service; shared across workers |
| Ingestion trigger | In-process poller, 2s interval (services/splunk_poller.py) |
Queue-backed workers (Splunk → Kafka/webhook → queue); poller becomes a consumer; at-least-once semantics already designed in |
| Concurrency | Sequential per poll cycle; thread-isolated by thread_id |
Load-test concurrent incidents; per-tenant sharding |
| RAG quality | Pure vector, top-5, no threshold (graph/nodes/rag.py) |
Hybrid BM25+vector, reranking, similarity cutoff, retrieval evals in CI |
Tier 3 — Hardening
| Item | Gap today | Production design |
|---|---|---|
| RBAC on the FIX action | Any dashboard user can approve (api/routes/incidents.py accepts reviewer as a string) |
AuthN/AuthZ (SSO/OIDC), role-gated approval, four-eyes on high-severity, immutable approval audit with real identity |
| Secret management | GITHUB_TOKEN injected into remote URL (tools/github_tool.py) — can persist in git config |
Credential helper / deploy key / GitHub App installation tokens; secrets from a vault (Vault/AWS SM), never in URLs or .env |
| PII redaction at ingestion | Only our own telemetry redacts (telemetry.py redact()); arbitrary app logs pass through |
Redaction/PII scrubbing layer before any log text reaches a prompt; per-tenant data policies |
| Automated fix review | Diff is human-reviewed only | LLM-as-reviewer + static analysis + run the repo's tests on the fix branch before the PR is even opened |
| Guardrail generalization | operational_rca regexes target our demo vocabulary (incident_presentation.py) |
Configurable scrub rules per deployment; claim-level citation checking (verify cited docs support their sentences) |
| Model routing | One model for all three LLM nodes (llm/client.py) |
Small/cheap model for recommend, stronger for rca/fix; fallback chain on provider outage |
| Cost controls | None | Per-tenant token budgets, caching of identical RCA inputs (same correlation key → same incident, already deduped) |
| Multi-tenancy | Single index/repo/thresholds (config.py) |
Per-tenant config: Splunk index, target repo, confidence thresholds, approval policy |
What we would NOT change
- The graph shape — 7 nodes, static fan-out, one router, one interrupt. It maps 1:1 to the operational reality (observe → diagnose → prescribe → decide → act).
- The HITL gate placement — between prescription and action, with the eligibility router as a second, machine-side check.
- The tool boundaries — candidates-only fixes, no code execution, branch-not-main. These are the safety properties; they scale as-is.
- Code-as-enforcement — every new prompt rule ships with its validator.
The 15-year-architect one-liner
"This POC made the right trade-offs: it spent its complexity budget on the agentic core and safety boundaries, and used boring, replaceable infrastructure (SQLite, in-memory Chroma, a poller) behind interfaces. Production is a sequence of swaps at existing seams — Postgres, a queue, a real vector service, RBAC, and above all evals and tracing — not a redesign."