15. Production roadmap

If you only remember 3 things 1. The POC choices that were features (in-memory Chroma, SQLite, in-process poller, URL tokens) become liabilities at enterprise scale — each has a named replacement below. 2. Priority order: evals and tracing first (you can't improve what you don't measure), then durability/scaling, then hardening (RBAC, secrets, guardrails). 3. Every item maps to an existing seam in the code — this roadmap is evolution, not rewrite. That's the sentence that wins architecture Q&A.

Tier 1 — Measure before you scale

Item Why first Seam today
LLM output-quality evals RCA correctness is currently guardrails + human; a golden-set eval harness turns prompt tuning into measurement node_overrides + smoke_graph.py already run deterministic pipelines; add scoring (10-testing.md proposal)
Tracing (LangSmith or OpenTelemetry) Per-node spans, token counts, latency, and prompt-version regression in one place structlog JSON is already machine-parseable (monitoring/logging.py); LangGraph natively supports tracing callbacks
Latency metrics No latency data exists at all; add per-node timers/histograms NodeEvent already timestamps every node (graph/state.py) — compute durations from it
Durable metrics Counters reset on restart; export to Prometheus/VictoriaMetrics metrics.py counter abstraction isolates the backend swap

Tier 2 — Durability and scale

Item POC choice Production choice
Checkpointer SqliteSaver → backend/checkpoints.sqlite (graph/checkpointer.py) Postgres checkpointer; same interface
Domain DB sqlite3 → backend/app.sqlite3 (db/repository.py) Postgres + a real migration tool (repository is already plain SQL, no ORM lock-in)
Vector store Chroma EphemeralClient, rebuilt per process (rag/store.py) Persistent Chroma or a managed vector service; shared across workers
Ingestion trigger In-process poller, 2s interval (services/splunk_poller.py) Queue-backed workers (Splunk → Kafka/webhook → queue); poller becomes a consumer; at-least-once semantics already designed in
Concurrency Sequential per poll cycle; thread-isolated by thread_id Load-test concurrent incidents; per-tenant sharding
RAG quality Pure vector, top-5, no threshold (graph/nodes/rag.py) Hybrid BM25+vector, reranking, similarity cutoff, retrieval evals in CI

Tier 3 — Hardening

Item Gap today Production design
RBAC on the FIX action Any dashboard user can approve (api/routes/incidents.py accepts reviewer as a string) AuthN/AuthZ (SSO/OIDC), role-gated approval, four-eyes on high-severity, immutable approval audit with real identity
Secret management GITHUB_TOKEN injected into remote URL (tools/github_tool.py) — can persist in git config Credential helper / deploy key / GitHub App installation tokens; secrets from a vault (Vault/AWS SM), never in URLs or .env
PII redaction at ingestion Only our own telemetry redacts (telemetry.py redact()); arbitrary app logs pass through Redaction/PII scrubbing layer before any log text reaches a prompt; per-tenant data policies
Automated fix review Diff is human-reviewed only LLM-as-reviewer + static analysis + run the repo's tests on the fix branch before the PR is even opened
Guardrail generalization operational_rca regexes target our demo vocabulary (incident_presentation.py) Configurable scrub rules per deployment; claim-level citation checking (verify cited docs support their sentences)
Model routing One model for all three LLM nodes (llm/client.py) Small/cheap model for recommend, stronger for rca/fix; fallback chain on provider outage
Cost controls None Per-tenant token budgets, caching of identical RCA inputs (same correlation key → same incident, already deduped)
Multi-tenancy Single index/repo/thresholds (config.py) Per-tenant config: Splunk index, target repo, confidence thresholds, approval policy

What we would NOT change

The 15-year-architect one-liner

"This POC made the right trade-offs: it spent its complexity budget on the agentic core and safety boundaries, and used boring, replaceable infrastructure (SQLite, in-memory Chroma, a poller) behind interfaces. Production is a sequence of swaps at existing seams — Postgres, a queue, a real vector service, RBAC, and above all evals and tracing — not a redesign."