9. Robustness and observability
If you only remember 3 things 1. Retry lives in one place:
with_retries(tenacity, 3 attempts, exponential backoff 1–10s) wraps every LLM call, and the client is built withmax_retries=0on purpose — the outer wrapper owns the retry budget (backend/app/llm/client.py). 2. The poller is at-least-once with a safe cursor: a failed event is not marked processed, so the cursor doesn't advance and the event retries next cycle; dedupe by stable event ID makes the reprocessing harmless (backend/app/services/splunk_poller.py). 3. Observability is real but minimal: structlog JSON logs, a thread-safe metrics counter with/api/metrics, and a per-runNodeEventaudit trail — all in-memory; metrics reset on restart and there are no latency metrics. Say this before the evaluator does.
Retry and backoff (backend/app/llm/client.py)
model = ChatOpenAI(model=..., temperature=0.0, timeout=60, max_retries=0)
# "The outer retry wrapper owns the retry budget" (comment in file)
@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=1, max=10),
reraise=True)
def with_retries(...)
Why max_retries=0 on the client: if both the SDK and our wrapper retried, a
single failure could multiply into 3×3 = 9 attempts with uncoordinated backoff —
slower, noisier, and impossible to reason about. One owner, one budget: the wrapper
gives us 3 attempts with exponential waits (1s, 2s, 4s… capped at 10s), and
reraise=True surfaces the final exception to the node, which increments
rca_failed/recommend_failed and lets the poller's retry semantics take over.
Poller retry semantics (backend/app/services/splunk_poller.py)
poll_oncefetches rows, groups them by correlation key, and processes the last event per group (one incident per outage, not one per log line).- Per event: normalize → dedupe (
is_event_processed) → correlate → run/resume graph → only on successmark_event_processed→ update analysis status. - On failure: the event is not marked processed, so
since(the cursor) does not advance past it — next cycle re-fetches (with the 5-minute overlap window) and retries. Tested explicitly: "cursor keeps since on retry" inbackend/tests/unit/services/test_splunk_poller.py. - Malformed rows (e.g., missing message →
ValueErrorin the normalizer) are skipped, not fatal. - Pipeline status machine:
starting → polling → processing_errors | configuration_required | connection_error— surfaced at/api/pipeline/statusand/api/health(backend/app/main.py).
Graceful failure in code_fix (backend/app/graph/nodes/code_fix.py)
A fix failure never fails the incident: the node catches exceptions and writes
FixResult(status="failed", summary=<error>) into state, increments fix_failed, and
the graph ends normally. The incident, RCA, and approval are all already persisted —
a failed fix is a recorded outcome, not a lost run. (Verified by
backend/tests/unit/graph/test_code_fix_metrics.py.)
Best-effort telemetry (backend/app/telemetry.py)
The POS → Splunk forwarder is deliberately best-effort: a bounded queue
(maxsize=500), a daemon sender thread with a 2s timeout, and rate-limited failure
logging (one warning per 30s) so a Splunk outage can't crash the POS or flood the
logs. Why: telemetry must never take down the app it observes. The trade-off is
possible log loss under outage — acceptable, and the poller's overlap window
(earliest = since - 5min) gives a recovery grace period.
Structured logging (backend/app/monitoring/logging.py)
- structlog with
ProcessorFormatteremitting JSON lines — machine-parseable, grep-able, no regex log scraping. - Config is idempotent (safe to call twice) and is invoked at import time in
backend/app/main.py. - Added in commit
2e6d4b3("structlog + metrics") — the alignment review's Robustness finding.
Metrics (backend/app/monitoring/metrics.py, backend/app/api/routes/metrics.py)
A thread-safe Counter (a dict + threading.Lock) with 10 named counters:
incidents_processed, incidents_failed, rca_success, rca_failed,
recommend_success, recommend_failed, approval_approved, approval_rejected,
fix_created, fix_failed.
Exposed at GET /api/metrics as {"counters": {...}, "pipeline": {...}}.
Incremented at every node outcome (e.g., rca.py increments rca_success/rca_failed
around the LLM call). Tested in backend/tests/unit/monitoring/test_metrics.py and
test_metrics_endpoint.py.
Per-run audit trail (backend/app/graph/state.py)
Every node appends a NodeEvent(node, detail, timestamp) via make_event()
to state.events (an operator.add list). The final state therefore carries a
complete, ordered record of which node ran, when, and how it ended — visible in
the smoke output ("nodes run: 7") and stored in the checkpoint. This is the
forensic trail for "what did the pipeline actually do for this incident?"
Real vs. missing — the honest table
| Capability | Status | Where / gap |
|---|---|---|
| JSON structured logs | ✅ real | monitoring/logging.py |
| Outcome counters + endpoint | ✅ real | monitoring/metrics.py, /api/metrics |
| Per-run node audit trail | ✅ real | state.events, make_event() |
| LLM retry with backoff | ✅ real | llm/client.py |
| Poller at-least-once + safe cursor | ✅ real | splunk_poller.py + tests |
| Graceful fix failure | ✅ real | code_fix.py + test |
| Metrics persistence | ❌ in-memory, reset on restart | metrics.py (Counter is a plain dict) |
| Latency metrics | ❌ none | no timers/histograms anywhere |
| Distributed tracing | ❌ none | no LangSmith/OTel (roadmap) |
| Alerting | ❌ none | — |
The one-liner for "your metrics reset on restart, so is that monitoring?": "It's operational telemetry for the demo, not a durable monitoring system — the durable record is the per-incident audit trail in SQLite plus the NodeEvent history in each checkpoint. For production we'd export counters to Prometheus and add latency histograms; the counter abstraction already isolates that change."