9. Robustness and observability

If you only remember 3 things 1. Retry lives in one place: with_retries (tenacity, 3 attempts, exponential backoff 1–10s) wraps every LLM call, and the client is built with max_retries=0 on purpose — the outer wrapper owns the retry budget (backend/app/llm/client.py). 2. The poller is at-least-once with a safe cursor: a failed event is not marked processed, so the cursor doesn't advance and the event retries next cycle; dedupe by stable event ID makes the reprocessing harmless (backend/app/services/splunk_poller.py). 3. Observability is real but minimal: structlog JSON logs, a thread-safe metrics counter with /api/metrics, and a per-run NodeEvent audit trail — all in-memory; metrics reset on restart and there are no latency metrics. Say this before the evaluator does.

Retry and backoff (backend/app/llm/client.py)

model = ChatOpenAI(model=..., temperature=0.0, timeout=60, max_retries=0)
# "The outer retry wrapper owns the retry budget" (comment in file)
@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=1, max=10),
       reraise=True)
def with_retries(...)

Why max_retries=0 on the client: if both the SDK and our wrapper retried, a single failure could multiply into 3×3 = 9 attempts with uncoordinated backoff — slower, noisier, and impossible to reason about. One owner, one budget: the wrapper gives us 3 attempts with exponential waits (1s, 2s, 4s… capped at 10s), and reraise=True surfaces the final exception to the node, which increments rca_failed/recommend_failed and lets the poller's retry semantics take over.

Poller retry semantics (backend/app/services/splunk_poller.py)

Graceful failure in code_fix (backend/app/graph/nodes/code_fix.py)

A fix failure never fails the incident: the node catches exceptions and writes FixResult(status="failed", summary=<error>) into state, increments fix_failed, and the graph ends normally. The incident, RCA, and approval are all already persisted — a failed fix is a recorded outcome, not a lost run. (Verified by backend/tests/unit/graph/test_code_fix_metrics.py.)

Best-effort telemetry (backend/app/telemetry.py)

The POS → Splunk forwarder is deliberately best-effort: a bounded queue (maxsize=500), a daemon sender thread with a 2s timeout, and rate-limited failure logging (one warning per 30s) so a Splunk outage can't crash the POS or flood the logs. Why: telemetry must never take down the app it observes. The trade-off is possible log loss under outage — acceptable, and the poller's overlap window (earliest = since - 5min) gives a recovery grace period.

Structured logging (backend/app/monitoring/logging.py)

Metrics (backend/app/monitoring/metrics.py, backend/app/api/routes/metrics.py)

A thread-safe Counter (a dict + threading.Lock) with 10 named counters:

incidents_processed, incidents_failed, rca_success, rca_failed, recommend_success, recommend_failed, approval_approved, approval_rejected, fix_created, fix_failed.

Exposed at GET /api/metrics as {"counters": {...}, "pipeline": {...}}. Incremented at every node outcome (e.g., rca.py increments rca_success/rca_failed around the LLM call). Tested in backend/tests/unit/monitoring/test_metrics.py and test_metrics_endpoint.py.

Per-run audit trail (backend/app/graph/state.py)

Every node appends a NodeEvent(node, detail, timestamp) via make_event() to state.events (an operator.add list). The final state therefore carries a complete, ordered record of which node ran, when, and how it ended — visible in the smoke output ("nodes run: 7") and stored in the checkpoint. This is the forensic trail for "what did the pipeline actually do for this incident?"

Real vs. missing — the honest table

Capability Status Where / gap
JSON structured logs ✅ real monitoring/logging.py
Outcome counters + endpoint ✅ real monitoring/metrics.py, /api/metrics
Per-run node audit trail ✅ real state.events, make_event()
LLM retry with backoff ✅ real llm/client.py
Poller at-least-once + safe cursor ✅ real splunk_poller.py + tests
Graceful fix failure ✅ real code_fix.py + test
Metrics persistence ❌ in-memory, reset on restart metrics.py (Counter is a plain dict)
Latency metrics ❌ none no timers/histograms anywhere
Distributed tracing ❌ none no LangSmith/OTel (roadmap)
Alerting ❌ none —

The one-liner for "your metrics reset on restart, so is that monitoring?": "It's operational telemetry for the demo, not a durable monitoring system — the durable record is the per-incident audit trail in SQLite plus the NodeEvent history in each checkpoint. For production we'd export counters to Prometheus and add latency histograms; the counter abstraction already isolates that change."