10. Testing and verification

If you only remember 3 things 1. 103 tests pass, 2 skipped (verified: uv run pytest -q), covering graph topology, router policy, HITL resume, poller semantics, metrics, redaction, and an offline end-to-end dashboard flow — zero live LLM/Splunk/GitHub calls. 2. The two testability seams are node_overrides (swap any node with a stub at build time — backend/app/graph/builder.py) and SplunkClient as a Protocol (inject fakes — backend/app/tools/splunk_query.py). 3. What is NOT tested: LLM output quality. There are no evals measuring whether the RCA is correct — only that it's well-formed. We have a minimal proposal below; say this gap out loud before the evaluator finds it.

Layout

backend/tests/
├── unit/
│   ├── graph/        test_builder.py, test_state.py, test_recommend.py,
│   │                 test_approval_metrics.py, test_code_fix_metrics.py
│   ├── services/     test_splunk_poller.py, test_splunk_normalizer.py,
│   │                 test_splunk_connection.py, test_incident_grouping.py
│   ├── monitoring/    test_metrics.py, test_metrics_endpoint.py
│   └── ...           test_config.py, test_pos.py, test_telemetry.py
└── integration/
    ├── test_stub_flow.py          # full graph with stubbed LLM nodes
    ├── test_rca_dashboard.py     # offline e2e: ingest→retrieve→RCA parse→persist→APIs
    └── test_application_lifecycle.py

Run: make backend-test (→ uv run pytest), or uv run pytest -q from backend/.

What is covered (with the test that proves it)

Behavior Test file How
7-node topology, fan-out/fan-in edges unit/graph/test_builder.py explicit edge assertions
Parallel branches actually execute unit/graph/test_builder.py spy on both branch nodes in one run
fix_eligibility policy (4 cases) unit/graph/test_builder.py TestFixEligibility
route_after_approval (4 cases) unit/graph/test_builder.py TestRouteAfterApproval
Low confidence blocks code_fix (e2e) unit/graph/test_builder.py full run with stubbed RCA
node_overrides rejects unknown names unit/graph/test_builder.py ValueError
Interrupt payload + approve/reject resume integration/test_stub_flow.py real graph, stub LLM nodes
Pre-seeded auto-approval path integration/test_stub_flow.py state seeded before run
Poller: dedupe, grouping, cursor-on-retry, checkpoint recovery, cancellation unit/services/test_splunk_poller.py FakeGraph + FakeSplunkClient
Normalizer severity mapping + stable IDs unit/services/test_splunk_normalizer.py table-driven
Recommend: skip w/o RCA (no LLM call), persistence, demo scrub, failure metrics unit/graph/test_recommend.py stubbed model
Offline e2e: ingestion → retrieval → RCA parsing → persistence → read APIs integration/test_rca_dashboard.py OfflineEmbedding (deterministic vectors)
Metrics counters + endpoint unit/monitoring/test_metrics*.py direct + API
Telemetry redaction (13 cases) unit/test_telemetry.py SENSITIVE keys never forwarded
POS behavior (8) unit/test_pos.py API-level
Config (5), Splunk connection (4), incident grouping (6), lifecycle (3) respective files —

How the seams work

node_overrides (backend/app/graph/builder.py)

build_graph(..., node_overrides={...}) replaces named nodes with your callables before wiring edges; unknown names raise ValueError (itself tested). This is how test_stub_flow.py runs the real graph, real router, real checkpointer with fake LLM nodes — the orchestration logic is tested without any API key.

SplunkClient Protocol (backend/app/tools/splunk_query.py)

get_splunk_client() / set_splunk_client() let tests and the smoke script inject a fake poller client. The poller depends on the Protocol, never the HTTP class.

OfflineEmbedding (integration/test_rca_dashboard.py)

A deterministic embedding function (vectors keyed on "checkout"/"payment" keywords) lets the real Chroma store and retrieval run offline with predictable rankings.

make graph-smoke (backend/scripts/smoke_graph.py)

A zero-dependency end-to-end run: MemorySaver checkpointer, mock Splunk client, node_overrides for deterministic LLM outputs, a seeded CO-500-001 event. Verified output (run 2026-09-27): "ingested 10 error-doc chunk(s)", "nodes run: 7", "fix result: fix/demo-001 (created)". The script defaults to --decision approve (so bare make graph-smoke exercises the fix branch); the Makefile help suggests ARGS="--decision reject" to exercise rejection. Use this in the demo if anything live fails.

What is NOT tested

  1. LLM output quality — no eval asserts an RCA is correct, only that it parses and passes guardrails. No golden set, no regression detection on prompt changes.
  2. Live integrations — Splunk REST, OpenAI API, GitHub push are all faked/stubbed in tests. (Live behavior is UNVERIFIED.)
  3. Load/concurrency — no test drives parallel incidents through the graph.
  4. Frontend — no component tests (Incidents.jsx is verified by the offline integration test at the API level only).

Minimal eval proposal (what we'd add in a week)

  1. Golden incident set: 10–15 seeded incidents (the three demo errors plus variants) with human-written expected RCAs (root cause, severity, key evidence IDs).
  2. LLM-judge + exact-match hybrid: for each golden incident, run the real graph with the real LLM; score (a) root_cause/impacted_component exact-or-synonym match, (b) evidence-ID precision (cited IDs ⊆ expected IDs), (c) severity match, (d) confidence calibration (high-confidence runs should be the correct ones).
  3. Retrieval eval: hit-rate@5 for the golden queries against the 10-doc corpus.
  4. Regression gate: run in CI with a cheap model; alert if scores drop on prompt changes. This turns "we think the prompt is good" into "we measured it".