10. Testing and verification
If you only remember 3 things 1. 103 tests pass, 2 skipped (verified:
uv run pytest -q), covering graph topology, router policy, HITL resume, poller semantics, metrics, redaction, and an offline end-to-end dashboard flow — zero live LLM/Splunk/GitHub calls. 2. The two testability seams arenode_overrides(swap any node with a stub at build time —backend/app/graph/builder.py) andSplunkClientas a Protocol (inject fakes —backend/app/tools/splunk_query.py). 3. What is NOT tested: LLM output quality. There are no evals measuring whether the RCA is correct — only that it's well-formed. We have a minimal proposal below; say this gap out loud before the evaluator finds it.
Layout
backend/tests/
├── unit/
│ ├── graph/ test_builder.py, test_state.py, test_recommend.py,
│ │ test_approval_metrics.py, test_code_fix_metrics.py
│ ├── services/ test_splunk_poller.py, test_splunk_normalizer.py,
│ │ test_splunk_connection.py, test_incident_grouping.py
│ ├── monitoring/ test_metrics.py, test_metrics_endpoint.py
│ └── ... test_config.py, test_pos.py, test_telemetry.py
└── integration/
├── test_stub_flow.py # full graph with stubbed LLM nodes
├── test_rca_dashboard.py # offline e2e: ingest→retrieve→RCA parse→persist→APIs
└── test_application_lifecycle.py
Run: make backend-test (→ uv run pytest), or uv run pytest -q from backend/.
What is covered (with the test that proves it)
| Behavior | Test file | How |
|---|---|---|
| 7-node topology, fan-out/fan-in edges | unit/graph/test_builder.py |
explicit edge assertions |
| Parallel branches actually execute | unit/graph/test_builder.py |
spy on both branch nodes in one run |
fix_eligibility policy (4 cases) |
unit/graph/test_builder.py |
TestFixEligibility |
route_after_approval (4 cases) |
unit/graph/test_builder.py |
TestRouteAfterApproval |
| Low confidence blocks code_fix (e2e) | unit/graph/test_builder.py |
full run with stubbed RCA |
node_overrides rejects unknown names |
unit/graph/test_builder.py |
ValueError |
| Interrupt payload + approve/reject resume | integration/test_stub_flow.py |
real graph, stub LLM nodes |
| Pre-seeded auto-approval path | integration/test_stub_flow.py |
state seeded before run |
| Poller: dedupe, grouping, cursor-on-retry, checkpoint recovery, cancellation | unit/services/test_splunk_poller.py |
FakeGraph + FakeSplunkClient |
| Normalizer severity mapping + stable IDs | unit/services/test_splunk_normalizer.py |
table-driven |
| Recommend: skip w/o RCA (no LLM call), persistence, demo scrub, failure metrics | unit/graph/test_recommend.py |
stubbed model |
| Offline e2e: ingestion → retrieval → RCA parsing → persistence → read APIs | integration/test_rca_dashboard.py |
OfflineEmbedding (deterministic vectors) |
| Metrics counters + endpoint | unit/monitoring/test_metrics*.py |
direct + API |
| Telemetry redaction (13 cases) | unit/test_telemetry.py |
SENSITIVE keys never forwarded |
| POS behavior (8) | unit/test_pos.py |
API-level |
| Config (5), Splunk connection (4), incident grouping (6), lifecycle (3) | respective files | — |
How the seams work
node_overrides (backend/app/graph/builder.py)
build_graph(..., node_overrides={...}) replaces named nodes with your callables
before wiring edges; unknown names raise ValueError (itself tested). This is how
test_stub_flow.py runs the real graph, real router, real checkpointer with fake
LLM nodes — the orchestration logic is tested without any API key.
SplunkClient Protocol (backend/app/tools/splunk_query.py)
get_splunk_client() / set_splunk_client() let tests and the smoke script inject a
fake poller client. The poller depends on the Protocol, never the HTTP class.
OfflineEmbedding (integration/test_rca_dashboard.py)
A deterministic embedding function (vectors keyed on "checkout"/"payment" keywords) lets the real Chroma store and retrieval run offline with predictable rankings.
make graph-smoke (backend/scripts/smoke_graph.py)
A zero-dependency end-to-end run: MemorySaver checkpointer, mock Splunk client,
node_overrides for deterministic LLM outputs, a seeded CO-500-001 event. Verified
output (run 2026-09-27): "ingested 10 error-doc chunk(s)", "nodes run: 7", "fix
result: fix/demo-001 (created)". The script defaults to --decision approve
(so bare make graph-smoke exercises the fix branch); the Makefile help suggests
ARGS="--decision reject" to exercise rejection. Use this in the demo if
anything live fails.
What is NOT tested
- LLM output quality — no eval asserts an RCA is correct, only that it parses and passes guardrails. No golden set, no regression detection on prompt changes.
- Live integrations — Splunk REST, OpenAI API, GitHub push are all faked/stubbed in tests. (Live behavior is UNVERIFIED.)
- Load/concurrency — no test drives parallel incidents through the graph.
- Frontend — no component tests (
Incidents.jsxis verified by the offline integration test at the API level only).
Minimal eval proposal (what we'd add in a week)
- Golden incident set: 10–15 seeded incidents (the three demo errors plus variants) with human-written expected RCAs (root cause, severity, key evidence IDs).
- LLM-judge + exact-match hybrid: for each golden incident, run the real graph
with the real LLM; score (a)
root_cause/impacted_componentexact-or-synonym match, (b) evidence-ID precision (cited IDs ⊆ expected IDs), (c) severity match, (d) confidence calibration (high-confidence runs should be the correct ones). - Retrieval eval: hit-rate@5 for the golden queries against the 10-doc corpus.
- Regression gate: run in CI with a cheap model; alert if scores drop on prompt changes. This turns "we think the prompt is good" into "we measured it".