6. Workflow design: sequential, parallel, router
If you only remember 3 things 1. All three flow shapes live in one graph: sequential (rca → recommend → approval), parallel (intake → {retrieval ∥ rag} → rca), and conditional routing (
route_after_approval+fix_eligibility) —backend/app/graph/builder.py. 2. The router policy: approved +confidence ≥ CODE_FIX_MIN_CONFIDENCE (0.5)+severity ≠ "low"→code_fix; anything else → END. 3. The human interrupt always fires before the router — the router only decides what happens after the human decides (backend/app/graph/nodes/approval.py).
1. Sequential flow
Code (backend/app/graph/builder.py):
graph.add_edge("rca", "recommend")
graph.add_edge("recommend", "approval")
graph.add_edge("code_fix", END)
Why sequential here: recommend consumes the completed rca; approval consumes
the recommendation; code_fix consumes the approval. Each step's input is the previous
step's output — a strict dependency chain, so parallelism would be wrong.
Rubric line: "Workflow Design — sequential flow" (see 11-rubric-map.md).
2. Parallel flow (fan-out / fan-in)
Code (backend/app/graph/builder.py):
graph.add_edge(START, "intake")
graph.add_edge("intake", "retrieval") # fan-out
graph.add_edge("intake", "rag") # fan-out
graph.add_edge("retrieval", "rca") # fan-in
graph.add_edge("rag", "rca") # fan-in
Why parallel: retrieval (Splunk log window) and rag (Chroma similarity) both
depend only on error_event and not on each other. Running them concurrently cuts
their wall-clock sum to their max. Commit 9f65393 records the alternative considered:
"No Send() needed since it's a static fan-out, not per-item map-reduce."
Why the merge is safe: the two branches write different fields
(retrieved_logs vs similar_error_docs); the only shared fields (errors,
events) use operator.add append reducers (backend/app/graph/state.py). LangGraph
waits at rca until both branches complete (fan-in semantics).
Proof it actually runs in parallel: backend/tests/unit/graph/test_builder.py
contains a test that spies on both branch nodes executing in one run, and the edge
assertions check the fan-out/fan-in wiring explicitly.
Rubric line: "Workflow Design — parallel/branching flow".
3. Conditional routing (the router)
Code (backend/app/graph/builder.py):
def fix_eligibility(rca: RCAResult | None) -> tuple[bool, str]:
if rca is None:
return False, "no RCA available"
if rca.confidence < get_settings().code_fix_min_confidence:
return False, f"confidence {rca.confidence:.2f} below threshold — needs manual fix"
if rca.severity == "low":
return False, "low severity — straight-through suggestion, no fix needed"
return True, f"severity={rca.severity} confidence={rca.confidence:.2f} — eligible for auto-fix"
def route_after_approval(state: IncidentState) -> str:
if state.approval and state.approval.status == "approved":
eligible, _reason = fix_eligibility(state.rca)
if eligible:
return "code_fix"
return END
graph.add_conditional_edges("approval", route_after_approval, ["code_fix", END])
(CODE_FIX_MIN_CONFIDENCE = 0.5 — backend/app/config.py, .env.example.)
Why a router and not just "approved → fix": approval answers "did a human say yes?"; eligibility answers "is the RCA trustworthy enough to act on?" These are different questions. A human clicking FIX on a 0.2-confidence RCA shouldn't produce an automated code change. The router is the machine's second opinion on the human's decision — a policy layer between the human and the actuator.
Rubric line: "Workflow Design — conditional routing logic to direct data or tasks".
Walk-through: one incident down each router branch
Setup: cashier adds Greek Yogurt and clicks Total with
DEMO_BE_CHECKOUT_ERROR=true (backend/app/demo_errors.py, README.md). The POS
backend returns HTTP 500, logs it to Splunk via HEC (backend/app/telemetry.py), and
the poller picks it up (backend/app/services/splunk_poller.py).
The run proceeds identically for all branches: intake normalizes → retrieval+rag run in
parallel → rca produces RCAResult → recommend fills recommended_action → approval
interrupts with the review payload. Then:
Branch A — approved AND high-confidence, non-low severity
- The RCA cites the ZeroDivisionError in checkout pricing with confidence 0.9,
severity high (plausible for the seeded
CO-500family —data/knowledge/error_docs.txt,data/sample_app/checkout_service.py). - Reviewer clicks FIX →
POST /api/incidents/{id}/fix→graph.invoke(Command(resume=ApprovalDecision(status="approved", ...)))(backend/app/api/routes/incidents.py). - Router: approved ✓, confidence 0.9 ≥ 0.5 ✓, severity high ≠ low ✓ →
code_fix. code_fixsearches the sample app, proposes the guard fix, creates branchfix/inc-...-<timestamp>, pushes ifGITHUB_TOKENis set (backend/app/graph/nodes/code_fix.py,backend/app/tools/github_tool.py).- Final state:
FixResult(status="created", push_status="pushed"|"local"). - This is the
make graph-smokehappy path (verified output: "fix result: fix/demo-001 (created)", "nodes run: 7").
Branch B — approved BUT low-confidence (or low severity)
- Suppose the RCA says the internal cause cannot be determined from the logs
(the prompt explicitly allows this —
backend/app/graph/nodes/rca.py) and sets confidence 0.3. Note:operational_rcaalso caps confidence at 0.3 whenever it detects demo-flavored reasoning (backend/app/services/incident_presentation.py) — so this branch fires automatically for unproven claims. - Reviewer clicks FIX → approval = approved.
- Router: approved ✓, but
fix_eligibility→ confidence 0.3 < 0.5 → END. - Final state: human-reviewed recommendation, no code change attempted. The
incident is still fully recorded (RCA persisted at
rcatime —backend/app/db/repository.py). - Tested end-to-end: "low confidence blocks code_fix" in
backend/tests/unit/graph/test_builder.py; the guardrail demotion is covered in the commit-098a6dftest notes.
Branch C — rejected
- Reviewer decides the RCA is wrong or the fix is unnecessary. There is no reject
button in the UI (the dashboard only has FIX —
frontend/src/Incidents.jsx); a rejection is resuming withApprovalDecision(status="rejected"), which is exercised bymake graph-smoke ARGS="--decision reject"(backend/scripts/smoke_graph.py) and by tests (backend/tests/integration/test_stub_flow.py). - Router: status ≠ approved → END. No code change, no branch.
- The run stays checkpointed; the incident remains queryable.
Honest note for Q&A: the missing UI reject button is a real gap — say "rejection is a first-class path in the graph, tests, and smoke script; the dashboard currently only wires the approve action; adding a reject button is a one-line API call plus a button."
The fix_eligibility policy, stated precisely
| Condition | Result |
|---|---|
rca is None |
not eligible (defensive; can't happen on the normal path) |
confidence < 0.5 (CODE_FIX_MIN_CONFIDENCE) |
not eligible — "RCA isn't trustworthy enough to drive a code change" (commit 098a6df) |
severity == "low" |
not eligible — "not consequential enough to automate" (commit 098a6df) |
| otherwise | eligible → code_fix |
The threshold is configurable (CODE_FIX_MIN_CONFIDENCE in .env.example) — a
policy knob, not a magic number.