6. Workflow design: sequential, parallel, router

If you only remember 3 things 1. All three flow shapes live in one graph: sequential (rca → recommend → approval), parallel (intake → {retrieval ∥ rag} → rca), and conditional routing (route_after_approval + fix_eligibility) — backend/app/graph/builder.py. 2. The router policy: approved + confidence ≥ CODE_FIX_MIN_CONFIDENCE (0.5) + severity ≠ "low" → code_fix; anything else → END. 3. The human interrupt always fires before the router — the router only decides what happens after the human decides (backend/app/graph/nodes/approval.py).

1. Sequential flow

Code (backend/app/graph/builder.py):

graph.add_edge("rca", "recommend")
graph.add_edge("recommend", "approval")
graph.add_edge("code_fix", END)

Why sequential here: recommend consumes the completed rca; approval consumes the recommendation; code_fix consumes the approval. Each step's input is the previous step's output — a strict dependency chain, so parallelism would be wrong.

Rubric line: "Workflow Design — sequential flow" (see 11-rubric-map.md).

2. Parallel flow (fan-out / fan-in)

Code (backend/app/graph/builder.py):

graph.add_edge(START, "intake")
graph.add_edge("intake", "retrieval")   # fan-out
graph.add_edge("intake", "rag")        # fan-out
graph.add_edge("retrieval", "rca")     # fan-in
graph.add_edge("rag", "rca")           # fan-in

Why parallel: retrieval (Splunk log window) and rag (Chroma similarity) both depend only on error_event and not on each other. Running them concurrently cuts their wall-clock sum to their max. Commit 9f65393 records the alternative considered: "No Send() needed since it's a static fan-out, not per-item map-reduce."

Why the merge is safe: the two branches write different fields (retrieved_logs vs similar_error_docs); the only shared fields (errors, events) use operator.add append reducers (backend/app/graph/state.py). LangGraph waits at rca until both branches complete (fan-in semantics).

Proof it actually runs in parallel: backend/tests/unit/graph/test_builder.py contains a test that spies on both branch nodes executing in one run, and the edge assertions check the fan-out/fan-in wiring explicitly.

Rubric line: "Workflow Design — parallel/branching flow".

3. Conditional routing (the router)

Code (backend/app/graph/builder.py):

def fix_eligibility(rca: RCAResult | None) -> tuple[bool, str]:
    if rca is None:
        return False, "no RCA available"
    if rca.confidence < get_settings().code_fix_min_confidence:
        return False, f"confidence {rca.confidence:.2f} below threshold — needs manual fix"
    if rca.severity == "low":
        return False, "low severity — straight-through suggestion, no fix needed"
    return True, f"severity={rca.severity} confidence={rca.confidence:.2f} — eligible for auto-fix"

def route_after_approval(state: IncidentState) -> str:
    if state.approval and state.approval.status == "approved":
        eligible, _reason = fix_eligibility(state.rca)
        if eligible:
            return "code_fix"
    return END

graph.add_conditional_edges("approval", route_after_approval, ["code_fix", END])

(CODE_FIX_MIN_CONFIDENCE = 0.5 — backend/app/config.py, .env.example.)

Why a router and not just "approved → fix": approval answers "did a human say yes?"; eligibility answers "is the RCA trustworthy enough to act on?" These are different questions. A human clicking FIX on a 0.2-confidence RCA shouldn't produce an automated code change. The router is the machine's second opinion on the human's decision — a policy layer between the human and the actuator.

Rubric line: "Workflow Design — conditional routing logic to direct data or tasks".

Walk-through: one incident down each router branch

Setup: cashier adds Greek Yogurt and clicks Total with DEMO_BE_CHECKOUT_ERROR=true (backend/app/demo_errors.py, README.md). The POS backend returns HTTP 500, logs it to Splunk via HEC (backend/app/telemetry.py), and the poller picks it up (backend/app/services/splunk_poller.py).

The run proceeds identically for all branches: intake normalizes → retrieval+rag run in parallel → rca produces RCAResult → recommend fills recommended_action → approval interrupts with the review payload. Then:

Branch A — approved AND high-confidence, non-low severity

Branch B — approved BUT low-confidence (or low severity)

Branch C — rejected

Honest note for Q&A: the missing UI reject button is a real gap — say "rejection is a first-class path in the graph, tests, and smoke script; the dashboard currently only wires the approve action; adding a reject button is a one-line API call plus a button."

The fix_eligibility policy, stated precisely

Condition Result
rca is None not eligible (defensive; can't happen on the normal path)
confidence < 0.5 (CODE_FIX_MIN_CONFIDENCE) not eligible — "RCA isn't trustworthy enough to drive a code change" (commit 098a6df)
severity == "low" not eligible — "not consequential enough to automate" (commit 098a6df)
otherwise eligible → code_fix

The threshold is configurable (CODE_FIX_MIN_CONFIDENCE in .env.example) — a policy knob, not a magic number.