5. RAG pipeline

If you only remember 3 things 1. Corpus = 10 curated error-family entries in data/knowledge/error_docs.txt (verified count), chunked on \n---\n, embedded with Chroma's default ONNXMiniLM_L6_V2, stored in an in-memory Chroma collection rebuilt on every startup (backend/app/rag/ingest.py, backend/app/rag/store.py). 2. Retrieval is context, not proof: the RCA prompt explicitly says "Historical similarity is not proof" (backend/app/graph/nodes/rca.py), and similarity is a display-only approximation clamp(1 - distance) (backend/app/graph/nodes/rag.py). 3. Honest weakness: the corpus is tiny and demo-shaped; there are no retrieval-quality evals and no similarity threshold. We know; we have an answer (below).

Terms

Corpus

data/knowledge/error_docs.txt — 82 lines, 10 entries (verified: grep -c "^Error:" = 10; the smoke run prints "ingested 10 error-doc chunk(s)"). Entries include:

Entry format (convention documented in docs/architecture/rag-pipeline.md): Error / Component / Symptoms / Cause / Evidence / Solution / Validation, separated by a line containing ---. Several entries carry the epistemic rule "Cause: only if confirmed by request/response" — the corpus itself is written to avoid overclaiming.

Ingestion (backend/app/rag/ingest.py)

  1. Read ERROR_DOCS_PATH (default data/knowledge/error_docs.txt, backend/app/config.py).
  2. Split on CHUNK_DELIMITER = "\n---\n" — one chunk per entry; no sliding window, no overlap (appropriate because entries are already small and self-contained).
  3. Drop empty chunks; dedupe via dict.fromkeys.
  4. Reset the collection and rebuild — the text file is the source of truth, Chroma is derived state (never manually edited; docs/architecture/rag-pipeline.md §3).
  5. IDs are content-derived: "error-doc-" + sha256(content)[:16] (ingest.py) — stable across restarts, so evidence IDs in RCAs are reproducible.

Called from the FastAPI lifespan (backend/app/main.py); raises RuntimeError if 0 chunks — the app refuses to start with an empty knowledge base.

Vector store and embedding (backend/app/rag/store.py)

Retrieval (backend/app/graph/nodes/rag.py)

How retrieved docs reach the output

  1. rag writes similar_error_docs into state.
  2. rca's user prompt includes each doc's id + content (_build_user_prompt, backend/app/graph/nodes/rca.py).
  3. The RCA prompt allows citing those IDs in evidence_ids.
  4. The whitelist filter keeps only supplied IDs (see 04-prompt-engineering.md).
  5. The approval interrupt payload includes similar_error_docs, so the human reviewer sees exactly what historical context the model saw (backend/app/graph/nodes/approval.py).

Honest quality assessment

Dimension State Evidence
Corpus size Thin — 10 entries, demo-shaped data/knowledge/error_docs.txt
Corpus freshness Manual: edit file + restart (by design) ingest.py rebuild-on-start
Embedding Default MiniLM — decent for English prose, untested for our domain store.py
Retrieval quality No evals — no golden queries, no threshold on weak matches absence in backend/tests/
Hybrid search / reranking Not implemented — pure vector similarity rag.py
Multi-worker Broken by design — each process holds its own in-memory collection EphemeralClient
Persistence None — restart wipes and rebuilds (acceptable: source of truth is the file) store.py

The one-line answer when asked: "The corpus is deliberately small because this is a POC with a fixed demo scope — the pipeline is what matters. The architecture is corpus-agnostic: the file is the source of truth, ingestion is automatic, and growing it is an editorial task, not an engineering one. What we'd change for production is below."

What we would change for production

  1. Retrieval evals first (golden set of error → expected doc pairs, hit-rate@k and MRR tracked per change) — you cannot tune what you don't measure.
  2. Hybrid search: BM25 keyword + vector, because error codes (CO-500) are exact tokens that embeddings can dilute.
  3. A similarity threshold: currently even a weak match is returned; production should drop results below a tuned cutoff so RAG can say "no similar history".
  4. Embedding model choice: evaluate a domain-tuned or larger model against the eval set; pin the model artifact (versioned, vendored) so re-embedding is explicit.
  5. Persistent, shared vector store (Chroma PersistentClient or a vector service) once corpus size or multi-worker deployments justify it.
  6. Corpus growth + freshness: a review process for new entries (the file is already version-controlled), plus metadata (component, error code, revision) for filtered retrieval — listed as future work in docs/architecture/rag-pipeline.md §14.
  7. Claim-level citation checking: verify each cited doc actually supports the sentence it's attached to (also §14) — today we validate ID membership only.