5. RAG pipeline
If you only remember 3 things 1. Corpus = 10 curated error-family entries in
data/knowledge/error_docs.txt(verified count), chunked on\n---\n, embedded with Chroma's defaultONNXMiniLM_L6_V2, stored in an in-memory Chroma collection rebuilt on every startup (backend/app/rag/ingest.py,backend/app/rag/store.py). 2. Retrieval is context, not proof: the RCA prompt explicitly says "Historical similarity is not proof" (backend/app/graph/nodes/rca.py), and similarity is a display-only approximationclamp(1 - distance)(backend/app/graph/nodes/rag.py). 3. Honest weakness: the corpus is tiny and demo-shaped; there are no retrieval-quality evals and no similarity threshold. We know; we have an answer (below).
Terms
- RAG (Retrieval-Augmented Generation): before asking the LLM, fetch relevant documents and put them in the prompt, so the model answers from supplied text instead of memory.
- Embedding: turning text into a vector of numbers so "similar meaning" becomes "nearby vectors".
- Vector store: a database that finds nearest vectors. Ours is ChromaDB.
- Chunking: splitting the source file into pieces — one per entry here.
Corpus
data/knowledge/error_docs.txt — 82 lines, 10 entries (verified: grep -c "^Error:"
= 10; the smoke run prints "ingested 10 error-doc chunk(s)"). Entries include:
CO-500 CHECKOUT_SERVICE_CRASH(matches the deliberate ZeroDivisionError indata/sample_app/checkout_service.py)CH-502 CHILLER_SERVICE_CRASH,PAY-408 PAYMENT_TIMEOUT,INV-503 INVENTORY_SERVICE_FAILURE- Six POS validation/outage families (HTTP 422 quantity/lines/amount, 503 PLU lookup, 500 checkout, 502 card payment)
Entry format (convention documented in docs/architecture/rag-pipeline.md):
Error / Component / Symptoms / Cause / Evidence / Solution / Validation, separated by
a line containing ---. Several entries carry the epistemic rule "Cause: only if
confirmed by request/response" — the corpus itself is written to avoid overclaiming.
Ingestion (backend/app/rag/ingest.py)
- Read
ERROR_DOCS_PATH(defaultdata/knowledge/error_docs.txt,backend/app/config.py). - Split on
CHUNK_DELIMITER = "\n---\n"— one chunk per entry; no sliding window, no overlap (appropriate because entries are already small and self-contained). - Drop empty chunks; dedupe via
dict.fromkeys. - Reset the collection and rebuild — the text file is the source of truth, Chroma
is derived state (never manually edited;
docs/architecture/rag-pipeline.md§3). - IDs are content-derived:
"error-doc-" + sha256(content)[:16](ingest.py) — stable across restarts, so evidence IDs in RCAs are reproducible.
Called from the FastAPI lifespan (backend/app/main.py); raises RuntimeError if 0
chunks — the app refuses to start with an empty knowledge base.
Vector store and embedding (backend/app/rag/store.py)
chromadb.EphemeralClient— in-memory only, nothing persisted to disk.- Collection name:
error_docs. - Embedding function: Chroma's default
ONNXMiniLM_L6_V2(a local MiniLM model, runs via ONNX — no embedding API call, no per-query cost). Explicitly documented as a deployment consideration indocs/architecture/rag-pipeline.md§6.
Retrieval (backend/app/graph/nodes/rag.py)
- Query:
f"{error_code} {event} {message}"— e.g."CO-500 CHECKOUT_SERVICE_CRASH ZeroDivisionError while calculating checkout discount". top_k = min(RAG_TOP_K, collection_count)withRAG_TOP_K=5(backend/app/config.py).- Each result becomes an
ErrorDoc(doc_id, content, similarity, source_section)in state (backend/app/graph/state.py). similarity = clamp(1.0 - distance, 0, 1)— the code comments call it "approximate, display only"; it is not a calibrated probability.
How retrieved docs reach the output
ragwritessimilar_error_docsinto state.rca's user prompt includes each doc'sid+content(_build_user_prompt,backend/app/graph/nodes/rca.py).- The RCA prompt allows citing those IDs in
evidence_ids. - The whitelist filter keeps only supplied IDs (see 04-prompt-engineering.md).
- The approval interrupt payload includes
similar_error_docs, so the human reviewer sees exactly what historical context the model saw (backend/app/graph/nodes/approval.py).
Honest quality assessment
| Dimension | State | Evidence |
|---|---|---|
| Corpus size | Thin — 10 entries, demo-shaped | data/knowledge/error_docs.txt |
| Corpus freshness | Manual: edit file + restart (by design) | ingest.py rebuild-on-start |
| Embedding | Default MiniLM — decent for English prose, untested for our domain | store.py |
| Retrieval quality | No evals — no golden queries, no threshold on weak matches | absence in backend/tests/ |
| Hybrid search / reranking | Not implemented — pure vector similarity | rag.py |
| Multi-worker | Broken by design — each process holds its own in-memory collection | EphemeralClient |
| Persistence | None — restart wipes and rebuilds (acceptable: source of truth is the file) | store.py |
The one-line answer when asked: "The corpus is deliberately small because this is a POC with a fixed demo scope — the pipeline is what matters. The architecture is corpus-agnostic: the file is the source of truth, ingestion is automatic, and growing it is an editorial task, not an engineering one. What we'd change for production is below."
What we would change for production
- Retrieval evals first (golden set of error → expected doc pairs, hit-rate@k and MRR tracked per change) — you cannot tune what you don't measure.
- Hybrid search: BM25 keyword + vector, because error codes (
CO-500) are exact tokens that embeddings can dilute. - A similarity threshold: currently even a weak match is returned; production should drop results below a tuned cutoff so RAG can say "no similar history".
- Embedding model choice: evaluate a domain-tuned or larger model against the eval set; pin the model artifact (versioned, vendored) so re-embedding is explicit.
- Persistent, shared vector store (Chroma PersistentClient or a vector service) once corpus size or multi-worker deployments justify it.
- Corpus growth + freshness: a review process for new entries (the file is already
version-controlled), plus metadata (component, error code, revision) for filtered
retrieval — listed as future work in
docs/architecture/rag-pipeline.md§14. - Claim-level citation checking: verify each cited doc actually supports the sentence it's attached to (also §14) — today we validate ID membership only.