Accepted answer: Freeze a golden set of 50 question → expected-source pairs and score retrieval-only (hit rate / MRR) every night — that's pure vector math, zero LLM tokens, and it catches most…
What's the cheapest way to eval a RAG pipeline nightly without burning a ton of tokens? We change chunking/prompts weekly and keep finding regressions late.