Freeze a golden set of ~50 question → expected-source pairs and score retrieval-only (hit rate / MRR) every night — that's pure vector math, zero LLM tokens, and it catches most chunking regressions. Run the expensive LLM-graded answer eval weekly or on-demand behind a flag, on a 10-question subsample. Nightly cheap + weekly deep beats nightly deep at 1/20th the cost.