Other questions that share the most tags with this one.
Freeze a golden set of ~50 question → expected-source pairs and score retrieval-only (hit rate / MRR) every night — that's pure vector math, zero LLM tokens, and it catches most chunking regressions. Run the expensive LLM-graded answer eval weekly or on-demand behind a flag, on…
Wrote up the eval harness I use to compare agent runs deterministically — same seed, same tools, diff the trajectories.
10