Every RAG prototype clears the same low bar: it answers the three or four questions you personally tried while building it. That bar has almost no relationship to whether it will hold up in production, and teams keep getting surprised by the gap between the two.
When you build a retrieval system, you unconsciously test the queries you expect it to handle well. You already know roughly what’s in the documents, so you phrase questions the way the documents phrase answers. The system looks great — because you graded your own homework.
Real users don’t do this. They ask questions the way they think about the problem, which is rarely how your documents are written. The first hundred real queries will look nothing like the ten you used to convince yourself it worked.
Closing the gap isn’t about a fancier retriever. It’s about building the feedback loop before you need it:
None of this is exotic — it’s closer to unit testing than research. Teams skip it anyway because the demo already “worked,” and building an eval harness feels like overhead on top of a system that appears done. It isn’t overhead. It’s the difference between a system you can trust to ship and one you’re hoping holds up.
If you’re staring at a prototype that looked great in the demo and want a second opinion on whether it’ll survive contact with real traffic, that’s exactly the kind of review worth doing before launch, not after.