The eval gap: why your RAG demo works and your RAG system doesn't

Every RAG prototype clears the same low bar: it answers the three or four questions you personally tried while building it. That bar has almost no relationship to whether it will hold up in production, and teams keep getting surprised by the gap between the two.

The demo only tests what you already believe

When you build a retrieval system, you unconsciously test the queries you expect it to handle well. You already know roughly what’s in the documents, so you phrase questions the way the documents phrase answers. The system looks great — because you graded your own homework.

Real users don’t do this. They ask questions the way they think about the problem, which is rarely how your documents are written. The first hundred real queries will look nothing like the ten you used to convince yourself it worked.

What an actual eval loop looks like

Closing the gap isn’t about a fancier retriever. It’s about building the feedback loop before you need it:

  • A held-out question set you didn’t write from the documents. Pull real questions from support tickets, sales calls, or a domain expert who wasn’t involved in building the system.
  • A retrieval metric, not just an answer metric. Did it find the right passage? An LLM can generate a fluent, wrong answer from the wrong passage just as easily as a right one from the right passage — you won’t tell the difference downstream unless you check retrieval separately.
  • Failure modes categorized, not averaged away. “83% accuracy” hides whether the 17% is scattered noise or one entire question type your system can’t handle. The second is a design problem; the first is tuning.
  • A cadence for re-running it. Every change to the prompt, the chunking strategy, or the embedding model should re-run against the same held-out set, so you know whether you actually improved things or just moved the failures around.

The real cost of skipping this

None of this is exotic — it’s closer to unit testing than research. Teams skip it anyway because the demo already “worked,” and building an eval harness feels like overhead on top of a system that appears done. It isn’t overhead. It’s the difference between a system you can trust to ship and one you’re hoping holds up.

If you’re staring at a prototype that looked great in the demo and want a second opinion on whether it’ll survive contact with real traffic, that’s exactly the kind of review worth doing before launch, not after.

Have a system that needs a second opinion?