Vector search is a crude proxy for relevance

To find that a document or product is relevant to a query, we encode both the query and the document as a vector, dense or sparse, and test mathematical similarity.

This is one of the most common ways of testing relevance.

This assumes that the encoding/embedding captures all the information a human would use to judge relevance.

That assumption is flawed.

Embeddings capture only some aspects of a query-document relationship: syntax, semantics in a specific context, or other narrower signals.

But there is no universal embedding. Every encoding misses some part of user intent, which makes it an incomplete proxy for true relevance.

A user may mean:

  • a specific constraint, a hidden intent, a domain nuance
  • a tradeoff the query never states explicitly

That is exactly where vector search starts to break. So the mistake is not using vector search. The mistake is promoting it from retrieval tool to relevance model.

So what works better?

  • A better proxy is a re-ranker. Instead of relying only on vector similarity, a re-ranker examines the query-document pair more carefully and judges whether the match is actually relevant. It can also be tuned to a specific domain or context.
  • A general-purpose proxy is an LLM. It can evaluate whether a document is truly relevant to a query and explain the match with supporting evidence. For better performance, the model can be fine-tuned for a specific domain.
  • Another approach is Query Understanding. Decompose the query into domain-specific sub-intents. Map each sub-intent to an attribute of the document or product, and then evaluate relevance at that level.

Similarity helps retrieve. It does not decide relevance.

Embeddings are a convenient approximation, not a theory of relevance. Treating them as one leads directly to brittle search systems.

Similarity gets you candidates. Relevance requires judgment. Use embeddings to search faster, not to pretend you have solved relevance.

First published in the newsletter: read it on Substack.

Have a system that needs a second opinion?