A post proposes storing fewer search vectors and rebuilding them for reranking
The author suggests using a smaller, pooled set of vectors to find candidate results, then a learned decoder to regenerate the original vectors and rerank those results.
TLDR
A post argues that late-interaction search models contain highly redundant vectors—the numerical representations used for matching—and that their geometry could be exploited. It suggests storing a smaller, pooled set to find candidate results, then decoding that set to regenerate the original vectors for reranking. The author sees this as a promising direction that starts to bridge the gap between ColBERT and MICE.
Combined views
24
1 Source, first seen 19d ago