Researcher Shares Findings on Embedding Compression for Retrieval Model
Tony Wu outlines methods to reduce embedding sizes while preserving retrieval performance on a benchmark.
TLDR
Tony Wu posted about a retrieval model called NeoMME Retriever. The post states that high-resolution late-interaction embeddings start large per page. Token pooling together with asymmetric quantization can shrink those embeddings substantially. The resulting smaller embeddings retain most of the baseline performance level on the ViDoRe benchmark. The update also notes that the model reaches the highest score among compared systems when using late-interaction embeddings. Omar Khattab reposted the thread.
Combined views
9.6K
9 Sources, first seen 27d ago