• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    LLMs Outperform Embedding Models at Much Higher Cost

    Researcher shares findings comparing LLMs to specialized embedding models on benchmarks.

    JH
    NM
    RP
    7 Sources, 41d ago, first seen 41d ago

    TLDR

    Niklas Muennighoff posted results from a new paper finding that LLMs now beat embedding models on tasks such as MTEB. The controlled study compared ten LLMs across six families to 26 embedding models. Andy Konwinski replied that LLMs can perform semantic embedding but cost 1000x more and run slower. He asked how long it will take before LLMs become cheaper and faster at the task. The paper is titled The Embedder's Dilemma and examines when to select each option.

    Combined views

    90.5K

    7 Sources, first seen 41d ago

    Combined views

    90.5K

    7 Sources, first seen 41d ago

    829 likes
    829 likes
    47 comments
    562 saves
    122 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    47 comments
    562 saves
    122 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    7 Sources

    @Muennighofffor retrieval tasks, especially if reasoning-heavy, the extra cost can be worth it for classification/sts tasks usually best to stick with embeddings. Clustering a bit mixed but as embedding is usually cheaper, LLM may not be worth it.
    @andykonwinskisure you can use an LLM to do your semantic embedding. it's only 1000x more expensive and slower. the bigger question is: how long before LLMs can do it better cheaper and faster? Will they ever?
    @tomaarsenAt 1400x cheaper, I know I'm sticking to embeddings (dense, sparse, multi-vector), plus hybrid (incl. bm25) and rerankers. Perhaps I'd even use listwise cross-encoders, they seem interesting.
    @jeremyphowardRT @tomaarsen: At 1400x cheaper, I know I'm sticking to embeddings (dense, sparse, multi-vector), plus hybrid (incl. bm25) and rerankers.…
    @rohanpaul_aiNew Harvard + Stanford paper says, don’t replace your embedding model with an LLM. Use embeddings for the cheap, fast first pass. And bring in an LLM only when the retrieval problem actually requires reasoning. An LLM costs up to 1,431X more than an embedding model of comparable quality. Across 10 LLMs and 26 embedding models on 37 tasks, Gemini 3.1 Pro scored 77.6 versus 77.2 for Octen-8B, a statistical tie. The cost was nowhere close: $154.14 for the LLM benchmark pass versus $0.11 for Octen-8B, or 1,431× more. The task split explains when that extra spend can make sense. LLMs led retrieval by 8.5 points, while embeddings led classification by 5.6; clustering, semantic similarity, and pair classification were effectively tied. Embeddings encode documents once and reuse vectors, while an LLM can read multiple documents together with the query and reason across them. So overall recommendation, embeddings for candidate retrieval, LLMs only where the shortlist actually needs reasoning. – arxiv. org/abs/2608.12875 Title: "The Embedder's Dilemma: LLMs Are Better, but at What Cost?"

    7 Sources

    @Muennighofffor retrieval tasks, especially if reasoning-heavy, the extra cost can be worth it for classification/sts tasks usually best to stick with embeddings. Clustering a bit mixed but as embedding is usually cheaper, LLM may not be worth it.
    @andykonwinskisure you can use an LLM to do your semantic embedding. it's only 1000x more expensive and slower. the bigger question is: how long before LLMs can do it better cheaper and faster? Will they ever?
    @tomaarsenAt 1400x cheaper, I know I'm sticking to embeddings (dense, sparse, multi-vector), plus hybrid (incl. bm25) and rerankers. Perhaps I'd even use listwise cross-encoders, they seem interesting.
    @jeremyphowardRT @tomaarsen: At 1400x cheaper, I know I'm sticking to embeddings (dense, sparse, multi-vector), plus hybrid (incl. bm25) and rerankers.…
    @rohanpaul_aiNew Harvard + Stanford paper says, don’t replace your embedding model with an LLM. Use embeddings for the cheap, fast first pass. And bring in an LLM only when the retrieval problem actually requires reasoning. An LLM costs up to 1,431X more than an embedding model of comparable quality. Across 10 LLMs and 26 embedding models on 37 tasks, Gemini 3.1 Pro scored 77.6 versus 77.2 for Octen-8B, a statistical tie. The cost was nowhere close: $154.14 for the LLM benchmark pass versus $0.11 for Octen-8B, or 1,431× more. The task split explains when that extra spend can make sense. LLMs led retrieval by 8.5 points, while embeddings led classification by 5.6; clustering, semantic similarity, and pair classification were effectively tied. Embeddings encode documents once and reuse vectors, while an LLM can read multiple documents together with the query and reason across them. So overall recommendation, embeddings for candidate retrieval, LLMs only where the shortlist actually needs reasoning. – arxiv. org/abs/2608.12875 Title: "The Embedder's Dilemma: LLMs Are Better, but at What Cost?"