• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Perplexity introduces contextual embedding preview with compatibility warning

    Perplexity claims leading Answer and Evidence retrieval on context-bench, while its earlier model leads two Document measures. The preview’s embeddings may be incompatible with future releases.

    AS
    MP
    PE
    12 Sources, ,

    TLDR

    Perplexity’s contextual embedding preview encodes document chunks together and trains on relevance signals meant to retrieve answers and supporting context. The company claims leading Answer and Evidence retrieval on context-bench, but its earlier model is slightly ahead on Document@3 and Document@5. The model card warns that future versions may break compatibility and advises against mixing preview embeddings with later releases.

    Combined views

    264.3K

    12 Sources, first seen 9h ago

    Combined views

    264.3K

    12 Sources, first seen 9h ago

    1.5K likes

    Useful links

    Perplexity AI

    pplx-embed: State-of-the-Art Embedding Models for Web-Scale Retrieval

    Byte Goose AI. · YouTube

    Perplexity AI Multilingual Open-Weight Retrieval Models. Late Chunking and Context Aware Embeddings.

    Karan Prasad - Founder, Obvix Labs

    Your RAG Pipeline Has a Context Problem. Perplexity Just Open-Sourced the Fix.

    AISoftScope · YouTube

    Perplexity Just Beat Google's Embedding Model — And Released It for Free
    Today's Rank

    #10

    Today's Rank

    #10

    9h ago
    first seen 9h ago
    1.5K likes
    96 comments
    826 saves
    105 reposts

    Perplexity has introduced pplx-embed-v2-context-9b-preview, a contextual embedding model that processes document chunks together so each chunk’s representation reflects surrounding context. The company announced the preview on Sept. 30.

    Its announcement thread describes the aim as retaining context that can be lost when retrieval systems split a long document into isolated chunks.

    Training beyond a single answer chunk

    Perplexity describes a training signal drawn from its query-aware context compression model. That model scores document tokens against a query; those scores are aggregated into chunk-level targets. The company says this trains the embedder to retrieve both answer chunks and supporting material.

    On turbopuffer’s privately held context-bench, Perplexity claims the preview leads Answer and Evidence retrieval at every cutoff. The same update qualifies the comparison: its earlier context v1 4B model is slightly ahead on Document@3 and Document@5. The claimed lead therefore does not cover every document-retrieval measure.

    A preview with compatibility limits

    The model card documents 2,048-dimensional INT8 embeddings and training at both 1,024 and 2,048 dimensions. It also requires different encoding methods for queries and document chunks, warning that using the document method for queries degrades retrieval quality.

    Perplexity warns that this is a preview rather than a final model. Weights, embeddings and the interface may change without backward compatibility. Its model card advises against mixing embeddings produced by this preview with those from a future release.

    Perplexity
    Featured Source
    96 comments
    826 saves
    105 reposts

    13 Sources

    huggingfaceperplexity-ai/pplx-embed-v2-context-9b-preview · Hugging Face
    @perplexity_aiWe built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage
    @bo_wangboin the past months we worked closely with tpuf team, iterating our models beyond web search. During the journey, we built a new way of training contextual embedding models, which connects our previous work, the context pruner. We're sharing how we train the model, and how it performs on @turbopuffer 's private bench. great work from @ESL_Sarah , Markus, @antoine_chaffin , Louis, Max and special thanks to @n0riskn0r3ward , who helps evaluate, iterate again and again 🐐. We'll work closely together to continue improve it, and eventually offer it to all.
    @AravSrinivasWe’re open sourcing our state-of-the-art contextual embedding models, which perform best in turbopuffer’s context-bench.
    @denisyaratswe developed a new way to train contextual embedding models. instead of embedding each chunk on its own, the model encodes the whole document once and pools chunk vectors afterward, so every chunk sees the full document. the new part is the training signal: instead of one labeled "gold" chunk per query, we distill relevance from our context compression model, which scores every document token against the query. this teaches the model to retrieve both the answer and the context that supports it. pplx-embed-v2-context-9b-preview sets a new SOTA on ConTEB and on turbopuffer's new private context-bench, which it was evaluated on blind: +14.4 points answer recall@10 over voyage-context-4. and at 1024 dims in int8 (1KB per vector), it still beats voyage-context-4 at 8KB per vector.
    @turbopufferwe maintain several internal benchmarks to guide our customers toward better search relevance @perplexity's new model tops context-bench, our internal contextual embedding benchmark, and boosts document recall@10 by 50%+ over traditional SOTA embedding models
    @MParakhinPrediction: "in context of ..." encoding will be more and more popular. Back in the day I was making fun of the rudimentary "split into paragraphs" approach for RAG, popularized in one famous tutorial - it is getting better now :-) Turbopuffer founder is Shopify alumni, btw

    Sentiment

    Positive92.3%7.7%Negative

    Summary

    Useful Links

    Perplexity AI

    pplx-embed: State-of-the-Art Embedding Models for Web-Scale Retrieval

    Related Videos

    Sentiment

    Positive92.3%7.7%Negative

    Many accounts welcomed Perplexity’s open-sourced contextual embedding models for their document-level training and practical efficiency gains, while some replies questioned the SOTA claims on a privately held benchmark.

    Based on 68 sentiment-bearing replies from 61 accounts across 6 conversations.

    Karan Prasad - Founder, Obvix Labs

    Your RAG Pipeline Has a Context Problem. Perplexity Just Open-Sourced the Fix.
    Perplexity AI Multilingual Open-Weight Retrieval Models. Late Chunking and Context Aware Embeddings.Byte Goose AI. · YouTube
  • Perplexity Just Beat Google's Embedding Model — And Released It for FreeAISoftScope · YouTube
  • 13 Sources

    huggingfaceperplexity-ai/pplx-embed-v2-context-9b-preview · Hugging Face
    @perplexity_aiWe built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage
    @bo_wangboin the past months we worked closely with tpuf team, iterating our models beyond web search. During the journey, we built a new way of training contextual embedding models, which connects our previous work, the context pruner. We're sharing how we train the model, and how it performs on @turbopuffer 's private bench. great work from @ESL_Sarah , Markus, @antoine_chaffin , Louis, Max and special thanks to @n0riskn0r3ward , who helps evaluate, iterate again and again 🐐. We'll work closely together to continue improve it, and eventually offer it to all.
    @AravSrinivasWe’re open sourcing our state-of-the-art contextual embedding models, which perform best in turbopuffer’s context-bench.
    @denisyaratswe developed a new way to train contextual embedding models. instead of embedding each chunk on its own, the model encodes the whole document once and pools chunk vectors afterward, so every chunk sees the full document. the new part is the training signal: instead of one labeled "gold" chunk per query, we distill relevance from our context compression model, which scores every document token against the query. this teaches the model to retrieve both the answer and the context that supports it. pplx-embed-v2-context-9b-preview sets a new SOTA on ConTEB and on turbopuffer's new private context-bench, which it was evaluated on blind: +14.4 points answer recall@10 over voyage-context-4. and at 1024 dims in int8 (1KB per vector), it still beats voyage-context-4 at 8KB per vector.
    @turbopufferwe maintain several internal benchmarks to guide our customers toward better search relevance @perplexity's new model tops context-bench, our internal contextual embedding benchmark, and boosts document recall@10 by 50%+ over traditional SOTA embedding models
    @MParakhinPrediction: "in context of ..." encoding will be more and more popular. Back in the day I was making fun of the rudimentary "split into paragraphs" approach for RAG, popularized in one famous tutorial - it is getting better now :-) Turbopuffer founder is Shopify alumni, btw
    Summary

    Many accounts welcomed Perplexity’s open-sourced contextual embedding models for their document-level training and practical efficiency gains, while some replies questioned the SOTA claims on a privately held benchmark.

    Based on 68 sentiment-bearing replies from 61 accounts across 6 conversations.

    Useful Links

    Perplexity AI

    pplx-embed: State-of-the-Art Embedding Models for Web-Scale Retrieval

    Karan Prasad - Founder, Obvix Labs

    Your RAG Pipeline Has a Context Problem. Perplexity Just Open-Sourced the Fix.

    Related Videos

    • Perplexity AI Multilingual Open-Weight Retrieval Models. Late Chunking and Context Aware Embeddings.Byte Goose AI. · YouTube
    • Perplexity Just Beat Google's Embedding Model — And Released It for FreeAISoftScope · YouTube

    Related

    Perplexity Computer adds Automations for ongoing work

    Perplexity says Automations can take action on a schedule or in response to event-based triggers. They work with memory, skills and connected apps such as Slack, Gmail, Outlook and Linear.

    Bumblebee and Numbat feed findings to Computer, with humans approving rule changes

    Perplexity says Bumblebee scans macOS and Linux developer machines for risky packages, extensions and AI tool configs. Its Numbat tool monitors coding agents and can block dangerous actions.

    Portable Computer keeps its model on-device and sandboxes approved actions

    Perplexity says its local-first agent uses an on-device 27B model. The company reports that its harness scored 82.6% on real knowledge work, ahead of open-source harnesses Pi and Hermes.