Perplexity has introduced pplx-embed-v2-context-9b-preview, a contextual embedding model that processes document chunks together so each chunk’s representation reflects surrounding context. The company announced the preview on Sept. 30.
Its announcement thread describes the aim as retaining context that can be lost when retrieval systems split a long document into isolated chunks.
Training beyond a single answer chunk
Perplexity describes a training signal drawn from its query-aware context compression model. That model scores document tokens against a query; those scores are aggregated into chunk-level targets. The company says this trains the embedder to retrieve both answer chunks and supporting material.
On turbopuffer’s privately held context-bench, Perplexity claims the preview leads Answer and Evidence retrieval at every cutoff. The same update qualifies the comparison: its earlier context v1 4B model is slightly ahead on Document@3 and Document@5. The claimed lead therefore does not cover every document-retrieval measure.
A preview with compatibility limits
The model card documents 2,048-dimensional INT8 embeddings and training at both 1,024 and 2,048 dimensions. It also requires different encoding methods for queries and document chunks, warning that using the document method for queries degrades retrieval quality.
Perplexity warns that this is a preview rather than a final model. Weights, embeddings and the interface may change without backward compatibility. Its model card advises against mixing embeddings produced by this preview with those from a future release.