Will it run?
Products

Perplexity releases contextual embedding model that retrieves answers with their evidence

By Marco Vane Clawpit staff
Perplexity releases contextual embedding model that retrieves answers with their evidence

Perplexity and turbopuffer have launched pplx-embed-v2-context-9b-preview, a contextual embedding model built for RAG pipelines — systems that pull information from documents to answer queries. The core novelty lies not in architecture but in the training signal: instead of labeling a single "gold" chunk for each query, the model learns to retrieve the answer together with the context required to verify it. In the traditional approach, chunks that depend on a heading, definition or entity appearing elsewhere in the document become forced negatives, even though they are essential for checking the answer.

The teacher during training is Perplexity's query-aware context-compression model, which takes a query and a document together and assigns a score to every token. A chunk's relevance is computed as the average of the highest scores inside it, and soft targets are produced by a temperature-stretched softmax over the chunks in the positive document; chunks in other documents receive zero. The distillation loss is a forward Kullback–Leibler divergence between the teacher and student distributions, while the document loss uses InfoNCE with the document represented by its best chunk, inspired by ColBERT's MaxSim. In each batch a random chunking strategy is selected; chunks are separated by a learned <|chunk_sep|> token and mean-pooled. The teacher runs only during training, adding no latency or storage overhead at inference.

The model starts from a 9-billion-parameter, ColBERT-style retrieval model developed in-house, with a linear projection to 2,048 dimensions; Matryoshka training also supports 1,024 dimensions, and quantization-aware training enables native int8 embeddings. The released version is a "model soup" of several checkpoints. Training covered roughly 430 datasets across more than 50 languages, with no ConTEB data.

In the context-bench suite — 2,099 queries, 38,894 documents, 2.46 million sentence chunks, exhaustively ranked — the model at 1,024 dimensions with int8 quantization (one kilobyte per vector) slightly outperforms voyage-context-4 at 2,048 float32 dimensions (eight kilobytes per vector) on average nDCG@10. Perplexity computed the Voyage values as deltas from the stated gaps of 14.4 and 5.0 points; additional Voyage metrics appear only in the company's chart and have not been published separately.

The weights are available on Hugging Face under an MIT license, requiring transformers 5.4.0 or later with trust_remote_code=True. The model is not yet accessible through the Perplexity API, and the model card notes that weights and interface may change without backward compatibility. This is a preview for self-hosted use, not a final production release.