Perplexity releases pplx-embed-v2-context-9b-preview contextual embedding model
Perplexity Research and turbopuffer have released pplx-embed-v2-context-9b-preview, a contextual embedding model for RAG pipelines that embeds each chunk with the full document in view. Training uses token-level distillation from a query-aware teacher instead of a single gold passage. Weights are on Hugging Face under the MIT license, but the model is not yet in the Perplexity API.
- 9B-parameter model outputs 2048-dim embeddings, with Matryoshka support for 1024 dims and native int8
- On context-bench at K=10: 45.5% answer recall and 40.6% evidence recall
- Leads voyage-context-4 by 14.4 points on answer recall and 5.0 on evidence recall
- Trained on roughly 430 datasets across 50+ languages; MIT weights require transformers>=5.4.0
Read next
AI