Perplexity releases embeddings designed to retrieve answers and supporting evidence

The preview model encodes document chunks together, and Perplexity says it leads on two retrieval benchmarks while using smaller vectors than Voyage's competing model.

By · Published

Primary source: Perplexity AI on X

Why it matters

Retrieval determines which evidence an AI system can use. Perplexity is betting that document-aware embeddings can improve answer and evidence retrieval while keeping vector storage smaller, though its headline results remain company-reported and one benchmark is private.

A compact bundle of document excerpts rests on an open source document, with one excerpt slightly lifted.

Perplexity, co-founded by CEO Aravind Srinivas, released a new embedding model designed to retrieve answers alongside the passages that explain or verify them. In an October 1st post on X, Perplexity introduced pplx-embed-v2-context-9b-preview, which encodes document chunks with the rest of the document in view.

For Srinivas, that work extends a core Perplexity product problem: getting a system to answer questions using relevant source material. He studied electrical engineering at IIT Madras and computer science at UC Berkeley, then worked as a research scientist at OpenAI. The new model applies that retrieval focus to a layer users may never see, but that can determine whether an AI answer is grounded in the right passage.

Most retrieval systems split long documents into chunks and embed each chunk on its own. That makes search more granular, but a sentence can lose the context that identifies a person, document, date, or definition. Perplexity's approach encodes all chunks in a document together, then returns a separate vector for each chunk. Perplexity says this adds document context without requiring a context-compression model at search time; that model serves as a training teacher instead.

Training for evidence, not just the answer

The training method addresses a common shortcut in retrieval benchmarks: labeling one passage as the correct answer, or the "gold chunk." That label can teach a model to find an answer sentence without surfacing the other passages needed to establish what the answer refers to or why it is credible.

Perplexity says its context-compression model scores document tokens for relevance to a query. The company aggregates those token-level scores into chunk-level training targets, giving the embedding model a graded signal for answer passages and supporting evidence. The teacher operates during training; Perplexity says the released model produces contextual chunk embeddings without adding a separate compression or reranking step at inference.

That distinction is useful for search products that must do more than return a plausible sentence. A system searching contracts, filings, or technical manuals needs to identify the correct document and retrieve enough surrounding evidence for a user or another model to check the answer. Perplexity's technical write-up describes context-bench as testing document disambiguation, answer retrieval, and evidence retrieval across long documents.

Perplexity reports that context-bench contains 2,099 queries and 38,894 documents. The benchmark is privately held by turbopuffer, whose founders built search infrastructure after working together on Shopify's systems. The private test is intended to make benchmark contamination harder, and Perplexity says it submitted the model for blind evaluation. Perplexity reports that the preview leads on Answer and Evidence Recall at every tested cutoff, while its earlier 4B context model remains slightly ahead on Document@3 and Document@5.

Those scores are company-reported. A private benchmark can make it harder for a model developer to train directly against its questions, but outsiders cannot inspect the full test set or independently reproduce the ranking from the benchmark itself. Perplexity's blog supplies benchmark details and score comparisons; it does not make the underlying evaluation independently verifiable.

On ConTEB, a public contextual-embedding benchmark, Perplexity says the preview has the highest average nDCG@10 among the models it tested, but does not lead every task. Perplexity also says the model beats Voyage AI's voyage-context-4 on chunk retrieval while using about 1 KB per vector at 1,024 dimensions with int8 quantization, versus about 8 KB for Voyage's 2,048-dimensional float32 vectors. That storage comparison is relevant to large indexes, where vector size affects memory and infrastructure costs, but it is a vendor comparison rather than an independently reproduced result.

A preview with production caveats

The model is available on Hugging Face. Its model card labels it a preview and warns that weights, embeddings, and the interface may change without backward compatibility. It also requires separate methods for queries and documents: encode_queries for queries and encode for document chunks. Using the document method for queries silently degrades retrieval quality, according to the card.

The release follows Perplexity's earlier description of its embedding and ranking stack. Together, those disclosures show Perplexity working on retrieval infrastructure beneath its answer product, where better ranking and smaller vectors could help answer systems find useful material at manageable index costs. The preview's benchmark claims give a reason to test the model; its compatibility warning means production teams will need to plan for re-embedding if the model changes.

Reader comments

Conversation for this story loads after sign-in.