On September 30, 2026, Perplexity Research and turbopuffer published Contextual embedding beyond the gold passage. The artifact is pplx-embed-v2-context-9b-preview: a 9B contextual embedding model that produces one vector per chunk after encoding the whole document.
If you already read explainx.ai's embeddings and vector-search guide, the pain is familiar. You split a 10-K, a lease, or a clinical protocol into windows. The sentence that answers the query no longer names the company, the revision, or the table header. Independent chunk embeddings then collide with near-duplicate sentences from the wrong sibling document.
The preview tries to train that away. The API is not the story this week. Hugging Face is. Vendor charts are still vendor charts.
TL;DR — should you touch this preview?
| Question | Direct answer |
|---|---|
| What shipped? | Research post + HF preview of pplx-embed-v2-context-9b-preview (Sep 30, 2026) |
| What is new vs v1? | Teacher is Perplexity's context compression model, not a single gold-chunk LLM label |
| Inference shape | One document pass, learned chunk-separator token, mean-pool tokens per chunk |
| Dims / dtype | Matryoshka 2048 and 1024; native int8 via QAT; released as a checkpoint soup |
| Where to run it | Hugging Face preview. Perplexity says it is working on the API |
| Benches named | Private turbopuffer context-bench; public ConTEB; internal Q2C/Q2D suites |
| Headline vendor numbers | At K=10 on context-bench: 45.5% Answer, 40.6% Evidence, 31.1% All-Evidence; Document@1 15.2%, Document@10 61.6% |
| Should you rebuild production? | Only after a private eval. Do not treat ConTEB average as a buying decision |
What people are asking after the headline
"Is this just another pplx-embed checkpoint?"
No. Perplexity already ships independent pplx-embed-v1 models and contextual pplx-embed-context-v1 models in 0.6B and 4B sizes. Those are the ones in the contextualized embeddings API docs (pplx-embed-context-v1-0.6b at $0.008 / 1M tokens and pplx-embed-context-v1-4b at $0.05 / 1M, 32K context). This v2 preview is a 9B ColBERT-initialized student with a different training objective.
The serving story you already have on explainx.ai is Fast Embeddings on GPUs: Ivy, Tulip, ROSE, CUDA graphs. That post is how Perplexity runs embedders. This post is how they train a contextual one. Do not collapse the two.
"What is a gold passage and why would I care?"
Classic dense retrieval labels one chunk as relevant. Contrastive training then treats every other chunk in the same document as a negative. That is convenient for MS MARCO-style datasets. It is a bad match for documents where the answer sentence is true only after you also retrieve the definition, the revision banner, or the table header.
Perplexity's own sentence from the post:
A gold passage is often not enough in practice. It may contain the answer but lack the supporting context needed to understand or verify it.
That is the whole product pitch. If your RAG reader is an agent that should cite and verify, you want Answer@K and Evidence@K, not a single gold hit.
"Did they just invent late chunking?"
No. Late chunking — encode the document once, pool afterwards — is the standard contextual-embedding pattern. Perplexity says prior work (ConTEB and their own pplx-embed-context-v1) already showed gains over independent chunks on long documents. The claimed novelty is how labels are generated: distill token-level scores from a query-aware context compressor, then aggregate those scores onto whatever chunk boundaries you sampled that batch.
That is a training-time teacher. At inference there is no extra compressor and no extra reranker in the recipe they describe. Storage is still one vector per chunk.
"Can I call this from the Perplexity API this week?"
Perplexity writes, in the release section:
A preview of our model is available on Hugging Face, and all results in this post are reported for this preview. We are working on making the model available through the Perplexity API.
Until that lands, production callers stay on v1 IDs. Treat the 9B checkpoint as an eval toy unless you already self-host large embedders.
How the training recipe actually works
The student is an in-house 9B ColBERT retriever plus a linear projection to 2048 dimensions. Chunk boundaries are marked with a learned <|chunk_sep|> token. Chunk vectors are mean-pooled token embeddings. Queries go through the same model and are also mean-pooled.
Two losses share a batch:
- Document-level InfoNCE. Document score is the max chunk–query cosine (ColBERT's MaxSim idea, applied to chunks). In-batch documents are negatives.
- Chunk-level distillation. The compressor scores every token in the positive document. A chunk's teacher score is the mean of its top-n token scores. Softmax those scores into a target; other documents' chunks get zero. The student matches that distribution with forward KL.
Combined objective: weighted sum of the two. Batches are sampled from a single dataset so in-batch negatives are hard. Chunking strategy is resampled per batch, which is the point of token-level teachers: you do not re-annotate when you change window size.
Training data: roughly 430 public and in-house query–document datasets, 50+ languages. In-house pairs come from PII-filtered production data and synthetic queries over long documents. Perplexity says no ConTEB training data and no context-bench data during development. The released weights are a model soup of checkpoints from the same run.
That last sentence matters more than the marketing average. Soup + private bench + self-reported ConTEB is a lot of degrees of freedom. Reproduce on your leases before you throw out BM25.
What context-bench actually measures
turbopuffer holds context-bench private: 2,099 queries, 38,894 documents, sentence chunks totaling 2,458,072 vectors, 21 domains. Median primary target document is about 6,100 tokens; 1,061 of 1,197 distinct primary targets have at least 200 sentences. Eval is exhaustive over all chunks, so index config is not the confounder.
Three metrics, not one:
| Metric | What it asks |
|---|---|
| Document@K | Did the right long document beat near-duplicate siblings? |
| Answer@K | Did an accepted answer sentence land in the top K chunks corpus-wide? |
| Evidence Recall@K | After dropping the top answer chunk, and only if the gold doc is in the top 10, how many evidence groups are recovered? |
The lease example in the paper is the useful one. Many files share the sentence "Monthly rent is $X." The model has to use address and year context that may sit thousands of tokens away. That is closer to enterprise RAG than a Wikipedia gold passage.
Perplexity reports the preview leading Answer@K and Evidence Recall@K at every cutoff they plotted, plus All-Evidence@10. At K = 10: 45.5% answer recall, 40.6% evidence recall, 31.1% all-evidence. Document@1 15.2%, Document@10 61.6%. Versus voyage-context-4 at K=10: +14.4 points answer, +5.0 evidence.
Those are Perplexity's numbers on a bench turbopuffer holds. Other labs can submit to contextbench@turbopuffer.com. You cannot download the labels and rerun tonight. That is the contamination argument, and it is also why you should not paste 45.5% into a board deck as independent science.
ConTEB, domain suites, storage — still vendor benches
On ConTEB, Perplexity says the preview has the highest average nDCG@10 among the models in their figure. It does not win every task: they say pplx-embed-context-v1-4B is higher on NarrativeQA and Nemotron-3-8B is highest on COVID-QA (where lexical matching on medical terms can beat context). Average-without-per-task is how vendor plots get shared. Read the task bars.
On their internal query-to-chunk (Q2C) and query-to-document (Q2D) suites built from public datasets (LegalBench, FinanceBench, plus extra tasks), they say the preview leads average chunk retrieval and some domains, while voyage-context-4 is higher in finance and multilingual, and Nemotron-3-8B leads conversation. On document retrieval the preview is slightly below voyage-context-4 on average. If you only needed document IDs, a contextual 9B is not an automatic upgrade.
Storage claim, still theirs: contextualization does not multiply vector count versus the same chunker. Truncating 2048 → 1024 and using int8 drops bytes per vector. They say 1024-d int8 is 1 KB/vector and slightly beats voyage-context-4 at 2048-d float32 (8 KB/vector) on their chunk-retrieval plot. 2048-d int8 is 2 KB. Those figures are vector payload only, not HNSW graphs or metadata.
Chunk-size sweep: mean nDCG@10 across 74 MTEB tasks falls from 81.0% at 64-token chunks to 79.9% at 512 tokens. Modest. Useful if your indexer already picked a window for other reasons.
All of those evals encode documents up to 32,768 tokens in one pass. If your PDFs are longer, you still have a packing problem. The compressor teacher does not magically appear at query time.
How this sits next to agentic RAG and Cohere Embed 5
explainx.ai's RAG vs agentic RAG argument still holds for code: grep and structure beat naive chunk vectors. Contextual embeddings are an attempt to make chunk vectors less naive on long prose, filings, and manuals — the corpora where agents still retrieve then read.
Q2D-Web is a different axis: web-scale first-stage retrieval under agent-rewritten queries. A model that wins private long-doc evidence recall can still lose Recall@1000 on 190M web docs. Do not pick an embedder from one plot.
Same week, Cohere shipped Embed 5 Pro and Fast with a shared embedding space, 128K context, and vendor ViDoRe V3 scores. That is a hosted multimodal API with list prices. Perplexity's v2 is an open preview of a late-chunk text model with API forthcoming. Different buying motion. If you need images and a SLA this week, Cohere is the product. If you need to inspect a 9B contextual checkpoint, Perplexity is the lab drop.
For a closed-vs-open shortlist that is already stale on dates, see the top 10 embedding models roundup and refresh it with your own gold set. HyDE and Sentence Transformers v6 ColBERT remain the technique companions: hypothetical queries and late interaction are not replaced by one 9B soup.
What to do this week (if you actually retrieve documents)
- Do not wait for the API if you only wanted a paper. Read the post, note the teacher, note the private bench.
- If you self-host, pull the Hugging Face preview, encode a 100–200 query slice of your long docs, and score Answer plus a cheap evidence checklist (did the retrieved window contain the entity name?). Compare against independent chunks from the same backbone class and against v1 contextual if you already pay Perplexity.
- Match query encoding. Contextual doc models still need queries encoded the way the card says — v1 docs wrap queries as a one-element inner list. Confirm the v2 card before you mix spaces.
- Keep hybrid search. Contextual vectors do not retire BM25 on SKUs, error codes, or clause IDs. See the semantic vs hybrid search guide.
- Budget GPU like an LLM prefill. A 9B bidirectional pass over 32K tokens is not a MiniLM call. Perplexity's own GPU serving writeup is the honest infra companion.
- Log agent queries separately. If your users are agents, add a Q2D-Web-style split rather than only human FAQ queries.
Example shape for a contextualized v1 API call (still the documented production path, not the v2 preview):
curl -X POST https://api.perplexity.ai/v1/contextualizedembeddings \
-H "Authorization: Bearer $PPLX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "pplx-embed-context-v1-4b",
"input": [[
"Northlake is the registrant in this 10-K.",
"The company repurchased $12.8 billion of shares."
]]
}'
The v2 preview will not accept that ID. When the API ID lands, re-read the model card for encoding_format, Matryoshka dimensions, and whether queries share the contextual endpoint.
Honest limitations
- Preview tag. Weights can move. Numbers in the blog are for this soup.
- Private primary bench. You cannot audit label quality or contamination from the outside.
- Teacher is Perplexity's compressor. Distillation quality is capped by that model. Errors in token relevance become chunk supervision.
- 32K single pass. Longer docs still need packing, hierarchical retrieval, or an agent that greps.
- Not multimodal. Page images, slides, and scanned tables are Cohere Embed 5 / ColPali territory, not this checkpoint.
- ColBERT init, single-vector out. You get chunk vectors, not token-level late interaction at query time. Storage looks like a bi-encoder index.
- Vendor domain plots. Finance and conversation losses versus named baselines are in their own figures. Copy those caveats into your eval notes.
Related on explainx.ai
- Cohere Embed 5 Pro vs Fast and ViDoRe V3
- Perplexity Fast Embeddings on GPUs — Ivy, Tulip, ROSE
- Q2D-Web: agentic RAG retrieval benchmark
- What are embeddings? Vector search complete guide
- RAG vs agentic RAG
- Top 10 open and closed embedding models
- What is an embedding? Examples
- Sentence Transformers v6 ColBERT
Official sources: Contextual embedding beyond the gold passage, Perplexity contextualized embeddings docs, pplx-embed Hugging Face collection.
Model IDs, benches, and API status are as of October 1, 2026, from Perplexity's September 30, 2026 Hub post. Preview weights and vendor scores can change; re-check Hugging Face and the Hub before you freeze an index.
