Assaf Elovic's GPT Researcher tagged v3.7.0 on September 26, 2026 as "Jev update and many bug fixes." The headline feature is a context filter: after scrape, each chunk is scored by Jev (TypeSafe AI's System One model) on usefulness for the sub-query, not cosine similarity to the query embedding. The v3.7.0 release notes land feat #2161 and feat #2162. Official docs live at docs.gptr.dev/docs/gpt-researcher/gptr/context-filter (verified live: title "Context Filter | GPT Researcher").
If you saw an aggregator line like "73% relevant context," treat it as a compressed precision number, not a product-quality score. Primary eval notes say 73% of chunks Jev kept were judged relevant versus 46% for embeddings on 28 replayed tasks. Users without a TypeSafe key do not get that Jev precision; they get BM25.
explainx.ai already covers TypeSafe's Jev launch, wiring Jev into agent routing, and cheap verification checkpoints. This post is the research-agent instance of that pattern: score retrieved passages before the writer spends tokens on them.
TL;DR
| Question | Answer |
|---|---|
| What shipped? | GPT Researcher v3.7.0 (PyPI gpt-researcher 0.16.0 in the release notes) |
| Does Jev replace RAG? | No. Filter after scrape. Vector stores and CONTEXT_FILTER=embeddings still exist |
| Default with no TypeSafe key? | BM25 keyword, not embeddings (#2162) |
| What is 73%? | Kept-passage precision for Jev vs 46% embeddings on 28 tasks, not report accuracy |
| Keyword vs embeddings? | BM25 matched or beat embeddings on the published measures; directional 14–6 head-to-head |
| Python? | 3.12+ required |
| How to opt into old path? | CONTEXT_FILTER=embeddings |
| Where is the eval? | evals/context_filter/ |
How the filter sits in the pipeline
Search still finds URLs. Scrapers still pull pages. The new step is select_context() in gpt_researcher/context/select.py: each page is capped at 50,000 characters, split into 1,000-character chunks with 100 characters of overlap, then ranked.
Jev and embeddings keep up to 10 chunks per sub-query. Keyword keeps up to 25. Sub-queries run concurrently; joined passages go to the writer. Report chat uses the same function, treating the finished report as one page and the latest message as the query.
Small page sets (under 8,000 characters total, and no more pages than the chunk budget) skip filtering in every mode.

That is closer to a reranker than to "RAG is dead." If you already think in RAG versus MCP retrieval terms: GPT Researcher still retrieves; it changed how it prunes.
Does Jev replace RAG?
No. Three facts from the docs and PRs:
- Jev never generates the report. It returns a 0–3 usefulness score (probability-weighted rubric). The writer LLM still writes.
- Embeddings are optional, not deleted.
CONTEXT_FILTER=embeddingsrestores cosine similarity. A caller-suppliedvector_store=still uses embeddings. Detailed-report section overlap still uses embeddings when a model is available, otherwise keywords. - Without
TYPESAFE_API_KEY, you are not on Jev.autofalls through to BM25. That is the opposite of "everyone now runs a System One RAG stack."
The useful comparison is Jev as a typed usefulness judge versus embedding similarity as a topical neighbor search. TypeSafe's own product story is the same split explainx.ai documented at launch: Jev does not replace chat or long-form generation. GPT Researcher used that primitive on chunks, the same family of decision as agent routing and post-retrieval checkpoints.
What people are asking: the 73% headline
Official copy: Jev's kept context is "59% more relevant than embeddings, at the same cost (73% of kept passages relevant, against 46%)." Relative lift: (73 − 46) / 46 ≈ 59%. Absolute: 73 vs 46, not "the product is 73% accurate."
| Claim you might see | What the eval actually says |
|---|---|
| "73% relevant context" | Precision of kept chunks, judged by gpt-5.4-mini |
| Implies all users | Only Jev mode with a working TypeSafe key |
| Implies BM25 is 73% | Keyword shipped at 51% precision (relative-threshold, up to 25) |
| Implies better answers | SimpleQA 18–20 / 20 for every filter — saturated |
| Implies huge eval | 28 tasks: 20 SimpleQA + 8 open-ended, one writer (gpt-5.4), one report type |
Head-to-head (blind, both orders, win only if both agree): Jev 15 · 10 · 3 vs embeddings (sign test on 18 non-ties, p ≈ 0.008 in the write-up). Keyword 14 · 8 · 6 (p ≈ 0.12 — directional, not as strong).
Unranked control (10 chunks round-robin, no ranking): 33% precision. Forced Jev top-10 with no threshold: 50% — the 1.5 min score is where most of the Jev gain comes from. Cosine similarity cannot say "nothing else on this page is worth including."
Do not treat 73% as a substitute for reading evals/context_filter/README.md. Caveats in that README: 28 questions, same model family writes and judges, detailed/deep research not replayed.
Cost and latency: Jev vs BM25 vs embeddings
Numbers below are from the shipped docs table and README footnotes (median over the 28-task replay unless noted). Writer time was about 45 seconds for every filter; total run time moves mainly with the filter step. Docs put a typical research run around two minutes, with Jev adding about 0.7s at the median versus embeddings.
| Mode | Kept-passage precision | Filter time (median) | Context sent (median) | Cost per report |
|---|---|---|---|---|
| Jev | 73% | 1.7s (p90 3.6s at concurrency 64) | 4.5k tokens | $0.115 |
| Keyword (BM25) | 51% | 0.02s | 6.4k tokens | $0.116 |
| Embeddings | 46% | 1.0s (p90 1.3s) | 7.6k tokens | $0.117 |
| Unranked control | 33% | ~0.01s | 8.0k tokens | $0.117 |
| No filter | n/a | — | 33.4k tokens (max 108k) | $0.192 |
Jev API cost. Docs: Jev bills $0.042 per million input tokens, output free — the same published rate explainx.ai used in checkpoint cost math. Each chunk is its own POST https://api.typesafe.ai/v1/systemone call. Filtering added about half a cent per report in the benchmark; README tracked $0.59 of Jev across all configurations in the full eval spend (~$41 writing/filtering plus judging).
Throughput. Default JEV_CONCURRENCY=64 (shared per event loop). Warm call ~0.33s; ~0.45s at 13k tokens; 64 concurrent calls ~1.1s. Raising to 128 did not make filtering faster. HTTP 429/529 retry up to three times; other failures raise JevError and select_context falls back to keyword.
BM25. Pure Python in gpt_researcher/context/lexical.py, no extra dependency. ~20ms per typical sub-query. k1 = 1.5, b = 0.75. Keep chunks scoring at least 50% of the best (KEYWORD_RELATIVE_THRESHOLD), up to 25. Plain top-10 BM25 lost to embeddings (40% precision, 7–12 head-to-head). The relative threshold is the shipped default for a reason.
No filter. Wins open-ended pairwise vs embeddings (8–0 on the eight open-ended tasks) because more material reaches the writer, but costs ~65% more per report on average (~83% on open-ended) and is not faster. Fact-finding (SimpleQA) does not benefit.
For TypeSafe's broader speed/cost marketing, keep explainx.ai's Jev claims fact-check in mind: GPT Researcher's table is a narrow filter benchmark, not a 200× LLM replacement study.
How to set CONTEXT_FILTER
export TYPESAFE_API_KEY=... # enables jev under default auto
export CONTEXT_FILTER=auto # auto | jev | keyword | embeddings | none
export JEV_MIN_SCORE=1.5 # 0-3 usefulness a chunk needs
export JEV_CONCURRENCY=64
export JEV_CHUNK_SIZE=1000
export JEV_MODEL=jev-latest
export KEYWORD_RELATIVE_THRESHOLD=0.5
export KEYWORD_MAX_RESULTS=25
export SIMILARITY_THRESHOLD=0.42 # embeddings mode only
export COMPRESSION_THRESHOLD=8000 # small page sets pass through
| Mode | How passages are chosen | Needs |
|---|---|---|
auto (default) | Jev if TYPESAFE_API_KEY is set, else keyword | nothing |
jev | Usefulness score ≥ JEV_MIN_SCORE, up to 10 | TYPESAFE_API_KEY |
keyword | BM25 with relative threshold, up to 25 | nothing |
embeddings | Cosine similarity vs query embedding | an EMBEDDING provider |
none | Full pages | nothing |
Fallback chain: jev → keyword on missing key, network error, rate limit after retries, or malformed response. embeddings → keyword if the embedding model cannot be built. Only embeddings mode constructs an embedding model for standard research.
Jev's rubric (from docs): unrelated; same topic but does not help; partially answers or supporting facts; directly answers with specific facts. The response score is the probability-weighted level from 0 to 3.
POST https://api.typesafe.ai/v1/systemone
Authorization: Bearer $TYPESAFE_API_KEY
{
"state": "<the 1,000-character chunk>",
"model": "jev-latest",
"questions": {
"usefulness": {
"type": "score",
"instructions": "How useful is this passage for answering the question: <sub-query>",
"criteria": [
"Unrelated to the question",
"Same topic, but does not help answer the question",
"Partially answers the question or gives useful supporting facts",
"Directly answers the question with specific facts"
]
}
}
}
For schema-level wiring outside this repo, use the Jev integration guide and how Jev's Score primitive works.
When to keep embeddings
Keep CONTEXT_FILTER=embeddings when:
- You already pay for an embedding endpoint and want no TypeSafe dependency.
- Queries are paraphrases of source language (BM25's light stemming will miss a lot of synonymy).
- You pass a custom
vector_store— that path is still embedding-backed. - You are A/B testing against a corpus where cosine already works and you do not want a 28-task web-research replay to decide for you.
Prefer keyword when you want zero extra APIs, fastest filter (~50× vs embeddings on the published p50), and the maintainers' claim that BM25 matched or beat embeddings on every measure in this replay — remembering the pairwise result is weaker than Jev's.
Prefer Jev when you have a TypeSafe key, accept network + ~1.7s median filter time, and want the thresholded usefulness behavior. It is the same "cheap structured check after retrieval" idea as pipeline checkpoints, not a new retrieval index.
Prefer none only if you have budget for ~65%+ higher writer cost and your report type is open-ended enough that more raw pages help. Deep/detailed research multiply sub-queries; the docs warn the token gap grows there.
Python 3.12+ and the rest of v3.7.0
The GitHub release lists Python 3.12+ as a change to note. Pin your runtime before pip install -U gpt-researcher.
Also in the tag: retriever plugins via the gpt_researcher.retrievers entry point (#2154), frontend WebSocket double-connect fix (#2159), multi-agent section overlap (#2158), self-hosted frontend assets without third-party CDNs (#2082), plus 25+ community retriever/scraper/deep-research fixes.
#2161 originally defaulted auto without a TypeSafe key to embeddings. #2162 replaced that fallback with keyword so no embeddings provider is required for standard research. If you upgraded from a mid-PR mental model, re-read auto.
Reproduce the 28-task replay
export OPENAI_API_KEY=... TAVILY_API_KEY=... TYPESAFE_API_KEY=...
python -m evals.context_filter.collect --out runs/
python -m evals.context_filter.replay --runs runs/ --out results/
python -m evals.context_filter.judge --results results/
Committed artifacts are scores and metrics in evals/context_filter/results/, not scraped pages or generated reports. Collect is documented at roughly $0.40 / 20 minutes in the README; the full configuration sweep was about $41 tracked plus $8–15 judging.
Honest limitations
- n = 28, one writer, one report type. Not a leaderboard for every research agent.
- LLM-as-judge precision and same-family pairwise judging. Self-preference is argued to apply equally, but it is still model-on-model.
- Jev needs TypeSafe uptime. Failures silently become BM25. That is good for availability; it is bad if you assumed every run was Jev-scored.
- Jev is not "never wrong." Schema-valid scores can still be bad relevance calls — the same distinction as launch coverage.
- Keyword is not Jev. Do not quote 73% for a default Docker box with no
TYPESAFE_API_KEY.
What this means for builders
If you run GPT Researcher as a research tool, v3.7.0 is a filter swap, not a RAG funeral. Set TYPESAFE_API_KEY if you want usefulness scoring; otherwise you already have a local BM25 default that the authors argue is good enough to drop mandatory embeddings. If you build agents, treat this as another production example of Jev as a typed gate after retrieval, next to routing and checkpoints — then measure on your corpus, because 28 web-research tasks will not match a private wiki.
Related on explainx.ai
- TypeSafe AI launches Jev
- How to wire Jev into agent routing
- Jev as cheap verification checkpoints
- How Jev works: RLCD and System One
- Jev speed and cost claims, fact-checked
- RAG vs MCP: complete comparison
- Jev waitlist is gone
Official: v3.7.0 release · Context Filter docs · PR #2161 · PR #2162 · evals/context_filter · TypeSafe Jev intro
Version numbers, CONTEXT_FILTER defaults, Python 3.12+ requirement, and eval tables reflect GPT Researcher v3.7.0 and the Context Filter docs as of September 29, 2026. The 28-task replay is the authors' own; rerun evals/context_filter before treating 73% as a property of your workload.
