explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • The four models, and what each is actually for
  • Darwin's numbers, against the field
  • Why Darwin isn't shipping broadly yet
  • What people are asking
  • Related reading on explainx.ai
← Back to blog

explainx / blog

OpenEvidence Darwin Hits 100% on MedQA — First AI to Do It

Medical AI, OpenEvidence, Model Releases, Benchmarks, Healthcare

OpenEvidence launched a four-model medical AI family on September 3, 2026. Darwin, its research-preview flagship, scored a perfect 100% on MedQA — the first AI in history to do so — and leads general models on every clinical benchmark tested.

Sep 4, 2026·7 min read·Yash Thakker
add explainx.ai
go deep
OpenEvidence Darwin Hits 100% on MedQA — First AI to Do It

OpenEvidence, the AI platform already widely used by clinicians for point-of-care medical search, released a four-model family on September 3, 2026 — and its flagship, Darwin, scored a perfect 100.0% on MedQA, which OpenEvidence describes as the first time any AI has done so. The launch is notable both for the benchmark number and for what OpenEvidence chose not to do with it: Darwin isn't shipping broadly. It's in application-only research preview while its safeguards get validated with partners, even as the three less powerful models in the family go live immediately.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.


TL;DR

table · 2 cols
QuestionDirect answer
What launched?A four-model OpenEvidence Model Family: Osler, Sackett, Snow (live now), and Darwin (research preview)
What's Darwin's headline result?100.0% on MedQA — the first AI, per OpenEvidence, to score perfectly
How does it compare to general frontier models?Ahead of Claude Fable 5, GPT-5.6 Sol, and Gemini 3.7 Flash on all four benchmarks OpenEvidence published
Can I use Darwin today?Only by application, in research preview — not a self-serve rollout
Can I use the other three models today?Yes — live now on web, iOS, and Android for verified clinicians
What differs between Osler, Sackett, and Snow?Primarily how long each model thinks and how deep it searches, not clinical-accuracy standard
Is the benchmark comparison independent?No — it's OpenEvidence's own published comparison, pending independent replication

The four models, and what each is actually for

OpenEvidence frames the family by clinical moment rather than raw capability tier — a deliberate choice, since it signals these are meant as fit-for-context tools rather than a single model dialed up or down:

table · 3 cols
ModelFramed forStatus
Osler"The hallway" — quick, in-the-moment lookupsLive now
Sackett"The consult" — deeper single-patient reasoningLive now
Snow"The tumor board" — multi-specialist reviewLive now
DarwinFlagship, most capableResearch preview, application-only

OpenEvidence's own description of the split is precise: "Every model in the family is held to the same standard of clinical accuracy. What varies is time: how long a model thinks, and how deep it searches." That's a meaningfully different design philosophy than most consumer AI product tiers, which typically trade off capability, not just latency, between a fast and slow option.

Darwin's numbers, against the field

OpenEvidence published a direct head-to-head comparison of Darwin against three general-purpose frontier models across four medical benchmarks:

table · 5 cols
BenchmarkOpenEvidence DarwinClaude Fable 5GPT-5.6 SolGemini 3.7 Flash
MedQA (USMLE-style, accuracy)100.0%99.7%99.1%99.2%
MedXpertQA (specialty boards, accuracy)72.8%64.3%57.6%65.1%
HealthBench Pro (open-ended rubrics)82.7%70.6%66.9%59.0%
NOHARM (harm-weighted F1)87.2%72.5%74.0%71.5%

Two things stand out in that table. First, the MedQA gap between Darwin and the general models is small in absolute terms — all four models are already clustered near the top of that benchmark's ceiling, which is a familiar pattern as a specific benchmark saturates across the field. Second, the gap widens substantially on the harder, more open-ended benchmarks — MedXpertQA, HealthBench Pro, and NOHARM — where Darwin's lead over the closest general model runs 8-24 percentage points depending on the test. That pattern is consistent with what you'd expect from a domain-specialized model versus general-purpose frontier models on a domain's hardest, least benchmark-saturated tasks.

Why Darwin isn't shipping broadly yet

OpenEvidence's own framing for restricting Darwin is unusually direct for a launch announcement: "Capability at that level cuts both ways in medicine, so Darwin is in research preview, by application only, while its safeguards are validated with partners like RareDiseases."

That's a notable departure from the more common pattern of shipping the most capable model first and adding restrictions reactively. Restricting the most capable model in a medical-AI family, specifically because of its capability, while shipping the less powerful models immediately, is a legible statement about where OpenEvidence sees the actual risk surface — not in whether the model gets clinical facts right (Darwin's benchmark scores are the strongest in the family), but in what happens when a highly capable model's output gets trusted without the same guardrails as the more constrained models.

What people are asking

Is OpenEvidence's benchmark comparison trustworthy? It's a first-party comparison, published by the company whose model is winning it, against three benchmarks (MedXpertQA, HealthBench Pro, NOHARM) that are less universally standardized than MedQA. That doesn't make the numbers wrong, but it does mean independent replication — by researchers or a third-party evaluator without a commercial stake — hasn't happened yet as of this post, and is worth watching for before treating the comparison as settled.

Are Claude Fable 5, GPT-5.6 Sol, and Gemini 3.7 Flash "bad" at medicine? No — all three score above 99% on MedQA in OpenEvidence's own comparison, which is itself a strong result for general-purpose models never specifically trained on OpenEvidence's clinical dataset or fine-tuning pipeline. The comparison illustrates the value of domain specialization on the harder, more open-ended benchmarks specifically, not a broad medical-AI capability gap.

Does a high MedQA score mean an AI is safe to use for real clinical decisions? No, and this is exactly the distinction OpenEvidence's own Darwin restriction seems to acknowledge. MedQA measures knowledge recall against board-style questions; NOHARM's harm-weighted F1 and HealthBench Pro's physician-written open-ended rubrics are closer proxies for real clinical judgment, and even there, a benchmark score is not equivalent to validated safety in live clinical use — which is precisely why Darwin remains gated behind partner validation rather than shipping on benchmark strength alone.

How is this different from a general chatbot answering medical questions? OpenEvidence is built specifically for point-of-care clinical use by verified clinicians, with models tuned and evaluated against medical-specific benchmarks and rubrics, rather than a general-purpose assistant that happens to also answer medical questions. The four-model family structure — differentiated by clinical context and thinking depth rather than a single one-size-fits-all model — reflects that specialization.


Related reading on explainx.ai

  • AI Benchmarks: Complete Guide — background on how benchmarks like MedQA and HealthBench Pro are constructed and their limits
  • Claude Protein Design and Analytical Chemistry — Anthropic's own domain-specialized scientific AI work, for comparison
  • Claude Team Plan for Scientists: 10,000 Seats — general frontier-model adoption in scientific and clinical research settings
  • AI Drug Discovery: Clinical Evidence and Benchmarks — a broader look at how AI clinical-benchmark claims hold up against independent scrutiny
  • Moderna and Merck: AI-Designed mRNA Cancer Vaccine Phase 3 — another example of specialized medical AI moving from benchmark to real clinical application
  • Gemini 3.7 Flash vs Grok 4.6 vs Sonnet 5 vs GPT-5.6 Comparison — where the general models cited in OpenEvidence's comparison stand outside medicine
  • What Are AI Agents? Complete Guide — background context on how domain-specific AI products differ from general-purpose assistants

Official source: OpenEvidence Model Family launch announcement

Benchmark figures in this post are OpenEvidence's own published comparison as of the September 3, 2026 launch and have not yet been independently replicated. Darwin's availability is restricted to research-preview, application-only access as of this post's publication.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 23, 2026

Claude Opus 5.5 Is Live: Every Benchmark, Price, and Reaction That Matters

Anthropic's first release since calling for "pacing the frontier" claims Fable 5.1-level performance at 40% lower cost, a rewritten communication style, and the strongest safety scores of any Claude model to date. Here is every number from the announcement, plus what developers who switched from Opus 5 are actually reporting in the first hours of real usage.

Sep 23, 2026

GPT-6 Sol and Luna Launch: 50% Price Cuts and Where They Actually Land

OpenAI cut API prices 50% on its mid-tier and small models the same day Anthropic launched Opus 5.5. GPT-6 Sol and Luna bring GPT-6 Astra's training advances to cheaper, faster models — but independent evaluators found Luna actually regressed slightly on coding benchmarks even as its price fell. Here's every number, and what developers found once they actually ran them.

Sep 7, 2026

"GPT-6 Pro" Reportedly Spotted in ChatGPT: What We Actually Know

Unverified reports circulating around September 6-7, 2026 describe a "GPT-6 Pro" label surfacing in the ChatGPT interface, alongside a separate claim from a prominent AI industry figure that a model called "Max" is the best model for math. Neither claim comes from an official OpenAI announcement. Here's a sober read on what a "Pro" tier would typically mean, why math leadership claims are especially contested right now, and how to verify a new tier yourself instead of trusting a screenshot.