explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR
  • "Absence of evidence" is the precise claim — read it precisely
  • The fair objection: the clock has not run out yet
  • Phase II is the only scoreboard that counts
  • The benchmark lesson: when does a benchmark predict the real world?
  • What the tools actually do in a lab today
  • Attribution bias: "when we win we're England, when we lose we're MCC"
  • "Doing what should be done" versus "doing what can be done"
  • The ladder: from ligand to drug, from demo to production
  • What to take away if you build with AI
  • Verdict
  • Related on explainx.ai
  • Primary sources
← Back to blog

explainx / blog

AI Drug Discovery Has an Evidence Problem — And a Benchmark Lesson for Everyone Else

A Nature Reviews Drug Discovery review finds evidence of clinically relevant AI impact is "disappointingly limited" — and explains when benchmarks stop predicting the real world.

Aug 16, 2026·19 min read·Yash Thakker
AI in HealthcareDrug DiscoveryAI BenchmarksAI EvaluationResearch
go deep
AI Drug Discovery Has an Evidence Problem — And a Benchmark Lesson for Everyone Else

The volume of AI-in-drug-discovery announcements has been extraordinary for about four years. The volume of demonstrated clinical impact has not kept pace — and a new review in Nature Reviews Drug Discovery says so bluntly.

The line that made the paper travel: "Although a wide variety of AI methods have been developed, applied and benchmarked, evidence of their clinically relevant impact is, so far, disappointingly limited."

Chemist Derek Lowe covered it on August 10, 2026 in In the Pipeline, Science.org's long-running medicinal chemistry column, and it climbed to 80 points and 42 comments on Hacker News the same week. Lowe's framing of why this particular paper is worth reading: it comes from "a number of people who know what they're talking about in the field and who (importantly) are not trying to sell you anything, either."

This post is not a takedown of AI in biology. The paper itself refuses that reading, and so should we. What makes it worth your time even if you never touch a molecule is its argument about benchmarks — specifically, why benchmarks translated cleanly to real-world performance in image and speech recognition and do not necessarily translate in domains where the labels themselves carry confounders. That is a portable idea, and it is the reason this post exists on a blog about building with AI rather than a pharma trade publication.

TL;DR

table · 2 cols
QuestionAnswer
What did the paper find?Evidence of clinically relevant AI impact in drug discovery is "disappointingly limited" so far.
Is that the same as saying AI does not work?No. The authors call it an "absence of evidence", explicitly not an "evidence of absence."
Is there a fair defence?Yes — drug development cycles exceed the window in which these tools have been usable. Measuring impact honestly takes years more.
Where would real impact show up?Phase II success rates. Everything upstream is a rounding error against trial cost.
Why should a non-pharma builder care?The benchmark critique: benchmarks predict reality only when labels are reproducible and condition-independent.
What do the tools actually do today?Make existing work much faster. Practitioners report no genuinely novel outputs yet.
What is the trap in reading success stories?Attribution bias — wins get credited to AI, losses quietly go unattributed.
What should the field do instead?Shift from "modelling data that is readily available" to generating the data that would actually settle the question.
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

"Absence of evidence" is the precise claim — read it precisely

The paper's own summary sentence is careful in a way that most coverage of it will not be:

"It can be concluded that we are still at a stage with an 'absence of evidence' (and not necessarily an 'evidence of absence') when it comes to the translation of AI into clinically relevant impact."

Those are different claims and the difference matters. "Evidence of absence" would mean somebody ran the study, and the AI-assisted programmes did no better. "Absence of evidence" means the study that would settle it has not been run, or has not had time to read out.

That distinction is the same one we apply across the research reality-check series on explainx.ai: a model beating a benchmark is an engineering result, and a change in patient outcomes is a clinical result, and no amount of the former automatically becomes the latter. The Berkeley debate over AI and cancer-risk prediction turned on exactly this gap.

Lowe's own reading is consistent with the paper's hedging. On decision quality specifically — target selection, disease-area selection, candidate selection, trial design — he writes:

"I go around telling people that these are some of the areas that current AI techniques are least equipped to help us out with, but I always need to emphasize that word 'current'. There's no fundamental reason why these things can't be improved, but such improvements are going to be slower and more expensive than claimed by many of the people who send out all those press releases."

The complaint is about the slope of the claims, not the possibility of the thing.

The fair objection: the clock has not run out yet

Before the critique lands, it deserves its strongest rebuttal, which came from the Hacker News thread rather than the paper.

A commenter pushing back on the framing made a point that is hard to argue with: drug development timelines typically exceed the interval in which these technologies have been available, or at least effective. A programme that started using modern generative and structure-prediction tooling in 2022 or 2023 is, at best, somewhere in early clinical development now. Phase II readouts for anything genuinely AI-shaped are mostly still ahead of us.

That is not a rhetorical dodge. It is the mechanical reason the evidence base looks thin, and it is a large part of why "absence of evidence" is the right phrase. Any honest scoreboard for this question is a 2030s scoreboard.

The same commenter also objected to the implication that drug discovery scientists do not think carefully about why they use their tools — they think about it constantly, and AI is simply another tool in a field that has absorbed many. That is worth holding onto: the critique is of an evidence culture, not of the competence of the people doing the work.

Phase II is the only scoreboard that counts

Here is the part that reframes the entire economic argument, and it is the part most AI-in-bio pitches skip.

Lowe: "All that stuff in early-stage discovery work is a roundoff error in time and money compared to the clinical trials, and Phase II is where you first run into the really large cross-section for failure in the clinic."

Phase II is where a compound first meets the question of whether it does anything useful in patients with the disease. It is where the attrition is brutal and where the spending has already become enormous. A method that halves the time to identify a hit series has saved a slice of a small number. A method that raises Phase II success rates has changed the industry.

Lowe's historical note is the sharp end of it: "pre-AI there have been all sorts of Capitalized Revolutionary Acronymic Protocols announced in various companies to try to get those success rates up. None of this has had much impact in the end. The AI era has just been more of the same, as far as I can tell so far."

The transferable version of this for any builder: find the stage in your pipeline where the cost and failure actually concentrate, and ask whether your AI intervention touches it. Speeding up a cheap step feels like progress and shows up beautifully in a demo. It does not move the number your business is actually judged on. This is the same trap as optimizing the fastest part of a system because it was the easiest part to instrument.

The benchmark lesson: when does a benchmark predict the real world?

This is the section worth the price of admission for readers who will never run an assay.

The paper's argument, quoted at length because the phrasing is precise:

"Another key problem is that the nature of the data at hand in drug discovery — its confounding factors, conditionality and epistemic opacity — may not be understood by everyone aiming to model such data, and hence the limitations of labelling data and using labelled data are not understood either. As a consequence of 'believing the labels', there is a considerable knock-on effect on benchmarks that are being used in the field. Benchmarks have been crucial to provide some estimate of progress in AI, given that otherwise no quantitative comparison of novel approaches to the existing 'state-of-the-art' (SOTA) would be possible. In some areas, such as image recognition and speech recognition, benchmarks by and large translated to better performance in real-world use cases. However, in the area of data used in drug discovery, this has to be viewed with caution."

Sit with the comparison the authors are drawing. Benchmarks did work for image recognition. ImageNet progress really did become better cameras, better search, better accessibility tooling. Speech benchmarks really did become usable transcription. So the failure mode being described is not "benchmarks are bad." It is that benchmarks have a validity precondition, and image and speech happen to satisfy it while assay data does not.

What image and speech had that assay data lacks

A photo labelled "cat" is labelled with a fact about the world that any independent human can regenerate, from the same input, and get the same answer. The label does not encode the camera, the photographer, the day, or the lab. Ground truth is directly observable and condition-independent.

An assay label is not like that. "IC50 = 12 nM" is not a property of the molecule. It is a property of the molecule plus the assay format, the cell line, the buffer, the incubation time, the readout chemistry, the plate, the operator, and the week. Run the same compound in two labs and you can get numbers that differ by more than the effect size you were trying to detect. The measurement conditions are inside the label, and they are usually not recorded as features the model can see.

Lowe's blunter version is that doing meaningful machine learning on the accumulated pile of historical assay data "is very likely impossible, due to the huge numbers of variables in the assay conditions and indeed, in biological systems themselves. We don't know how to clean this stuff up and categorize it for real ML/AI, and honestly, we don't even know if it can be."

A five-question diagnostic you can run on any benchmark

Generalize the argument and you get a test that works on evaluation sets far outside biology — LLM leaderboards, retrieval evals, recommender offline metrics, fraud models, medical imaging, anything.

table · 3 cols
QuestionBenchmark is predictive when…Benchmark degrades when…
1. Reproducibility of the labelAn independent party could regenerate the label from scratch and get the same valueThe label is a measurement whose value depends on how it was produced
2. Proxy distanceThe label is the outcome you care aboutThe label is several causal steps upstream of the decision it feeds
3. Hidden conditionalityMeasurement conditions are constant, or recorded as featuresConditions vary across rows and are not in the data
4. Error structureLabel noise is random with respect to the inputNoise correlates with something learnable — the lab, the vendor, the era, the annotator pool
5. Downstream cost asymmetryThe benchmark's decision and the expensive decision are the same decisionA cheap decision is being scored while an expensive one is being implied

Fail question 1 or 3 and your model can learn the provenance instead of the phenomenon — it identifies which lab produced a row and predicts that lab's house bias, scoring superbly on a held-out split drawn from the same labs and collapsing on new ones. That is not overfitting in the ordinary sense; the held-out set is contaminated by a variable nobody wrote down.

This failure has an LLM-native form too. A preference score from a rater pool encodes that rater pool. A coding benchmark's pass/fail encodes the coverage of its test suite, not correctness. A "helpfulness" label encodes the instructions the annotators were given on the day. Our guides on how to read an AI benchmark and Goodhart's law in 2026 benchmark contamination cover the leakage and gaming side of this problem; the paper adds the deeper one, which is that a benchmark can be perfectly uncontaminated and still measure the wrong thing because its labels were never the thing.

The practical upshot is not "distrust all benchmarks." It is: before you optimize against an eval, write one paragraph describing how each label was physically produced. If you cannot write that paragraph, you do not know what your leaderboard position means. This is the same discipline that separates a real internal eval from a vanity one — see our complete guide to AI benchmarks and the fact-check of 2026 benchmark claims.

What the tools actually do in a lab today

The most useful comment in the Hacker News thread came from a structural biologist at a mid-sized biotech, and it is worth quoting almost in full because it is the sober middle position that neither the press releases nor the backlash will give you:

"I use AI tools daily. They make accomplishing the same things I was able to accomplish before quite a lot faster and easier. They don't help me magically accomplish new things that I couldn't previously. For example, it helps me install academic software, debug things. It helps me take a large dataset and write scripts to ask questions. It helps me go through experiment drafts to see if I'm missing things. It helps me remember obscure formulas I use every 6 months. It has not, at least in my experience, come up with anything truly novel."

Their concrete example: AlphaFold is genuinely great for producing a starting model for a chimeric fusion. Work that meant one to two hours of fumbling through PDB or CIF files is now a quick prompt. That is a real, repeated, compounding saving. It is also emphatically not a new scientific capability.

That framing — faster at what you could already do, not new things you could not — matches what we found looking at Anthropic's Claude for Science workbench and at the rare disease research grants around the Monarch Initiative. It is also the honest ceiling on most agentic tooling in specialist domains right now, including the drug-discovery agent stacks being assembled around NVIDIA BioNeMo.

The downside the same practitioner named

One line deserves its own heading because it describes a cognitive failure mode that almost nobody quantifies:

"It's far too easy to leave one branch of reasoning now and jump to another whenever progress gets difficult."

When switching costs collapse, so does persistence. The friction of a hard path used to be what kept you on it long enough to get through. If the model can instantly generate a plausible alternative approach, the moment a line of reasoning becomes genuinely difficult — which is often exactly the moment before it becomes productive — you bail. This generalizes directly to software: the ability to instantly try a different library, a different architecture, a different prompt, is not free.

The upside they also named

The counterweight, from the same person:

"I've been putting myself to sleep at night by just asking it questions about expectation maximization and Bayesian statistics. This has seriously boosted my understanding of cryo-EM alignment algorithms in a way I couldn't do in grad school because there was no professor that understood enough to help me when I got stuck reading literature. So it's a double edged sword for sure."

That is the single most defensible claim anyone makes for these systems: an available tutor for hard technical material in a niche where no human expert was available to you. It is also, not coincidentally, what explainx.ai is built around — the tutoring loop is more reliably transformative than the discovery loop, at least today.

Attribution bias: "when we win we're England, when we lose we're MCC"

Lowe names a bias that will be instantly recognizable to anyone who has watched a methodology get adopted in their own field:

"There's a natural tendency to ascribe successful projects to keen insights, bold leadership, and relentless hard work, while unsuccessful ones, umm, is there some reason why you'd want to talk about those?... Any project that had any AI aspect to it likely succeeded because of that, don't you know, and the other projects that failed even though they had the Digital Laying On of Hands, well, moving right along..."

Mechanically: successes with an AI component get attributed to the AI component. Failures with the identical AI component get attributed to biology, funding, the market, or nothing at all. Nobody has to lie for this to produce a wildly inflated apparent hit rate — the denominator just quietly evaporates.

You have seen this exact structure elsewhere. It is why "we adopted [framework] and shipped 3x faster" case studies are worthless without the projects that adopted it and died. It is why every internal AI-productivity claim needs a pre-registered denominator, and why the Stanford AI Index is more useful than any single company's retrospective. If you are the one reporting, the fix is cheap and unpopular: decide before the project starts what you will count as an AI-attributable outcome, and publish the failures under the same label.

"Doing what should be done" versus "doing what can be done"

The paper's prescription is a direct challenge to how the field currently allocates effort:

"...the focus of AI in drug discovery must shift from doing what can be done — such as modelling data that is readily available, but that is unlikely to move the needle — to doing what should be done, even if this requires, for example, substantial data generation..."

That is the right answer and an expensive one. Lowe's reaction is the realistic counterweight: "It's a worthy goal, but I think that many involved in this work might be thinking, even unconsciously, 'You first'." Generating purpose-built, condition-controlled data at scale is a large capital commitment with murky odds, and it implicitly means telling investors and the press that things will take longer than they are used to hearing.

Note what this is really a critique of: data convenience as a research strategy. The available dataset determines the paper, the paper determines the benchmark, the benchmark determines what counts as progress, and none of that chain was ever pointed at the decision that matters. This is a close cousin of the argument that AI is flattening scientific discovery by concentrating effort on whatever is already well represented in the training corpus.

The same critique applies to your own work with almost no translation. If your evaluation set exists because it was easy to assemble rather than because it represents the decision you care about, you are optimizing convenience.

The ladder: from ligand to drug, from demo to production

Lowe's escalation test for a compound is worth memorizing as a general-purpose pattern:

"if you say you have a good compound in a cloned isolated protein, I want to know how it is in cells. If it works in cells, tell me about the rodents. If it works there, I want to know about two-week dog tox, and if it gets through that, then why aren't you putting it into human patients now?"

Every rung is a different system with different failure modes, and success on rung N carries surprisingly little information about rung N+1. The isolated-protein result is real; it is just not a drug.

The software version writes itself. A prompt that works in a playground is not a feature. A feature that works on your ten favourite examples is not a product. A product that works for internal users is not one that survives adversarial ones. A demo that works on the happy path is not a system that works at 3am under load with malformed input. Each rung is a genuinely different environment, and the discipline is refusing to let a result on one rung be reported as a result on a higher one.

What to take away if you build with AI

  1. Locate the expensive decision. If your AI intervention does not touch the stage where cost and failure concentrate, it is a demo improvement. Phase II is pharma's; find yours.
  2. Interrogate your labels before your model. Write down how each label was physically produced. If the production conditions vary and are not features, your benchmark is measuring provenance.
  3. Assume attribution bias in every success story, including your own. Pre-register what counts as attributable, and count failures under the same banner.
  4. Distinguish "faster at what I could do" from "new capability." Both are valuable; only one justifies the claims usually attached to it.
  5. Watch for the abandonment failure mode. Cheap branch-switching quietly erodes the persistence that difficult problems require.
  6. Keep "absence of evidence" and "evidence of absence" apart. Most arguments about AI capability are people conflating the two in opposite directions.

Verdict

The paper is not an obituary. It is an audit finding: the field has produced an enormous quantity of methods, applications and benchmark results, and has not yet produced the clinical evidence that would justify the surrounding narrative. Given drug development timelines, that evidence could not plausibly have arrived yet — which makes the honest verdict "unproven," and makes the loud verdict "revolutionary" premature by roughly a decade.

For everyone else, the durable takeaway is the benchmark argument. Benchmarks earned their reputation in image and speech recognition, domains where the label is a reproducible fact about the world. In domains where the label is a measurement that encodes the conditions that produced it, a leaderboard can improve indefinitely while the real-world decision it supposedly informs does not budge. That is not a pharma problem. That is a description of most interesting AI applications outside vision and speech.

Before you trust a number, find out how it was made.

Related on explainx.ai

  • Can AI cure cancer? A research-backed reality check
  • Can AI prevent the next pandemic?
  • How to read an AI benchmark
  • Goodhart's law and benchmark contamination: the 2026 receipts
  • The complete guide to AI benchmarks
  • Claude for Science: the AI workbench for scientists
  • Anthropic's rare disease research grants and the Monarch Initiative
  • Long-read genome sequencing for rare disease diagnosis
  • Is AI flattening scientific discovery?
  • John Jumper leaves Google DeepMind for Anthropic
  • NVIDIA BioNeMo agent toolkit for drug discovery
  • Demis Hassabis on the frontier AI framework
  • Stanford AI Index 2026: the takeaways that matter

Primary sources

  • Derek Lowe, "So How Is AI Drug Discovery Doing, Really?", In the Pipeline, Science.org, August 10, 2026 — science.org/blogs/pipeline
  • The review under discussion, Nature Reviews Drug Discovery, 2026 — nature.com/articles/s41573-026-01496-2
  • Hacker News discussion thread, August 2026 (80 points, 42 comments) — practitioner commentary quoted above

Quotes are reproduced from the sources listed above as published on August 10 and 11, 2026. Journal access, comment counts, and column URLs are accurate as of August 16, 2026 and may change.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

May 2, 2026

AI Benchmarks in 2026: The Complete Guide to MMLU, GPQA, SWE-bench, and Beyond

AI benchmarking in 2026 has reached a critical inflection point. Traditional benchmarks like MMLU and HellaSwag are saturated above 88% and 95%, while frontier models cluster within statistical noise. This comprehensive guide covers every major benchmark category—from language understanding to agent evaluation—the 37% lab-to-production gap, benchmark gaming vulnerabilities, and what actually matters for production AI systems.

Aug 14, 2026

Dognosis: Canine + Bayesian AI Hits 90.8% Cancer Detection Sensitivity

Indian startup Dognosis published a Phase II study in the Journal of Clinical Oncology showing 90.8% sensitivity and 0.962 AUC for multicancer breath detection — not from the dogs alone, but from a Bayesian model that fuses their individual indications into one calibrated score. Here's the methodology, the numbers, and what still needs to happen before this is a screening test.

Aug 6, 2026

Goodhart’s Law Comes for Every Benchmark You Trust: The 2026 Receipts

Public AI benchmarks aren't just theoretically gameable — 2026 research proves it with numbers. GSM1k found up to 13% accuracy drops on fresh math problems, an MMLU audit found a 6.49% error rate, and the Leaderboard Illusion paper caught Arena's best-of-N submission gaming with a controlled experiment. Here is the quantitative evidence behind Goodhart's law in AI evaluation.