An ICML 2026 Position Paper Track submission with a deliberately provocative title, "Position: LLMs Can't Jump," by Tom Zahavy, has been quietly circulating in AI discussion — including a citation dropped approvingly into an unrelated Hacker News thread about the map/territory problem in agentic coding. The paper was accepted (Accept, regular) after a full public peer-review cycle on OpenReview, and that review thread is arguably as interesting as the paper itself — three reviewers, a real rebuttal, and an author correction on overclaiming, all visible to anyone who wants to read how a claim about AI's limits actually gets pressure-tested.
The core argument: generative AI has mastered induction — statistical pattern matching across enormous training corpora — and is rapidly closing the gap on deduction — formal, step-by-step proof from given premises. But it has no working mechanism for abduction: generating a genuinely new explanatory hypothesis, or axiom, that didn't already exist in some form in the training data. The paper's proposed fix isn't a bigger LLM. It's physically grounded world models.
TL;DR
| Question | Short answer |
|---|---|
| What's the paper's core claim? | LLMs are strong at induction and improving fast at deduction, but structurally lack abduction — the creative leap to a new axiom with no precedent in training data |
| What's the case study? | Einstein's equivalence principle (acceleration mimics gravity), the intuitive jump behind general relativity, grounded in physical/sensory intuition an LLM doesn't have |
| Is this proven? | No — it's a position paper. The authors softened their conclusion from "confirms" to "suggests" during review, at a reviewer's request |
| What's the proposed bridge? | Physically-consistent, multimodal, action-controllable world models (like Genie) rather than LLMs, because they let an agent simulate and "feel" a scenario rather than just predict the next frame |
| Did reviewers buy it? | Yes — all three ICML reviewers rated Accept or Borderline Accept, and all raised their scores after the rebuttal |
| Does it mean AI can't do science? | No — it argues one specific creative step is out of reach for LLMs alone, not that AI has no role in scientific discovery |
| Where was this published? | ICML 2026 Position Paper Track, publicly reviewed on OpenReview |
Where does the title come from?
The title is a deliberate double entendre, and both readings matter for understanding the argument.
First, the colloquial sense: LLMs can't make the creative conceptual "leap" — the kind of jump a scientist makes when a genuinely new idea appears that wasn't a logical consequence of what came before.
Second, the literal sense: LLMs have no body. They have never felt acceleration, never lost their balance, never sat in a bathtub and noticed water rise as they got in. No sensory apparatus means no jumping, in the most physical reading of the word — and the paper's argument is that this isn't just a cute pun, it's the actual mechanism behind the first claim.
Einstein's own account of how discovery works
The paper's framework doesn't come from modern cognitive science — it comes from Einstein himself, in a 1952 letter to his friend Maurice Solovine. Einstein sketched scientific discovery as a cycle with two very different kinds of moves:
- E → A: An intuitive, non-logical "jump" from raw sensory experience (E) to a set of axioms (A). This step is not derivable by any formal rule. It's a creative act.
- A → theorems: Once the axioms are in hand, logical deduction takes over — theorems and predictions follow from the axioms by rigorous, checkable proof, eventually tested back against experience.
That second step — deduction from established premises — is exactly the kind of formal reasoning that language models, especially reasoning-tuned ones with extended chain-of-thought, have gotten dramatically better at over the past two years. The paper doesn't dispute this. It argues the opposite is true of the first step, the E → A jump, and that this asymmetry is being systematically underrated because deduction is the visible, benchmarkable part of "reasoning," while abduction rarely shows up in a leaderboard.
The case study: the equivalence principle
The paper's central example is Einstein's own path to general relativity, specifically the equivalence principle — the insight that acceleration and gravity are locally indistinguishable. Standing in an accelerating elevator with no windows feels identical to standing still in a gravitational field. There's no experiment you can run inside the elevator to tell the two apart.
This idea did not fall out of existing physics by deduction. Newtonian mechanics didn't imply it; nothing in the symbolic apparatus of physics at the time necessitated it. Einstein has described arriving at it via a specific, embodied thought experience — imagining himself in free fall, imagining the sensation of weightlessness. The paper's claim is that this is abduction in its purest form: a new axiom generated from lived, physical, sensory experience that has no antecedent in any symbolic system.
An LLM, the paper argues, can flawlessly deduce every consequence of general relativity once given the field equations. What it cannot do, on this account, is generate the equivalence principle itself from scratch, because it has never accelerated, never fallen, never felt weight — it has only ever seen the word "gravity" appear next to other words in text.
The second example added during review: Archimedes in the bath
One reviewer (vzJN) pushed the authors for a simpler example than general relativity — something more intuitive for readers without a physics background. The authors' response, added during the rebuttal period, was the Archimedes "eureka" story: stepping into a bath, noticing the water level rise in proportion to the volume of his submerged body, and abducing the principle that buoyant force equals the weight of displaced fluid.
The paper frames this as manipulative abduction — a hypothesis generated not from staring at data or manipulating existing symbols, but from physically sensing one's own body interacting with the world in real time. Like the equivalence principle, buoyancy isn't a deduction from prior axioms, and it isn't a statistical pattern extracted from a corpus of prior bath observations. It's a new causal relationship proposed on the spot, from embodied sensory feedback.
Isn't this just a prompting problem?
This was reviewer m94K's central push during discussion: if you give a model better context, better scaffolding, or a more clever prompt, doesn't that unlock the same kind of leap? It's a fair question, and one worth taking seriously given how much of the recent context engineering conversation has been about exactly this — that a lot of apparent capability gaps are really specification gaps.
The authors' rebuttal draws a sharp line: prompting and context engineering improve reasoning within an already-existing symbolic space. A better prompt can help a model find a proof it already had the pieces for, structure a deduction more reliably, or retrieve a relevant analogy from training data. What no amount of prompting supplies is sensory grounding — the actual physical experience that generated Einstein's or Archimedes' hypothesis in the first place. You can't prompt your way into having a body. The paper's position is that abduction of genuinely new axioms is bottlenecked on grounding, not on elicitation technique, and those are different problems with different fixes.
So is the paper right? What's the honest read
It's worth being precise about what kind of claim this is, because the paper's own review thread makes a point of it. ICML's Position Paper Track exists specifically for argued theses rather than proven empirical results — a position paper is judged on whether it makes a clear, well-supported, falsifiable argument, not on whether it settles the question.
Reviewer Ldv4 flagged exactly this distinction during review: the original submission's conclusion stated that the paper's analysis "confirms" LLMs cannot perform abduction. That's a strong empirical claim for a paper built on two historical case studies and a conceptual argument, not a controlled experiment. The authors agreed and revised the language to "suggests" — a small edit that matters, because it's the difference between an evidence-backed hypothesis and an overclaimed result. That's the epistemic posture worth carrying into how you read the paper: as a serious, well-argued position that two carefully chosen historical examples are consistent with, not as a proof that no future architecture will ever manage abductive leaps.
It's also worth noting what the paper does not claim. It doesn't say AI has no role in scientific discovery — plenty of recent AI-for-science work shows models accelerating hypothesis search, literature synthesis, and experimental design. The claim is narrower and more specific: the one creative step of generating a genuinely novel axiom with no symbolic precedent is, on current evidence, out of reach for a system that only ever sees text.
The proposed bridge: world models, not bigger LLMs
The paper's constructive half is where it gets concrete about what might close the gap, and the answer is a shift away from language models entirely, toward physically-consistent, multimodal world models.
The distinction the authors draw is precise and worth sitting with, because it's easy to conflate "video generation" with "world simulation" and they are not the same claim:
| Model type | What it optimizes for | Example | Still induction? |
|---|---|---|---|
| Passive video generator | The statistically most likely next frame given the prompt | Veo | Yes — still pattern matching over training data |
| Action-controllable world model | A physically consistent next state, given an agent's active intervention | Genie | No — provides a synthetic laboratory for testing physical hypotheses |
The argument is that a model like Veo, however visually convincing, is still doing induction — it's predicting what's most likely to come next based on what it's seen before, the same fundamental operation as a language model predicting the next token, just in pixel space instead of token space. An action-controllable world model like Google DeepMind's Genie is a different kind of system: it lets an agent actually intervene, take an action, and observe a physically consistent consequence. That's closer to what the paper calls a "synthetic laboratory" — a space where an agent could, in principle, run something analogous to Einstein's thought experiment or Archimedes' bath, testing a hypothesis against simulated physical feedback rather than retrieving the most probable next word.
This is the same distinction explainx.ai has covered from the infrastructure side in the NVIDIA Cosmos 3 guide — Cosmos 3 explicitly separates a Reasoner surface (text in, text out) from a Generator surface that includes action-conditioned forward dynamics, precisely because "predict the next frame" and "simulate the consequence of an action" are different capabilities with different training objectives. The same split shows up in NVIDIA's Alpamayo 2 Super, where a world model underneath a reasoning layer is used to generate physically grounded training signal a pure language model couldn't produce on its own.
The authors also point to a concurrent, non-LLM approach worth flagging for anyone chasing this thread further: a neuro-symbolic system for abductive reasoning that sidesteps the whole training-data-contamination question by not routing through a language model at all. It's cited in the paper as an interesting alternative path rather than a competing claim, and it underscores that the "abduction gap" is being attacked from more than one architectural direction right now, not just the world-model one this paper focuses on.
What people are asking
Does this mean LLMs are useless for science? No. The paper's target is one specific step — generating a wholly new axiom from scratch. LLMs remain useful for literature synthesis, deriving consequences of known theories, generating and checking proofs, and exploring large hypothesis spaces once a human (or, per the paper's proposal, a grounded world model) has supplied the seed idea.
Could a multimodal LLM with video/audio input already do this? The paper's distinction isn't about input modality alone — it's about whether the model can actively intervene and observe consequences (Genie-style) versus passively predict the next observation (Veo-style, even with rich multimodal conditioning). A multimodal LLM that only ever predicts the next frame is, by this paper's framing, still doing induction, just in a richer output space.
Is "abduction" a well-established term, or is the paper inventing a category? It's a well-established term from the philosophy of logic — usually credited to Charles Sanders Peirce as "inference to the best explanation," the third leg alongside induction and deduction. The paper is applying an existing philosophical distinction to current AI capability, not coining new vocabulary.
How does this connect to the "map is not the territory" argument that cited it? The Hacker News thread referencing this paper was making a related but distinct point in the context of agentic coding — that an agent works from the map (prompt, context, instructions) it's given, not the underlying territory. The Zahavy paper's abduction argument is a deeper, more fundamental version of the same intuition: an LLM can reason expertly within a given symbolic space (a map), but has no mechanism for generating a genuinely new map from raw, ungrounded experience (the territory itself). See explainx.ai's coverage of Thariq's map/territory field guide for the agentic-coding framing.
Update — September 4, 2026: A month after ICML acceptance, the paper broke out on X after a viral thread summarized it as proof AI "can never make real scientific discoveries." That framing overshoots what a position paper claims (see above), but it arrived alongside a second, independently authored paper making a related argument from a different angle, plus a fresh round of reader backlash against AI-written text — both covered below, along with explainx.ai's own read on what's compounding the abduction gap.
A second paper, different angle, same conclusion: Visual General Intelligence
Days after the "LLMs Can't Jump" discourse spread on X, a 14-author white paper titled "Visual General Intelligence: A White Paper" started circulating with overlapping authorship — contributors from AIST, Oxford's Visual Geometry Group, Google DeepMind, Stanford, CMU, Harvard, Princeton, Imperial College London, NYU, Cambridge, and OpenAI, among others. It doesn't cite Zahavy's paper, but it argues a structurally similar point from a completely different research tradition: computer vision, not philosophy of science.
The paper's claim is that language is a late, compressed, human-abstracted symbol system, while vision is close to a half-billion years old and may function as the actual "foundational intelligent function for capturing the structure of the world." Its argument in brief:
- GPT-3 showed that scaling next-token prediction over web-scale text produces broad capability without any images, audio, or embodied interaction — a striking result, but one built entirely on a compressed abstraction of reality, not reality itself.
- A baby understands gravity, object permanence, and spatial reasoning long before it strings together a sentence. On this view, vision isn't just an input modality alongside text — it's the more foundational one.
- The paper proposes visual general intelligence (VGI) — learning natively from images, video, and geometry, rather than starting with language and translating pixels into text — as a candidate pathway toward AGI, and lays out benchmarks and learning paradigms for pursuing it.
The overlap with Zahavy's argument is direct even though the two papers never reference each other: both conclude that text-only training is a structural ceiling, not a scale problem, and both point at physically grounded, non-linguistic experience — Zahavy's E→A "jump," VGI's vision-as-foundation — as the missing ingredient. Two independent research groups, one from AI safety/philosophy of science and one from computer vision, converged on the same diagnosis within weeks of each other. That convergence is worth more than either paper alone.
It's also worth noting the pushback this paper drew immediately on X: several replies pointed out that frontier LLMs have technically ingested images and video for two-plus years already, and that "natively multimodal" tokenization already treats pixels and words as the same currency. The VGI authors' rebuttal, in effect, is the same one Zahavy's paper makes about prompting — bolting a vision encoder onto a next-token predictor is not the same as building a system whose primary substrate is spatial and physical, any more than reading the word "gravity" is the same as falling.
How X reacted: viral overclaiming, real skepticism, and a separate but related complaint about AI writing
The "LLMs Can't Jump" paper picked up its widest audience not through ICML's own channels but through a viral X thread that flattened "suggests LLMs structurally lack one mechanism" into "Google DeepMind proved LLMs can never make real scientific discoveries. It is mathematically impossible." Neither claim is in the paper — Zahavy's team explicitly softened "confirms" to "suggests" during peer review (see above), and nothing in the paper is a mathematical impossibility proof. It's a position paper built on two historical case studies.
The replies split roughly three ways, and the split itself is informative:
- Agreement with the substance, skepticism of the framing. Commenters pointed out that "mathematically impossible" is a category error for an argued thesis, while still agreeing LLMs alone won't reach AGI without some non-text grounding — echoing the paper's actual, narrower claim.
- Direct technical pushback. Several replies noted that modern LLMs already train on images and video as tokens, arguing the "text-only" framing understates current multimodal training — the same objection the VGI paper's replies drew, and one both papers implicitly answer by distinguishing passive multimodal input from active, embodied grounding.
- A tangential but related complaint: AI-generated writing itself. The same news cycle carried commentary on reader backlash against AI-written content being dismissed as "slop," and a developer survey (via Cynthia Dunlop, covered on Substack) finding that while some writers feel AI assistance makes their prose better, readers who suspect AI authorship stop reading and downvote — a perception gap between "person writing" and "person reading" that explainx.ai has covered at length in its AI slop and Slopocalypse coverage. It's a separate phenomenon from the abduction argument — one is about scientific creativity, the other about writing quality — but both threads surfaced in the same 24 hours of discourse because they share an underlying question: what, exactly, can a system trained only on existing text originate versus merely recombine?
explainx.ai's take: the abduction gap and the data wall are the same problem, viewed from two angles
Everything above is what the papers and the discourse actually claim. This section is explainx.ai's own analysis, not a claim made by Zahavy's paper, the VGI authors, or anyone quoted on X — and it's worth being explicit about that boundary before making it.
Our read: the abduction gap Zahavy's paper describes isn't just a standalone architectural limitation. It's downstream of, and compounded by, a problem the field has been circling for over a year — the data wall. High-quality, human-generated text is a finite resource, and the frontier labs are visibly running low on the easy kind. That's not idle speculation; it's the reason synthetic data pipelines, RL-generated training signal, and now action-controllable world models like Genie have become the industry's default answer to "what do we train the next model on."
Here's the compounding effect we'd flag as a prediction, not an established fact:
- Induction needs a growing corpus. The corpus has stopped growing at the same rate. An LLM's pattern-matching strength scales with fresh, high-quality data it hasn't already seen. As the internet's stock of pre-AI, human-written text gets exhausted — and as more of the new text online is itself AI-generated, recursively diluting the pool — induction itself gets harder to improve, not just abduction.
- Abduction, per Zahavy's own framing, requires a jump from something outside the symbolic system. If a model's entire universe is scraped text, it has nothing ungrounded-in-data to jump from in the first place — there's no E in Einstein's E→A diagram, only ever more A. A genuinely fresh axiom requires contact with something the training corpus didn't already compress into words. That's precisely what a data wall guarantees a text-only model will never get more of.
- This is why both papers land on the same non-linguistic prescription. It's not a coincidence that Zahavy's proposed bridge (physically-consistent world models) and the VGI paper's proposed bridge (vision-as-foundation) both reach past text entirely, rather than proposing more text, better curated. If the diagnosis is "there's no more good data, and even the data that exists can't supply what's structurally missing," then the only two levers left are (a) generate physically grounded experience synthetically, via simulation and world models, or (b) source an entirely different substrate, like raw visual/embodied streams, that was never bottlenecked on internet text in the first place. Both papers pick one of those levers. Neither picks "scale the LLM further."
Our prediction, stated plainly: as the data wall bites harder over the next few model generations, expect the abduction gap this paper documents to become more visible, not less — because the standard fix for a capability gap in this industry has always been "more data," and that fix is running out precisely where it's needed most.
Related reading on explainx.ai
- Is AI out-thinking mathematicians, or just out-remembering them? — the same "reframing the problem" threshold, argued from working memory instead of logic
- The Map Is Not the Territory: Fable 5 field guide — the HN thread that surfaced this paper's citation
- NVIDIA Cosmos 3: Open Physical AI World Models — the Reasoner/Generator split mirrors the paper's induction/abduction distinction
- NVIDIA Alpamayo 2 Super: Open Reasoning Model for Robotaxis — world models as physically grounded training signal
- What Are World Models? Complete Guide — background on the world-model category the paper points toward
- Eight Myths About AI and Software Engineering — a similar genre of grounded, data-backed skepticism about LLM capability claims
- Tencent HY-World 2.0: 3D world models — persistent, editable world generation as another branch of the same field
- AI Models Hallucinate: Why and How to Catch It — related limits of pattern-matching-only systems
Official source: Position: LLMs Can't Jump — OpenReview (ICML 2026 Position Paper Track submission by Tom Zahavy) · Concurrent neuro-symbolic abduction paper (arXiv)
This post summarizes the publicly visible OpenReview submission, reviews, and rebuttal for "Position: LLMs Can't Jump" as of its ICML 2026 acceptance. As a position paper, its claims are an argued thesis, not an empirical result — the authors themselves revised their conclusion from "confirms" to "suggests" during peer review, and this post preserves that framing rather than overstating it.
