On October 8, 2026, Odyssey launched Odyssey-3, which it calls its most powerful foundation world model. The headline claim is a new state of the art on Physics-IQ Verified, a benchmark from Anates Labs and Google DeepMind: Odyssey-3 Pro reaches 66.1 on the video-to-video track, the highest reported score. Odyssey also reports first place in three of four WorldMark categories in its own evaluation, and shows the model adapted to robot arms, humanoids and a car driving on real roads.
We covered the earlier Odyssey-3 introduction in September, which had no numbers. This post covers what the October launch adds: the leaderboard, the cost-versus-accuracy math, the adaptation demos, and what to check before you rely on the claims. The source is Odyssey's launch post and the leaderboard graphic it published, dated October 7, 2026.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is it? | A foundation world model: an autoregressive diffusion transformer that predicts how scenes evolve |
| Headline score? | 66.1 on Physics-IQ Verified video-to-video (Odyssey-3 Pro, best-of-8) |
| Without best-of-8? | 63.4 for Pro; 61.6 for base Odyssey-3 with prompt enhancement |
| Image-to-video? | 54.7 for Pro with best-of-8 |
| Main rival on the board? | FLUX 3 large from Black Forest Labs at 64.35 with best-of-8 |
| Where is the catch? | Best-of-N sampling, cost figures that exclude prompt-rewriting fees, and company-run WorldMark tests |
| Can I try it? | A free research preview, with API access by contacting Odyssey |
| Is it a robot brain? | It is a foundation for one; the robot demos used small adaptation datasets |
What Odyssey-3 is
Odyssey describes Odyssey-3 as "a learned dynamical system, implemented as an autoregressive diffusion transformer, that predicts how objects move and interact through space and how situations evolve over time." In plain terms: it is a video model that continues a scene frame by frame, conditioned on what happened before and on actions or events you introduce. Developers use that learned knowledge two ways, to simulate environments and to train policies for physical systems.
The research preview shows first-person and third-person navigation plus independent camera movement. You prompt an environment, move through it, or introduce an event, and the model predicts how the world responds in real time. The real-time variant is a distilled few-step model, discussed in the training section below.
For background on the category, see our explainer on what world models are and the Odyssey Agora-2 multi-agent world model.
The Physics-IQ Verified leaderboard
Physics-IQ asks models to continue videos of real physical experiments, covering fluid dynamics, optics, solid mechanics, magnetism and thermodynamics, and compares the prediction with what actually happened. The leaderboard graphic Odyssey shared, for the video-to-video, verified-score, all-sampling, all-models view, ranks 11 entries:
| Rank | Model | Score | Sampling | Output FPS | Cost per video |
|---|---|---|---|---|---|
| 1 | Odyssey-3 Pro | 66.10% | Best of 8 | 16 | $3.484 |
| 2 | Odyssey-3 | 64.43% | Best of 8 | 16 | $1.430 |
| 3 | FLUX 3 large (Black Forest Labs) | 64.35% ± 0.22 | Best of 8 | 24 | $16.430 |
| 4 | Odyssey-3 Pro | 63.37% ± 0.63 | Single | 16 | $0.444 |
| 5 | Odyssey-3 | 61.56% ± 1.20 | Single | 16 | $0.187 |
| 6 | FLUX 3 large | 61.11% ± 0.55 | Single | 24 | $2.080 |
| 7 | Physical Registry + PRS (Microsoft Research Asia) | 52.61% | Best of 4 | 24 | not disclosed |
| 8 | Odyssey-3, no LLM, base prompts | 51.76% ± 0.69 | Single | 16 | $0.177 |
| 9 | Cosmos3 Super (NVIDIA, open source) | 50.80% ± 2.20 | Single | 24 | $0.823 |
| 10 | Physical Registry (Microsoft Research Asia) | 50.17% ± 0.72 | Single | 24 | not disclosed |
| 11 | Cosmos3 Nano (NVIDIA, open source) | 43.00% ± 2.00 | Single | 24 | $0.823 |
Microsoft Research Asia's entries are marked "Not Yet Available." The "±" is sample standard deviation; best-of-8 rows have none because they use one run. All Odyssey and FLUX rows use LLM prompt enhancement with custom prompts, except the Odyssey-3 row at 51.76, which uses base prompts and no LLM.
Odyssey also reports 54.7 on image-to-video for Odyssey-3 Pro with best-of-8, with Odyssey-3 at 52.8 best-of-8, 48.8 with prompt enhancement and 41.0 on base prompts. Its image-to-video chart compares against a long list of video generators including Seedance 2.5, MiniMax H3, Gemini Omni Flash, Veo 3.1, Grok Imagine Video, Sora 2, Wan 2.2 and others, but the numbers for those were not legible in the text we have, so we do not repeat them.
The cost math is as interesting as the score
Odyssey argues that Odyssey-3 "improves the measured tradeoff between physical accuracy and generation cost," meaning more simulations for the same compute budget. Using the table above:
| Comparison | Score gap | Cost ratio |
|---|---|---|
| Odyssey-3, best-of-8 (64.43) vs FLUX 3 large, best-of-8 (64.35) | about equal | FLUX costs about 11.5x more ($16.43 vs $1.43) |
| Odyssey-3 Pro, single (63.37) vs FLUX 3 large, single (61.11) | +2.26 points | FLUX costs about 4.7x more ($2.08 vs $0.444) |
| Odyssey-3 Pro best-of-8 (66.10) vs Odyssey-3 best-of-8 (64.43) | +1.67 points | Pro costs about 2.4x more |
| Odyssey-3 with prompt enhancement (61.56) vs base prompts (51.76) | +9.8 points | $0.187 vs $0.177 |
Two points deserve scrutiny.
Prompt enhancement does most of the work, and its fee is not in the cost. Going from base prompts to LLM-enhanced prompts adds about ten points for about a cent, but Odyssey's footnote says its cost "assumes $1 per MI355X GPU-hour, excluding prompt-rewriting fees." The comparison models' costs "include prompt fees where reported." That is an apples-to-oranges cost comparison, and the real Odyssey cost is higher than shown, by an amount the footnote does not give.
Best-of-N is a different kind of result. Best-of-8 generates eight samples and keeps the best one. The footnote says scores average four runs, while "best-of-8 uses 1 run with the same prompts." Best-of-N needs a way to pick the winner, and in a real application you may not have ground truth to select against. Rows without best-of-N are the fairer guide to what you get from one generation: 63.4 for Pro and 61.6 for base Odyssey-3, against 61.1 for FLUX 3 large.
Other footnotes: Odyssey-3 renders at 832×480 and Pro at 1280×720, Odyssey runs at 16 frames per second against 24 for FLUX 3 and Cosmos3, and Cosmos3's video-to-video cost uses compute-based pricing from October 1. All of that affects what "cost per video" means.
WorldMark: first in three of four splits, by Odyssey's own count
WorldMark measures control-following, visual quality and world memory. Odyssey evaluated models using the benchmark's own captions and the mean of its 13 reported metrics, so these are Odyssey-run numbers.
| Split | Odyssey-3 | Closest competitor |
|---|---|---|
| First-person stylized | 77.2 (1st) | LingBot-World 77.0, AlayaWorld 76.7 |
| First-person real | 80.6 (3rd) | Lyra 2.0 84.4, AlayaWorld 83.0 |
| Third-person real | 79.0 (1st) | HY-World 1.5 76.9 |
| Third-person stylized | 76.3 (1st) | HY-World 1.5 75.1 |
The first-person stylized lead is 0.2 points, which is within the range where evaluation noise could matter. Odyssey itself cautions that these results measure specific properties of generated worlds, and that applying the model to a physical system requires evaluating the behaviors that matter for that machine.
From video model to robot, car and agent
The most strategic part of the launch is the claim that one world model can be adapted to different bodies by training a small action decoder or policy on paired observations and actions.
| Demo | What Odyssey says | What to keep in mind |
|---|---|---|
| Robot arms | With only tens of hours of demonstrations, completed tasks such as pouring cereal and closing a screwbox, and showed recovery behaviors absent from the demonstrations | Small set of tasks; no success rates published in the post |
| Humanoids | Flexion built humanoid policies on Odyssey-3 that exceeded tested VLA baselines under environmental changes and kept working under lighting changes that broke those baselines | Third-party claim relayed by Odyssey; baselines unspecified |
| Driving | A policy trained on 20 hours of driving data in India, with the backbone frozen, predicts waypoints and drives in closed loop | Early result; no safety or intervention metrics given |
| Multi-sensor | A training checkpoint produced three-camera driving sequences after 100 training steps | An early experiment |
| Agent training | An agent pursues a natural-language goal inside the generated world | A demonstration, not a benchmark |
These are credible directions and a useful pitch, since data for physical systems is scarce. They are not independent evidence that the policies are robust. For comparison with another physical-AI stack, see NVIDIA Cosmos 3 and NVIDIA's SIGGRAPH Cosmos update, and for robot hardware with a physical API, 1X NEO.
How it was built
Odyssey's training data combines three sources: internet video with time-localized, schema-verified event annotations, gameplay recordings with time-aligned keyboard and mouse inputs, and simulated rigid-body interactions with captions and metadata. The goal is to connect observations with descriptions of what happens and, where available, the actions that caused it.
The recipe has three stages:
- A multi-step video diffusion transformer that uses temporally resolved prompts and controls to guide how the world unfolds.
- Autoregressive extension through teacher forcing and causal masking, so the model continues from earlier observations and predicts future states given actions.
- Post-training with distribution-matching and adversarial distillation, producing a few-step distilled variant fast enough for real-time interaction.
That last step is why "Flash" can run live while "Pro" is the quality tier.
What is confirmed and what is not
| Claim | Status |
|---|---|
| Physics-IQ Verified scores and ranking | From a leaderboard graphic dated October 7, 2026; benchmark from Anates Labs and Google DeepMind; submission and verification process not detailed in our sources |
| Costs | Odyssey's assumptions; exclude its prompt-rewriting fees |
| WorldMark ranks | Odyssey-run, using the benchmark's captions and mean of 13 metrics |
| Robot, humanoid, driving demos | Odyssey and partner claims, with small data budgets |
| Availability | Research preview; API by contact |
What this means for what you build
- If you build simulators or synthetic-data pipelines, try the preview and measure physical plausibility on your own scenes, not only the benchmark. The cost per video matters when you need thousands of rollouts.
- If you train robot policies, the interesting part is the adapter recipe: a frozen backbone plus a small decoder trained on tens of hours. Reproduce it on your hardware before assuming the headline holds.
- Read the sampling column. Compare single-generation scores with single-generation scores. Best-of-N is a tool for offline data generation, not a free accuracy boost.
- Ask for the full cost. Include prompt rewriting, resolution and frame rate when you compare per-video prices.
- Watch for independent replication. A benchmark leaderboard run by outside organizations is stronger evidence than a vendor's chart, and the other world models in this space, including Runway's GWM Worlds 2, ByteDance's Seedance world model and Tencent's HY-World 2, will be measured against the same bars.
Related reading on explainx.ai
- Odyssey-3 introduction: a unified foundation world model
- Odyssey Agora-2: a multi-agent world model for 20 players
- What are world models? Starchild-1 and Odyssey explained
- NVIDIA Cosmos 3: open physical AI world model
- FLUX 3 from Black Forest Labs: multimodal, video and robotics
- Runway GWM Worlds 2: an interactive world model
- ByteDance Seedance as a world model
- Tencent HY-World 2
Scores, costs and rankings come from Odyssey's October 8, 2026 launch post and the Physics-IQ Verified leaderboard graphic dated October 7, 2026, and may change. Many figures are company-run. We have not run the model. The leaderboard graphic itself is not reproduced here; the table is our own transcription of its text.
