Meet Odyssey-3: Our Most Powerful Foundation World Model

Oliver Cameron

Jeff Hawke
October 8th, 2026
Today we’re launching Odyssey-3, our most powerful foundation world model yet. Odyssey-3 Pro sets a new state of the art on Physics-IQ Verified’s benchmark, and ranks 1st in 3 of WorldMark’s 4 categories in our evaluations. Our research preview is available now, and if you’re a physical AI developer interested in building with Odyssey-3, please get in touch.
Odyssey-3 is a learned dynamical system, implemented as an autoregressive diffusion transformer, that predicts how objects move and interact through space and how situations evolve over time. It learns representations of physics, dynamics, and cause-and-effect from a broad collection of visual observations, and developers use that knowledge both to simulate environments and to train policies for different physical systems.
Our founding team spent a decade building driverless cars, where predicting the world was essential to determining what a car should do next. We founded Odyssey to pursue that idea far beyond the roads, to build a general-purpose technology that could bring learned world knowledge to all machines and tasks. With Odyssey-3, we believe this idea is now being realized.
Odyssey-3 Generates Environments in Real Time
Odyssey-3 generates embodied environments from a prompt and predicts in real time how they change as a person or agent takes actions or introduces events, using previous observations and the latest inputs. You can move through the environment or introduce an event during generation and observe how the model responds.
Today’s preview provides first-person and third-person navigation alongside independent camera movement, giving you different ways to interact with and inspect the model’s predictions.
Generating Environment
Generating Environment
Generating Environment
Generating Environment
Generating Environment
Generating Environment
Generating Environment
Generating Environment
Odyssey-3 Is State of the Art on Physics-IQ
Odyssey-3 Pro sets a new state of the art on Physics-IQ Verified’s video-to-video benchmark, achieving 66.1, the highest reported score. We have made physical accuracy a central focus of our research because a foundation model for physical intelligence must learn to predict how the world actually behaves.
Physics-IQ tests physical behavior across fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics, asking models to continue videos of real physical experiments and comparing their predictions with what actually happened. Odyssey-3 Pro also scores 54.7 in image-to-video, while Odyssey-3 improves the measured tradeoff between physical accuracy and generation cost, making it possible to generate more simulations within the same compute budget.
WorldMark measures control-following, visual quality, and world memory. In our evaluation, using the benchmark’s own captions and the mean of its 13 reported metric scores, Odyssey-3 ranks first in first-person stylized, third-person real, and third-person stylized environments. These results measure specific properties of generated worlds; applying the model to a physical system also requires evaluating the behaviors that matter for that machine and its tasks.
Odyssey-3's performance on Physics-IQ, a comprehensive benchmark dataset that can only be solved by acquiring a deep understanding of various physical principles, like fluid dynamics, optics, solid mechanics, magnetism and thermodynamics
Generating Environment
Scores average 4 runs; best-of-8 uses 1 run with the same prompts. Costs include prompt fees where reported. Odyssey assumes $1 per MI355X GPU-hour, excluding prompt-rewriting fees. Odyssey-3: 832×480; Pro: 1280×720. Cosmos3 V2V uses compute-based pricing from Oct 1. Source: Physics-IQ Verified leaderboard, Oct 7, 2026
WorldMark measures control-following, visual quality, and world memory. In our evaluation, using the benchmark’s own captions and the mean of its 13 reported metric scores, Odyssey-3 ranks first in first-person stylized, third-person real, and third-person stylized environments. These results measure specific properties of generated worlds; applying the model to a physical system also requires evaluating the behaviors that matter for that machine and its tasks.
WorldMark
Mean of the 13 WorldMark metrics (0–100) per split, using WorldMark’s own captions
First-Person Stylized
Odyssey-3
77.2
LingBot-World
77.0
AlayaWorld
76.7
Lyra 2.0
75.9
Matrix-Game 3.0
75.6
HY-World 1.5
75.2
Matrix-Game 2.0
73.5
SANA-WM
73.0
DreamX-World
70.6
Yume 1.5
63.3
HY-GameCraft 1.0
55.6
50
85
First-Person Real
Lyra 2.0
84.4
AlayaWorld
83.0
Odyssey-3
80.6
Matrix-Game 3.0
80.1
LingBot-World
79.6
Matrix-Game 2.0
77.1
HY-World 1.5
77.0
SANA-WM
75.2
DreamX-World
72.2
Yume 1.5
67.6
HY-GameCraft 1.0
59.5
50
85
Third-Person Real
Odyssey-3
79.0
HY-World 1.5
76.9
SANA-WM
74.8
DreamX-World
71.5
LingBot-World
69.5
Matrix-Game 2.0
68.6
60
85
Third-Person Stylized
Odyssey-3
76.3
HY-World 1.5
75.1
SANA-WM
72.8
DreamX-World
70.1
Matrix-Game 2.0
65.5
LingBot-World
60.8
55
85
Odyssey-3 Adapts to Physical Systems
Our introduction to Odyssey-3 showed how Odyssey-3’s learned world knowledge can be applied to different systems by training an action decoder or policy on paired observations and actions. These learned components translate that knowledge into the controls required by a particular machine, letting developers adapt the foundation to a new body or task.
Odyssey-3 Can Control Robot Arms
With tens of hours of robot demonstrations, Odyssey-3 completed manipulation tasks and showed recovery behaviors absent from those demonstrations, including reorienting a gripper after a missed grasp and retrieving a dropped object in an unusual position.
“Pour the cereal into the bowl”
Controlling Robot
“Close the screwbox”
Controlling Robot
Odyssey-3 Can Power Humanoids
Flexion has built humanoid control policies on Odyssey-3. The resulting policies exceeded the performance of the tested VLA baselines under environmental changes and continued to perform tasks under lighting changes that caused those baselines to fail.
Controlling Humanoid
“Open the blue container and take out the cardboard box”
“Move the plate to the center of the table and place the mug on top of it”
Controlling Humanoid
Odyssey-3 Can Drive Vehicles
We adapted Odyssey-3 to drive a car on real roads in India, training a driving policy on just 20 hours of driving data while keeping the Odyssey-3 backbone frozen. The policy uses the model’s visual representations to predict waypoints ahead of the car, allowing it to drive in closed loop.
“Take the first roundabout exit”
Driving Autonomously
“Drive along the road”
Driving Autonomously
Odyssey-3 Can Generate Multi-Sensor Data
We adapted Odyssey-3 to generate observations for particular sensor arrangements. In an early experiment using the front three cameras of an autonomous-driving dataset, an Odyssey-3 training checkpoint produced driving sequences with three camera views generated together after just 100 training steps.
Generating Environment
Generating Environment
Generating Environment
Generating Environment
Odyssey-3 Can Train and Evaluate Agents
An agent is an AI system that pursues a goal by observing its surroundings, choosing actions, and using what happens to decide what to do next. A world model can provide the environment in which those decisions are made, giving us a way to study how an agent responds to changing conditions and whether it can complete a task inside a world whose behavior is learned.
In our task-completion demonstration, an agent receives a natural-language goal and pursues it inside Odyssey-3, observing the generated world as it works toward the task.
How We Built Odyssey-3
Odyssey-3’s training data combines internet video with time-localized, schema-verified event annotations, gameplay recordings with time-aligned keyboard and mouse inputs, and simulated rigid-body interactions with captions and metadata. Together, these sources connect diverse observations with descriptions of what happens and, where available, the actions that produced it.
We begin with a multi-step video diffusion transformer using temporal rotary positional embeddings (RoPE) and causal masking, then extend it autoregressively through teacher forcing so it learns to continue from preceding observations while accepting prompts and controls during interaction. Post-training combines distribution-matching distillation, generative adversarial network discriminators, and reinforcement learning to produce a few-step model that responds quickly enough for real-time interaction.
Experience Odyssey-3 Today
We believe world models will power increasingly capable physical AI, generate environments in which other intelligences can train, and enable new kinds of human experiences, and we believe Odyssey-3 is a big leap towards this.
You can try Odyssey-3 in our research preview, prompting an environment, acting within it, and seeing how the world model responds. If you’re developing a robot, humanoid, self-driving car, drone, or any other autonomous system and want to explore how foundation world models can accelerate your work, please get in touch.




