Try Odyssey-3: An Experience Machine For Physical Intelligence

Oliver Cameron

Jeff Hawke
October 8th, 2026
Today, we’re launching Odyssey-3, our most powerful foundation world model yet, setting a new state of the art on Physics-IQ Verified and WorldMark. Odyssey-3 Pro achieves the highest Physics-IQ Verified video-to-video score at 66.1, while Odyssey-3 ranks first in three of WorldMark’s four evaluation categories.
We think of Odyssey-3 as an experience machine, drawing on what it has learned about physics, dynamics, and cause-and-effect to power physical AI, train other intelligences, and create new experiences for humans. In our introduction to Odyssey-3, we showed how this knowledge can be applied to robots, humanoids, self-driving cars, and drones, allowing them to acquire new capabilities with relatively little task-specific experience. That knowledge also allows Odyssey-3 to generate worlds that respond to the actions of their inhabitants, providing experience from which other intelligences can learn how to act.
Starting today, you can experience our research preview of Odyssey-3, seeing and interacting with the world knowledge behind these capabilities. If you’re building an autonomous machine and want to understand how Odyssey-3 can be applied to it, please get in touch.
Generating Imagined Experience
In real time, Odyssey-3 generates experiences from descriptions, allowing humans and agents to take actions and introduce events, while the model continues to predict how the world responds. An agent can try different ways of completing a task, while a person can change the conditions and examine how the model responds. In each case, Odyssey-3 generates the experience using its learned understanding of the world, making that knowledge available across different environments and tasks.
Today’s research preview lets you interact with this capability through first-person and third-person navigation, as well as independent camera movement. These controls give you different ways to act within and observe an experience, while the video on screen shows how the model predicts the world will respond.
Generating Environment
Generating Environment
Generating Environment
Generating Environment
Generating Environment
Generating Environment
What Makes a Useful Experience?
For experience to teach us something useful about the world, the consequences of our actions need to make sense. Objects need to behave according to physical principles, actions need to produce appropriate responses, and the environment needs to remain coherent as events unfold. These are the properties we evaluate with Physics-IQ Verified and WorldMark, where Odyssey-3 delivers state-of-the-art results.
Physics-IQ Verified measures how accurately generated events reproduce real-world physical behavior. Odyssey-3 Pro sets a new state of the art in video-to-video at 66.1 and scores 54.7 in image-to-video. Odyssey-3 also advances the frontier between physical accuracy and generation cost, making high-quality experience less expensive to generate and enabling more of it to be produced within the same generation budget.
Odyssey-3's performance on Physics-IQ, a comprehensive benchmark dataset that can only be solved by acquiring a deep understanding of various physical principles, like fluid dynamics, optics, solid mechanics, magnetism and thermodynamics
Generating Environment
Scores average 4 runs; best-of-8 uses 1 run with the same prompts. Costs include prompt fees where reported. Odyssey assumes $1 per MI355X GPU-hour, excluding prompt-rewriting fees. Odyssey-3: 832×480; Pro: 1280×720. Cosmos3 V2V uses compute-based pricing from Oct 1. Source: Physics-IQ Verified leaderboard, Oct 7, 2026.
WorldMark evaluates how reliably generated worlds respond to controls, alongside visual quality and world memory. It measures whether movement follows the requested direction, whether the world responds when commands change, and whether motion remains stable. Odyssey-3 ranks first in three of its four categories: first-person stylized, third-person real, and third-person stylized, demonstrating strong performance across different viewpoints and visual styles.
WorldMark
Mean of the 13 WorldMark metrics (0–100) per split, using WorldMark’s own captions
First-Person Real
Lyra 2.0
84.4
AlayaWorld
83.0
Odyssey-3
80.6
Matrix-Game 3.0
80.1
LingBot-World
79.6
Matrix-Game 2.0
77.1
HY-World 1.5
77.0
SANA-WM
75.2
DreamX-World
72.2
Yume 1.5
67.6
HY-GameCraft 1.0
59.5
50
85
First-Person Stylized
Odyssey-3
77.2
LingBot-World
77.0
AlayaWorld
76.7
Lyra 2.0
75.9
Matrix-Game 3.0
75.6
HY-World 1.5
75.2
Matrix-Game 2.0
73.5
SANA-WM
73.0
DreamX-World
70.6
Yume 1.5
63.3
HY-GameCraft 1.0
55.6
50
85
Third-Person Real
Odyssey-3
79.0
HY-World 1.5
76.9
SANA-WM
74.8
DreamX-World
71.5
LingBot-World
69.5
Matrix-Game 2.0
68.6
60
85
Third-Person Stylized
Odyssey-3
76.3
HY-World 1.5
75.1
SANA-WM
72.8
DreamX-World
70.1
Matrix-Game 2.0
65.5
LingBot-World
60.8
55
85
Adapting Experiences to Physical Systems
Experience also depends on how a system observes the world. A vehicle sees the road through several cameras, while a robot may observe its surroundings from its head or wrist, and those viewpoints shape the information available for making decisions. Applying generated experience to these systems means adapting both the situations they encounter and the views through which they encounter them.
In an early experiment, we adapted an Odyssey-3 training checkpoint to generate three camera views together, using data from the front three cameras of an autonomous-driving dataset. After just 100 training steps, the model produces multi-view driving sequences, showing how a general world model can be adapted to generate experience for a particular physical system.
Generating Environment
Generating Environment
Generating Environment
Generating Environment
The experiment demonstrates a practical connection between general world knowledge and the observations a specific machine needs. Extending this work to a wider variety of camera arrangements is part of making generated experience useful across robots, humanoids, vehicles, and drones, each with different bodies, tasks, and ways of observing the world.
“Pour the cereal into the bowl”
Controlling Robot
“Close the screwbox”
Controlling Robot
Controlling Humanoid
“Open the blue container and take out the cardboard box”
“Move the plate to the center of the table and place the mug on top of it”
Controlling Humanoid
“Take the first roundabout exit”
Driving Autonomously
“Drive along the road”
Driving Autonomously
Building an Experience Machine
We designed Odyssey-3 to generate many kinds of experience within a single model, accepting new actions and instructions throughout a generation and continuing to respond over extended interactions. This requires a model that can draw on what has already happened, predict what follows, and generate that response quickly enough for a person or agent to act again. We built Odyssey-3 as an autoregressive diffusion transformer to support this continuous interaction.
Our data pipeline brings together three complementary sources of experience. Internet video supplies a broad range of environments, events, and human interactions, which we process into time-localized, schema-verified annotations connecting descriptions to the moments when events occur. Gameplay recordings provide time-aligned keyboard and mouse inputs alongside world descriptions and events, connecting observations to the actions that produced them. Simulated rigid-body interactions add focused examples of physical behavior, together with corresponding captions and metadata.
Odyssey-3 begins as a multi-step video diffusion transformer, using temporal RoPE to represent position in time and causal masking to restrict access to future information. We extend this model autoregressively through teacher forcing, training it to continue an experience from preceding observations while accepting prompts and controls that arrive during interaction. This allows a participant’s next action to shape the model’s next prediction.
We then use a post-training pipeline that includes distribution-matching distillation, generative adversarial network discriminators, and reinforcement learning to turn the multi-step model into a few-step model. Reducing the number of generation steps makes it possible to deliver high-quality experience in real time, with participants observing a response and acting again as the world unfolds.
Agents Completing Tasks Within World Models
An agent can participate in this process by observing the generated world and choosing actions through the same harness. In our task-completion demonstration, a task described in natural language becomes a goal the agent pursues inside Odyssey-3: it takes an action, observes the result, and uses that observation to decide what to do next. The agent’s progress is visible in the generated environment, showing how another intelligence can act within a world produced by Odyssey-3.
This is the interaction at the center of training through experience, where an agent improves its behavior using feedback from what it does. Our work on PROWL-2 develops this further, training agents inside a learned world model while using the failures they uncover to improve the model itself. As each becomes more capable, it gives the other new situations to learn from.
Experience Odyssey-3
Odyssey-3 brings together the ability to generate experience and the world knowledge needed to make that experience useful. The same foundation supports physical systems learning new tasks, agents acting within generated environments, and people creating experiences of their own, connecting these applications through a shared understanding of how the world works.
You can experience Odyssey-3 in today’s research preview, choosing a world, acting within it, and seeing how the model responds. If you want to apply Odyssey-3 to your robot or another physical system, get in touch to explore how its learned world knowledge can support the tasks you are working on.




