Meet Odyssey-3: Our Most Powerful Foundation World Model

Oliver Cameron

Jeff Hawke

October 8th, 2026

Today we’re launching Odyssey-3, our most powerful foundation world model yet. Odyssey-3 Pro sets a new state of the art on Physics-IQ Verified’s benchmark, and ranks 1st in 3 of WorldMark’s 4 categories in our evaluations. Our research preview is available now, and if you’re a physical AI developer interested in building with Odyssey-3, please get in touch.

Odyssey-3 is a learned dynamical system, implemented as an autoregressive diffusion transformer, that predicts how objects move and interact through space and how situations evolve over time. It learns representations of physics, dynamics, and cause-and-effect from a broad collection of visual observations, and developers use that knowledge both to simulate environments and to train policies for different physical systems.

Our founding team spent a decade building driverless cars, where predicting the world was essential to determining what a car should do next. We founded Odyssey to pursue that idea far beyond the roads, to build a general-purpose technology that could bring learned world knowledge to all machines and tasks. With Odyssey-3, we believe this idea is now being realized.

Odyssey-3 Generates Environments in Real Time

Odyssey-3 generates embodied environments from a prompt and predicts in real time how they change as a person or agent takes actions or introduces events, using previous observations and the latest inputs. You can move through the environment or introduce an event during generation and observe how the model responds.

Today’s preview provides first-person and third-person navigation alongside independent camera movement, giving you different ways to interact with and inspect the model’s predictions.

Generating Environment

Generating Environment

Generating Environment

Generating Environment

Generating Environment

Generating Environment

Generating Environment

Generating Environment

Odyssey-3 Is State of the Art on Physics-IQ

Odyssey-3 Pro sets a new state of the art on Physics-IQ Verified’s video-to-video benchmark, achieving 66.1, the highest reported score. We have made physical accuracy a central focus of our research because a foundation model for physical intelligence must learn to predict how the world actually behaves.

Physics-IQ tests physical behavior across fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics, asking models to continue videos of real physical experiments and comparing their predictions with what actually happened. Odyssey-3 Pro also scores 54.7 in image-to-video, while Odyssey-3 improves the measured tradeoff between physical accuracy and generation cost, making it possible to generate more simulations within the same compute budget.

WorldMark measures control-following, visual quality, and world memory. In our evaluation, using the benchmark’s own captions and the mean of its 13 reported metric scores, Odyssey-3 ranks first in first-person stylized, third-person real, and third-person stylized environments. These results measure specific properties of generated worlds; applying the model to a physical system also requires evaluating the behaviors that matter for that machine and its tasks.

Odyssey-3's performance on Physics-IQ, a comprehensive benchmark dataset that can only be solved by acquiring a deep understanding of various physical principles, like fluid dynamics, optics, solid mechanics, magnetism and thermodynamics

Generating Environment

Physics-IQ Verified: Video-to-Video. Log-scale x axis: Generation cost per video in USD. Y axis: Physics-IQ Verified score, range 35 to 75. Odyssey-3 480p series: Odyssey-3 51.8, Odyssey-3 61.6, Odyssey-3 64.4. Odyssey-3 Pro 720p series: Odyssey-3 Pro 63.4, Odyssey-3 Pro 66.1. Other models are shown in gray by company, with a dotted best-of-other-models frontier. Source: physics-iq-verified.anates.ai, 7 Oct 2026.Physics-IQ Verified: Video-to-VideoScore Against Generation CostBlack Forest LabsNVIDIAOdyssey-3 Pro (720p)Odyssey-3 (480p)Best of Other Models354045505560657075$0.05$0.10$0.20$0.50$1$2$5$10$20Generation Cost per Video (USD, Log Scale)Physics-IQ Verified ScoreFLUX 3 [large]+ Best-of-8FLUX 3 [large]Prompt EnhancedCosmos3-SuperPrompt EnhancedCosmos3-SuperBase PromptsCosmos3-NanoBase PromptsOdyssey-3 51.8Base PromptsOdyssey-3 61.6Prompt EnhancedOdyssey-3 64.4+ Best-of-8Odyssey-3 Pro 63.4Prompt EnhancedOdyssey-3 Pro 66.1+ Best-of-8
Physics-IQ Verified: Image-to-Video. Log-scale x axis: Generation cost per video in USD. Y axis: Physics-IQ Verified score, range 20 to 60. Odyssey-3 480p series: Odyssey-3 41.0, Odyssey-3 48.8, Odyssey-3 52.8. Odyssey-3 Pro 720p series: Odyssey-3 Pro 50.0, Odyssey-3 Pro 54.7. Other models are shown in gray by company, with a dotted best-of-other-models frontier. Source: physics-iq-verified.anates.ai, 7 Oct 2026.Physics-IQ Verified: Image-to-VideoScore Against Generation CostBlack Forest LabsNVIDIAOdyssey-3 Pro (720p)Odyssey-3 (480p)Best of Other Models202530354045505560$0.10$0.20$0.50$1$2$5$10Generation Cost per Video (USD, Log Scale)Physics-IQ Verified ScoreFLUX 3 Large+ Best-of-8FLUX 3 LargePrompt EnhancedPhysis-Lang (Cosmos3-Super)Prompt EnhancedCosmos3-SuperPrompt EnhancedCosmos3-SuperBase PromptsPhysis-Lang(Cosmos3-Nano)Prompt EnhancedCosmos3-NanoBase PromptsSeedance 2.5MiniMax H3MiniMax H3 MaxGemini Omni 1.1 FlashGemini Omni FlashVeo 3.1 LiteVeo 3.1 FastGrok Imagine VideoHunyuan Video 1.5Wan 2.2 14BWan 2.2 5BCogVideoX-5BKandinsky-WM 1.0Sora 2P-VideoOdyssey-3 41.0Base PromptsOdyssey-3 48.8Prompt EnhancedOdyssey-3 52.8+ Best-of-8Odyssey-3 Pro 50.0Prompt EnhancedOdyssey-3 Pro 54.7+ Best-of-8
Scores average 4 runs; best-of-8 uses 1 run with the same prompts. Costs include prompt fees where reported. Odyssey assumes $1 per MI355X GPU-hour, excluding prompt-rewriting fees. Odyssey-3: 832×480; Pro: 1280×720. Cosmos3 V2V uses compute-based pricing from Oct 1. Source: Physics-IQ Verified leaderboard, Oct 7, 2026

WorldMark measures control-following, visual quality, and world memory. In our evaluation, using the benchmark’s own captions and the mean of its 13 reported metric scores, Odyssey-3 ranks first in first-person stylized, third-person real, and third-person stylized environments. These results measure specific properties of generated worlds; applying the model to a physical system also requires evaluating the behaviors that matter for that machine and its tasks.

WorldMark

Mean of the 13 WorldMark metrics (0–100) per split, using WorldMark’s own captions

First-Person Stylized

Odyssey-3

77.2

LingBot-World

77.0

AlayaWorld

76.7

Lyra 2.0

75.9

Matrix-Game 3.0

75.6

HY-World 1.5

75.2

Matrix-Game 2.0

73.5

SANA-WM

73.0

DreamX-World

70.6

Yume 1.5

63.3

HY-GameCraft 1.0

55.6

50

55

60

65

70

75

80

85

First-Person Real

Lyra 2.0

84.4

AlayaWorld

83.0

Odyssey-3

80.6

Matrix-Game 3.0

80.1

LingBot-World

79.6

Matrix-Game 2.0

77.1

HY-World 1.5

77.0

SANA-WM

75.2

DreamX-World

72.2

Yume 1.5

67.6

HY-GameCraft 1.0

59.5

50

55

60

65

70

75

80

85

Third-Person Real

Odyssey-3

79.0

HY-World 1.5

76.9

SANA-WM

74.8

DreamX-World

71.5

LingBot-World

69.5

Matrix-Game 2.0

68.6

60

65

70

75

80

85

Third-Person Stylized

Odyssey-3

76.3

HY-World 1.5

75.1

SANA-WM

72.8

DreamX-World

70.1

Matrix-Game 2.0

65.5

LingBot-World

60.8

55

60

65

70

75

80

85

Odyssey-3 Adapts to Physical Systems

Our introduction to Odyssey-3 showed how Odyssey-3’s learned world knowledge can be applied to different systems by training an action decoder or policy on paired observations and actions. These learned components translate that knowledge into the controls required by a particular machine, letting developers adapt the foundation to a new body or task.

Odyssey-3 Can Control Robot Arms

With tens of hours of robot demonstrations, Odyssey-3 completed manipulation tasks and showed recovery behaviors absent from those demonstrations, including reorienting a gripper after a missed grasp and retrieving a dropped object in an unusual position.

“Pour the cereal into the bowl”

Controlling Robot

“Close the screwbox”

Controlling Robot

Odyssey-3 Can Power Humanoids

Flexion has built humanoid control policies on Odyssey-3. The resulting policies exceeded the performance of the tested VLA baselines under environmental changes and continued to perform tasks under lighting changes that caused those baselines to fail.

Controlling Humanoid

“Open the blue container and take out the cardboard box”
“Move the plate to the center of the table and place the mug on top of it”

Controlling Humanoid

Odyssey-3 Can Drive Vehicles

We adapted Odyssey-3 to drive a car on real roads in India, training a driving policy on just 20 hours of driving data while keeping the Odyssey-3 backbone frozen. The policy uses the model’s visual representations to predict waypoints ahead of the car, allowing it to drive in closed loop.

“Take the first roundabout exit”

Driving Autonomously

“Drive along the road”

Driving Autonomously

Odyssey-3 Can Generate Multi-Sensor Data

We adapted Odyssey-3 to generate observations for particular sensor arrangements. In an early experiment using the front three cameras of an autonomous-driving dataset, an Odyssey-3 training checkpoint produced driving sequences with three camera views generated together after just 100 training steps.

Generating Environment

Generating Environment

Generating Environment

Generating Environment

Odyssey-3 Can Train and Evaluate Agents

An agent is an AI system that pursues a goal by observing its surroundings, choosing actions, and using what happens to decide what to do next. A world model can provide the environment in which those decisions are made, giving us a way to study how an agent responds to changing conditions and whether it can complete a task inside a world whose behavior is learned.

In our task-completion demonstration, an agent receives a natural-language goal and pursues it inside Odyssey-3, observing the generated world as it works toward the task.

How We Built Odyssey-3

Odyssey-3’s training data combines internet video with time-localized, schema-verified event annotations, gameplay recordings with time-aligned keyboard and mouse inputs, and simulated rigid-body interactions with captions and metadata. Together, these sources connect diverse observations with descriptions of what happens and, where available, the actions that produced it.

We begin with a multi-step video diffusion transformer using temporal rotary positional embeddings (RoPE) and causal masking, then extend it autoregressively through teacher forcing so it learns to continue from preceding observations while accepting prompts and controls during interaction. Post-training combines distribution-matching distillation, generative adversarial network discriminators, and reinforcement learning to produce a few-step model that responds quickly enough for real-time interaction.

Experience Odyssey-3 Today

We believe world models will power increasingly capable physical AI, generate environments in which other intelligences can train, and enable new kinds of human experiences, and we believe Odyssey-3 is a big leap towards this.

You can try Odyssey-3 in our research preview, prompting an environment, acting within it, and seeing how the world model responds. If you’re developing a robot, humanoid, self-driving car, drone, or any other autonomous system and want to explore how foundation world models can accelerate your work, please get in touch.

World Model

Odyssey-3

Our most powerful foundation world model yet, materially advancing the state-of-the-art in physical accuracy of world models

World Model

Agora-2

A multi-agent world model, enabling multiple participants—human or AI—to share and interact within the same world simulation in real-time

Reinforcement Learning

PROWL-2

PROWL-2 is the first framework in which a team of agents and its world model each learn from their own curriculum, continually and within a single training loop

World Model

Starchild-1

A step beyond world models that learn only from visual observation, toward systems that learn from richer multimodal interaction with the world

Information

World Models

World Model

Odyssey-3

Our most powerful foundation world model yet, materially advancing the state-of-the-art in physical accuracy of world models

World Model

Agora-2

A multi-agent world model, enabling multiple participants—human or AI—to share and interact within the same world simulation in real-time

Reinforcement Learning

PROWL-2

PROWL-2 is the first framework in which a team of agents and its world model each learn from their own curriculum, continually and within a single training loop

World Model

Starchild-1

A step beyond world models that learn only from visual observation, toward systems that learn from richer multimodal interaction with the world

Information

World Models

World Model

Odyssey-3

Our most powerful foundation world model yet, materially advancing the state-of-the-art in physical accuracy of world models

World Model

Agora-2

A multi-agent world model, enabling multiple participants—human or AI—to share and interact within the same world simulation in real-time

Reinforcement Learning

PROWL-2

PROWL-2 is the first framework in which a team of agents and its world model each learn from their own curriculum, continually and within a single training loop

World Model

Starchild-1

A step beyond world models that learn only from visual observation, toward systems that learn from richer multimodal interaction with the world

Information

World Models

World Model

Odyssey-3

Our most powerful foundation world model yet, materially advancing the state-of-the-art in physical accuracy of world models

World Model

Agora-2

A multi-agent world model, enabling multiple participants—human or AI—to share and interact within the same world simulation in real-time

Reinforcement Learning

PROWL-2

PROWL-2 is the first framework in which a team of agents and its world model each learn from their own curriculum, continually and within a single training loop

World Model

Starchild-1

A step beyond world models that learn only from visual observation, toward systems that learn from richer multimodal interaction with the world

Information

World Models