
Introducing PROWL-2: Jointly Learning Simulation and Decision-Making
PROWL-2 is the first framework in which a team of agents and its world model each learn from their own curriculum, continually and within a single training loop
Ahmet Hamdi Guzel
October 1st, 2026
World models were introduced in model-based reinforcement learning, as a learned dynamics model or ‘transition function’ predicting the next state. The motivation behind this is to enable reinforcement learning via a simulated environment, rather than needing to interact directly with an environment. World models also offer a path to agents that plan and reason about the consequences of their actions, rather than relying on scale alone. However, learning from imagined experience is only as good as the world model: where its predictions are wrong, the agent trains on hallucinated futures and learns behaviour that fails in the real environment.
Most world models so far imagine a world with a single agent in it, yet the experience that teaches the most is shared with others: teammates to coordinate with and rivals to outwit, each of whom raises the bar as they improve. This is why we argue that the era of experience is a multi-agent one. It is also where multi-agent imagined experience is the most difficult: as agents learn, they change the world the others experience. This means a world model must track a moving target, and every interaction it mispredicts can derail the rollout for the whole team.
To address these challenges, we've developed PROWL-2, the next generation of PROWL and, to our knowledge, the first framework in which a team of agents and its world model each learn from their own curriculum, continually and within a single training loop. Building on Unsupervised Environment Design (UED)—the automatic generation or selection of training scenarios that are most useful for improving an RL agent—agents train on rollouts generated entirely by the world model. These imagined trajectories are selected for their learning potential and organized into a curriculum that progresses from easier to harder scenarios as the agents improve. In sync with them, a developer agent explores the real environment to find where the world model breaks, building a curriculum for world-model repair, as we introduced in PROWL-1. A novel fidelity gate couples the two curricula. It periodically compares the world model's rollouts from the agents' highest-priority start states with the corresponding recorded real trajectories, withholds low-fidelity, hallucinated rollouts from policy training, and routes them to world-model repair instead, so that a learning signal produced by world-model error is not mistaken for a genuine one.
Across the StarCraft Multi-Agent Challenge v2 (SMACv2) and the Multi-Agent Quadruped Environment (MQE), PROWL-2 achieves the highest mean performance in all nine SMACv2 scenarios, with up to a 91% relative gain over its baseline world-model learner, while improving success on the hardest MQE tasks from 29.8% to 70.4% on Gate-3 and from 7.3% to 28.2% on Shepherd-Hard.
A Framework for Multi-Agent Learning from Imagined Experience
PROWL-2 is a framework for multi-agent learning from imagined experience, combining prioritized replay–inspired curriculum learning with adversarial world-model exploration. It is agnostic to both the policy-learning algorithm and the world model: rather than changing their architectures or objectives, it decides which data holds the highest learning potential for both the agents and the world model. This means it can be added on top of any existing model-based multi-agent learner. PROWL-2 combines three mechanisms—task-policy curriculum learning, world-model curriculum learning, and dual curriculum learning—within a single training loop to improve agent performance. We evaluated this in both discrete game scenarios and continuous robot coordination.
Task Policy Curriculum Learning
The agents train inside the world model on a curriculum of start states ranked by learning potential. As agents master easy situations, harder ones rise to the top. On its own this curriculum cannot tell a genuine weakness from a world-model error, and it can steer the agents towards hallucinations.
World-Model Curriculum Learning
As in PROWL-1, a developer agent explores the real environment for trajectories the world model mispredicts, rewarded for errors the world model can learn to fix. The world model continuously learns from these failures alongside its ordinary replay.
Dual Curriculum Learning
To connect the two, we developed a novel fidelity gate. It regularly replays the agents' highest-priority start states through the world model under the recorded actions and compares the prediction with the recorded real trajectory. States whose rollouts deviate too far from the recorded trajectory are temporarily removed from the task policy curriculum. Their real trajectories are added to the world-model curriculum, where the world model is fine-tuned on them. Once the fine-tuned world model predicts a withheld state accurately in consecutive checks, the state returns to the task policy curriculum. The gate thus eliminates hallucinated rollouts from the agents' training without discarding the informative states behind them.
Together, the three mechanisms carry unsupervised environment design into a learned world that keeps improving itself. The agents are continually challenged with scenarios with the highest learning potential, and the world model is continually repaired wherever its failures are exposed. This occurs through both the developer agent's directed exploration or by the high-learning-potential challenges the task-solving agents pursue. PROWL-2 thereby extends PROWL-1, where the world model improves only through an agent that actively searches for its failures, by also repairing it on the harder challenges the task-solving agents most want to learn from as well as training task-solving agents themselves
Learning Novel Coordinated Combat in StarCraft 2 Multi-Agent Challenge
We evaluate on SMACv2, a widely used benchmark for cooperative multi-agent combat in StarCraft II. SMACv2 procedurally generates every battle, randomising unit types and starting positions, so agents cannot memorise a fixed scenario and must learn closed-loop coordination that generalises to engagements they have never seen. We test PROWL-2 across all three races, Terran, Protoss and Zerg, at three combat scales: symmetric 5v5 and 10v10 battles, and a harder asymmetric 10v11 scenario in which the trained team of ten fights eleven enemy units.
Maintaining Terran Formations and Concentrating Fire
Terran 5v5, the evaluation seed setup is a symmetric SMACv2 mirror match: both teams start with three Marauders and two Marines in mirrored positions. The behavioral difference is visible in how the teams coordinate. PROWL-2 closes ranks during the approach and reaches contact in a much tighter formation. Four units then focus fire on the same target, securing the first elimination. The baseline enters with a longer, looser line and splits its fire across multiple targets, often losing units before gaining an advantage.
Across 40 evaluations, PROWL-2: 88%, baseline: 8% win rate.
Terran 5v5 · PROWL-2
Terran 5v5 · Baseline
Terran 10v10, the evaluation seed is a symmetric mirror match with seven Marauders and three Marines per side. PROWL-2 reorganizes into a tight firing line before contact, enabling concentrated fire and the first kill without losing a unit. It maintains this advantage and wins with three survivors. The baseline engages in fragmented groups, loses the first unit, and is wiped out after eliminating only three opponents.
Across 40 evaluations, PROWL-2: 70%, baseline: 8% win rate.
Terran 10v10 · PROWL-2
Terran 10v10 · Baseline
Terran 10v11, Terran starts with PROWL-2 at a one-unit disadvantage. PROWL-2 keeps nine of its ten units in a tight group, while a lone Marine draws enemy attention and creates an opening for the main force to secure a double kill. From there, PROWL-2 eliminates all eleven opponents while preserving all four Marauders. The baseline engages in three fragmented groups, loses three units before its first kill, and is wiped out after eliminating only three opponents.
Across 40 evaluations, PROWL-2: 55%, baseline: 2.5% win rate.
Terran 10v11 · PROWL-2
Terran 10v11 · Baseline
Coordinating Protoss Attacks and Tactical Diversions
Protoss 5v5, Both sides start with three Stalkers and two Zealots. A lone PROWL-2 Zealot draws both enemy Zealots away, while the remaining four units focus fire on a single Stalker and secure the first kill. PROWL-2 then eliminates all five opponents. The baseline advances in fragments, splits fire across three targets, and is wiped out after three kills.
Across 40 evaluations, PROWL-2: 70%, baseline: 18% win rate.
Protoss 5v5 · PROWL-2
Protoss 5v5 · Baseline
Protoss 10v10, Both sides start with one Colossus, four Stalkers, and five Zealots. PROWL-2 consolidates into two tight groups and keeps them coordinated through contact, eliminating all ten opponents with four units surviving. The baseline begins with one compact group, but the rest of its army fragments at engagement and it falls one kill short.
Protoss 10v10: Across 40 evaluations, PROWL-2: 65%, baseline: 10% win rate.
Protoss 10v10 · PROWL-2
Protoss 10v10 · Baseline
Protoss 10v11, evaluation seed starts with PROWL-2 at a one-unit disadvantage., PROWL-2 forms two tight groups before contact; when a wounded Stalker pulls back, five enemies chase it into the main firing line, allowing PROWL-2 to secure the first kill and eventually eliminate all eleven opponents with two survivors. The baseline remains highly fragmented, loses two units before its first kill, and is wiped out after eliminating seven opponents.
Across 40 evaluations, PROWL-2: 35%, baseline: 0% win rate.
Protoss 10v11 · PROWL-2
Protoss 10v11 · Baseline
Coordinating Zerg Attacks and Timing Baneling Strikes
Zerg 5v5, evaluation seed is a symmetric match with four Hydralisks and one Zergling per side. PROWL-2 keeps its Hydralisks together and focuses multiple units on one target at a time, eventually winning with a single Hydralisk remaining. The baseline fragments before contact, rarely concentrates fire, and is wiped out after eliminating only the enemy Zergling.
Across 40 evaluations, PROWL-2: 37.5%, baseline: 12.5% win rate.
Zerg 5v5 · PROWL-2
Zerg 5v5 · Baseline
Zerg 10v10, evaluation seed is a symmetric match with five Hydralisks, four Zerglings, and one Baneling per side. PROWL-2 commits its Baneling early, eliminating two enemy Zerglings, then keeps the remaining army coordinated and takes the opening exchange seven kills to two losses.
Across repeated 40 runs from the same setup, PROWL-2 commits its Baneling early, eliminating two enemy Zerglings, then keeps the remaining army coordinated and takes the opening exchange seven kills to two losses. It eventually wins with three units surviving. The baseline delays its Baneling, fragments its Zerglings from the Hydralisks, and is wiped out after seven kills.
Across 40 evaluations, PROWL-2: 90%, baseline: 0% win rate.
Zerg 10v10 · PROWL-2
Zerg 10v10 · Baseline
Zerg 10v11, evaluation seed starts with PROWL-2 at a one-unit disadvantage: eight Zerglings and two Hydralisks against nine Zerglings and two Hydralisks. PROWL-2 coordinates crossfire on isolated targets, securing three kills before losing a unit, then clears the remaining Zerglings and turns on the Hydralisks to win with four units surviving. The baseline engages less coherently, loses two units before its first kill, and is wiped out without damaging the enemy Hydralisks.
Across 40 evaluations, PROWL-2: 42.5%, baseline: 5% win rate.
Zerg 10v11 · PROWL-2
Zerg 10v11 · Baseline
Learning Robust Coordination on the Hardest MQE Tasks
To test PROWL-2 on multi-robot coordination, we use MQE. Unlike StarCraft combat, these are continuous-control tasks where success depends on precise physical coordination and timing between robots. In Gate, quadrupeds must pass through a narrow opening without colliding, requiring them to coordinate who moves through and when; we evaluate both two-robot and three-robot variants. In Shepherd, two quadrupeds must cooperatively herd a sheep robot into a target area, continuously adapting their movements to steer it in the desired direction; we evaluate both easy and hard variants.
To visualize how coordination develops over training, we compare checkpoints at 200K and 1.5M steps for Gate-2 and Shepherd-Easy, and at 600K and 1.5M steps for the more challenging Gate-3 and Shepherd-Hard tasks. At each checkpoint, each training run is evaluated over 20 held-out episodes, and we report success rates alongside representative rollouts.
Coordinating Two Robots Through a Narrow Passage
Gate-2 tests whether two quadruped agents can coordinate their movement through a narrow passage without colliding. The challenge is not simply reaching the other side, but learning a joint strategy that gets both robots through efficiently. At 200K training steps, PROWL-2 achieves a similar success rate to the baseline but discovers a faster solution. By 1.5M steps, both methods achieve near-ceiling success, yet PROWL-2 becomes faster still, while the baseline largely plateaus at a similar completion time.
Gate-2 · PROWL-2 · 200K steps
Gate-2 · Baseline · 200K steps
Gate-2 · PROWL-2 · 1.5M steps
Gate-2 · Baseline · 1.5M steps
Coordinating Three Robots Through a Narrow Passage
Gate-3 increases the coordination challenge by requiring three quadruped agents to pass through the same narrow opening without blocking or colliding with one another. At 600K training steps, the baseline fails across all 20 evaluation scenarios, while PROWL-2 already solves more than 20% of them. By 1.5M steps, PROWL-2 continues to improve, reaching around 70% success, while the baseline remains near 25%. This shows that PROWL-2 not only discovers coordinated traversal earlier, but continues to refine that behavior as training progresses.
Gate-3 · PROWL-2 · 600K steps
Gate-3 · Baseline · 600K steps
Gate-3 · PROWL-2 · 1.5M steps
Gate-3 · Baseline · 1.5M steps
Herding One Sheep Through a Narrow Gate
Shepherd-Easy tests whether two quadruped agents can coordinate to guide a sheep through a narrow gate. At 200K training steps, both methods achieve similar success rates, but PROWL-2 finds a faster herding strategy, reaching the goal in 3.6 s compared with 5.0 s for the baseline. By 1.5M steps, their success rates remain similarly high, yet PROWL-2 further reduces completion time to 2.0 s, compared with 3.2 s for the baseline.
Shepherd-Easy · PROWL-2 · 200K steps
Shepherd-Easy · Baseline · 200K steps
Shepherd-Easy · PROWL-2 · 1.5M steps
Shepherd-Easy · Baseline · 1.5M steps
Herding a Flock of Nine Sheep
Shepherd-Hard scales the challenge to a flock of nine sheep, making coordinated positioning and sustained herding substantially more difficult. At 600K training steps, the baseline has not yet learned a reliable herding strategy, while PROWL-2 already begins coordinating the two agents to move the full flock toward and through the gate. By 1.5M steps, PROWL-2 reaches close to a 30% success rate, while the baseline remains near 5%, showing a much larger gap on this harder coordination task.
Shepherd-Hard · PROWL-2 · 600K steps
Shepherd-Hard · Baseline · 600K steps
Shepherd-Hard · PROWL-2 · 1.5M steps
Shepherd-Hard · Baseline · 1.5M steps
Building Experience Machines for Multi-Agent and Frontier Intelligence
PROWL-2 shows that a team of agents and its world model can learn together, each from its own curriculum. The world model is repaired where it fails, and the agents train on increasingly complex behaviour only where the world model can be trusted. Across discrete StarCraft combat and continuous quadruped control, PROWL-2 shows considerable improvement over baselines in multi-agent reinforcement learning benchmarks.
We see world models as experience machines, places where agents can gain experience that the real world cannot safely or cheaply provide. To learn from that experience at scale, it must be both trustworthy and well ordered, progressing from easy to hard as the agents improve. PROWL-2 makes these two requirements reinforce each other. Each gain in the world model's fidelity opens harder situations to the agents, and each advance by the agents leads the world model into new situations it must learn to predict, so the challenges keep coming for both.
As games, robotics and swarm systems increasingly rely on agents trained in learned simulators, PROWL-2 can be added to these systems as they are, without redesigning their world models or learning algorithms. The same principle reaches beyond these domains to frontier models. We believe multimodal world models are the missing ingredient for the next leap of frontier models, and our research will continue to develop both the world models and the intelligence that acts within them. Foundation world models such as Odyssey-3 can supply that grounded experience, and PROWL is how we intend to make it trustworthy and progressively harder as the models learning from it improve. This kind of open-ended, co-evolving learning is, we believe, a critical part of the path to superintelligence, and PROWL-2 is a step toward that future.
The Team That Brought This to Life
PROWL-2 was made possible by the Odyssey team—Ahmet H. Güzel, Jenny Seidenschwarz, and Jeffrey Hawke—with valuable feedback from University of Basel and University College London advisor Ilia Bogunovic.







