Skip to content

Ch.08Part II · The seven familiesWorld simulators

World models as simulators, data engines and RL environments

The world model is not the policy: it generates futures, synthetic trajectories or imagined environments that other models train or plan in.

Math
L1

One boxed equation per idea

Sources
25

11 papers from 2026

Figures
4

Most are interactive

Exercises
3

Rust, with tests

Last checked
3 Oct 2026

The field moves monthly

Papers cited, by half-year of first arXiv version

Plannergets imagined futures, scoredData enginegets many synthetic trajectoriesRL environmentgets next frames and rewardsPolicy evaluatorgets a predicted success rateworld modelpredicts what happens nextgiven what is done

RL environment

Sends in the policy's actions; gets back next frames and rewards.

World4RL refines a policy entirely inside a frozen diffusion world model (September 2025); WMPO runs on-policy GRPO in imagination (November 2025); World-VLA-Loop co-trains the world model and the VLA (February 2026); VLA-MBPO adds multi-view consistency and branched rollouts (March 2026).

Where it goes wrong: The policy optimises the model, not the world, and finds the model's mistakes (figure 8.3).

Figure 8.1In this family the world model is never the policy: it is the place where planning, data, training and evaluation happen instead of on a real robot. One world model, four jobs. Select a job for what goes in, what comes back, which models do it, and where it goes wrong. Source: roles after the WAM survey (arXiv 2609.16074) and WorldArena (2602.08971); examples from each paper's abstract.

Why this chapter exists

The families so far all produce a policy: something that, given what the robot sees, says what to do. This chapter is about models that do not. A world model as a simulator answers a different question: if this were done, what would happen next? Other models then use the answer. A planner tries candidate actions in it. A data pipeline asks it to imagine thousands of demonstrations nobody recorded. A reinforcement learning (RL) algorithm trains a policy inside it, as if it were a video game. An evaluator runs a policy in it before letting the policy near a real robot.

The motivation is economic and practical. Real robot interaction is slow, expensive and, for RL, risky: a policy that is still learning breaks things. Physics simulators are fast and safe but narrow, and their images and contact physics never quite match reality. Learned world models promise the breadth of real video with the safety of a simulator, at a fidelity cost that is the subject of most of this chapter.

After this chapter you will be able to:

  • name the four roles a world model plays for robot learning, and place a new paper in one of them;
  • explain with figure 8.3 how a policy trained in imagination exploits the model's mistakes, and two ways to limit it;
  • tell a plausible rollout from a feasible one with figure 8.4, and say why image-quality metrics do not tell them apart.

Robotics moves fastest when researchers can build on open platforms, share code and test ideas on real machines.

Steve Cousins, executive director of the Stanford Robotics Center. NVIDIA Announces NVIDIA Isaac GR00T Reference Humanoid Robot for Academic Research, NVIDIA Newsroom, 31 May 2026.

The mechanism

A world model in this chapter is a forward dynamics model in the sense of chapter 4. It predicts the next observation from what has been seen and what is done:

o^t+1=W(o≤t, at)\boxed{\hat o_{t+1} = W(o_{\le t},\ a_t)}

Two features make the modern versions different from a physics engine. The observation oo is camera video, so the model can be trained on footage rather than on hand-built scenes. And WW is a large generative model, usually a video diffusion model, pretrained on internet-scale video and then adapted to robots. NVIDIA's Cosmos-Predict2.5 is the clearest open example: trained on 200 million curated video clips, released at 2 and 14 billion parameters, and built to support "synthetic data generation, policy evaluation, and closed-loop simulation" (NVIDIA et al. 2025, arXiv 2511.00062). Genie showed that the actions themselves can be latent: trained on unlabelled internet video, it lets a user act frame by frame in worlds it generates (Bruce et al. 2024, arXiv 2402.15391).

The idea of learning a world model and acting inside it is older than any of these. PlaNet planned in a learned latent model in 2018, Dreamer learned behaviours by imagining in it, and DreamerV3 outperformed specialised methods across more than 150 tasks with a single configuration, including collecting diamonds in Minecraft from scratch (Hafner et al. 2018, arXiv 1811.04551; Hafner et al. 2019, arXiv 1912.01603; Hafner et al. 2023, arXiv 2301.04104). What changed in 2025 and 2026 is the observation: real camera video of real robots, at a quality good enough that a policy trained on it can transfer.

Role 1: the data engine

The most widely used role is generating training data. Video models generate pixels, not actions, so every synthetic-data pipeline has the same shape: generate a video, then recover the actions that would have produced it with an inverse dynamics model,

a^t=IDM(ot, ot+1)\boxed{\hat a_t = \mathrm{IDM}(o_t,\ o_{t+1})}

and train the policy on the result as if it were a demonstration. DreamGen made the recipe explicit in four stages (Jang et al. 2025, arXiv 2505.12705), drawn in figure 8.2.

1. Fine-tune the video modelreal data in2. Generate videosworld model3. Label pseudo-actionsinverse dynamics4. Train the policyimitation5. Evaluatecloses the loopreal data is the seed;the world model multiplies it

1. Fine-tune the video model

In: teleoperated trajectories of the target robot. Out: a video model that knows this robot's body.

An off-the-shelf image-to-video model learns the robot's appearance, kinematics and dynamics.

Figure 8.2Real data is the seed and the world model multiplies it; the inverse dynamics step is what turns video back into training data. The synthetic data loop. DreamGen's four stages, with evaluation closing the loop. Select a stage for what goes in and comes out.schematic, not measured Source: Jang et al. 2025 (arXiv 2505.12705), section 1 and figure 2.

The results are what made the role popular. DreamGen reports that a humanoid learned 22 new behaviours, in seen and unseen environments, from teleoperation data of a single pick-and-place task in one environment (Jang et al. 2025, arXiv 2505.12705). NVIDIA's synthetic motion blueprint generated 780,000 trajectories, which it equates to 6,500 hours of human demonstration, in 11 hours, and reports improving GR00T N1's performance by 40 percent when combining them with real datavendor claim (NVIDIA Newsroom 2025). GigaWorld-0 is designed explicitly as a data engine, combining video generation with 3D reconstruction and physical system identification. Its makers report that GigaBrain-0, trained on its data, performs strongly on physical robots "without any real-world interaction during training" (GigaWorld Team et al. 2025, arXiv 2511.19861; GigaBrain Team et al. 2025, arXiv 2510.19430). Cosmos-Transfer does the related job of style transfer: it turns simulator renderings or segmentation maps into realistic video, for sim-to-real and real-to-real augmentation (NVIDIA et al. 2025, arXiv 2503.14492).

Role 2: the RL environment

Imitation learning copies what demonstrators did. To improve beyond them, a policy has to try things and learn from the outcome, which is RL. On a real robot that means thousands of attempts, resets and risk. Inside a world model it means compute. The policy acts, the model predicts what happens, a reward model scores it, and the policy updates. The catch fits in one line. The policy maximises the return the model predicts, J^(π)\hat J(\pi), not the real return J(π)J(\pi), and the difference is the model's bias:

model bias=J^(π)−J(π)\boxed{\text{model bias} = \hat J(\pi) - J(\pi)}

A policy that is good at maximising J^\hat J is, among other things, good at finding where J^\hat J is too high. Figure 8.3 shows how quickly that happens.

targetedgedashed: where the model says it stopssolid: where it really goes: off the edgelogged data-10100.20.40.60.81push strengthreward
Reward the policy believes
0.99
Reward it really gets
-1.00the block falls off the table
Figure 8.3A policy trained inside a world model optimises the model, and the first thing it finds is the model's blind spot. A robot learns how hard to push a block onto a mark near the table's edge, training only inside a world model fitted to logged pushes. Train it, then widen the logged data or keep the policy near it.toy simulation Source: this book's toy, in rust/ch08-simulators; the failure mode is the model bias named in WMPO (arXiv 2511.09515), World-VLA-Loop (2602.06508) and MBDPO (2605.26282).

With logged pushes only up to a strength of 0.4, the world model is a straight line through a curve that bends upwards, and it has never seen the block fall. Inside the model, the best push is about 0.85: the predicted slide lands on the mark and the imagined reward is 0.99. In the real world a push of 0.85 sends the block off the table. Nothing in the training signal warned the policy, because the training signal was the model.

The two fixes in the figure are the two families of fixes in the literature. Stay near the data: limit the policy to where the model has evidence, which is safe but stops short of the best push. Fix the data: collect where the policy now goes, including failures, and retrain the model. With logged pushes up to 1.0 the model learns where the edge is, and the policy lands close to the mark. World-VLA-Loop builds that second fix into the algorithm. It curates successful and "near-success" trajectories so the model learns subtle failures, predicts rewards from the same latents as the video, and feeds each improved policy's rollouts back into the world model (Liu et al. 2026, arXiv 2602.06508). VLA-MBPO attacks multi-view consistency and error compounding directly (Zhang et al. 2026, arXiv 2603.20607). A May 2026 study argues the deeper problem is a misalignment between search and value learning, and unifies them through diffusion policies (Cheng et al. 2026, arXiv 2605.26282).

The results so far are encouraging. World4RL refines pretrained policies entirely inside a frozen diffusion world model (Jiang et al. 2025, arXiv 2509.19080). WMPO runs on-policy RL for VLAs without touching the real environment and reports emergent self-correction (Zhu et al. 2025, arXiv 2511.09515). DreamPlan fine-tunes a vision-language planner inside a video world model trained on the planner's own suboptimal exploration (Jia et al. 2026, arXiv 2603.16860). WAM-OPD, from August 2026, swaps sparse-reward RL for dense supervision from a teacher on the student's own rollouts. Its authors present early RoboTwin results as "an initial capability proof rather than evidence of broad or uniform generalization" (Yang et al. 2026, arXiv 2608.22364).

Role 3: the evaluator

Evaluating a policy on a real robot is the slowest step in robot learning, and chapter 15 is about doing it well. A world model that predicted real success rates would make evaluation cheap. The Interactive World Simulator runs stable interactions for more than 10 minutes at 15 frames per second on a single RTX 4090. Its authors report that policies trained on its generated data perform comparably to those trained on the same amount of real data, and that success in the model correlates strongly with success in the real world (Wang et al. 2026, arXiv 2603.08546). DreamDojo lists policy evaluation among its applications (Gao et al. 2026, arXiv 2602.06949).

The warning comes from WorldArena, which evaluated 14 world models as data engines, policy evaluators and planners and found "a significant perception-functionality gap": high visual quality "does not necessarily translate into strong embodied task capability" (Shang et al. 2026, arXiv 2602.08971). Figure 8.4 shows the simplest version of that gap.

Role 4: the planner

Chapter 6 covered planning with a world model in detail: imagine candidate futures, score them, act on the best. The same risk applies as in RL, since the planner chooses whichever candidate the model scores highest. V-JEPA 2-AC, from chapter 7, plans in feature space towards image goals (Assran et al. 2025, arXiv 2506.09985), and Cosmos Policy can run as a policy or as a planner (Kim et al. 2026, arXiv 2601.16163).

Plausible is not feasible

Generated framest = 0.0 st = 0.1 st = 0.2 st = 0.3 st = 0.4 st = 0.5 st = 0.6 st = 0.7 st = 0.8 st = 0.9 sWhat an inverse dynamics model reads off them: gripper speed between framesspeed limit 1.2 m/s3.5 m/sEvery frame plausible, but between t = 0.4 and 0.5 s the gripper would have to move at 3.5 m/s.Decoded into commands, the motor either refuses or lurches. An image metric scores both rollouts alike.Violations of the robot’s limits: 3speed above 1.2 m/s, or acceleration above 10 m/s², counted per step
Figure 8.4A rollout can look right in every frame and still ask the robot for a motion it cannot make; only decoding it into actions reveals the difference. Two generated rollouts with the same start and end. Each frame of both is plausible. Switch between them and read the speeds an inverse dynamics model would extract.toy simulation Source: this book's toy, in rust/ch08-simulators; the executability gap is named and used as a reward in EVA (arXiv 2603.17808).

A video model is trained to produce frames that look like real video. Nothing in that objective says the motion between frames must be achievable by a particular robot. EVA calls the mismatch "the executability gap": visually coherent rollouts that "violate rigid-body and kinematic consistency" and decode into unstable or infeasible commands (Wang et al. 2026, arXiv 2603.17808). Its fix is to turn the inverse dynamics model into a reward. Generated videos are scored by the smoothness and legality of the actions they imply, which stays informative "even when generated videos contain severe visual artifacts" (Wang et al. 2026, arXiv 2603.17808). Exercise 8.3 implements the simplest version of that check.

Good at, bad at

What it is good at. It multiplies scarce robot data, enables RL without physical risk, and supports evaluation before deployment. It is also the one family whose output a person can watch: a generated rollout can be inspected frame by frame, which a latent vector (chapter 7) cannot.

What it is bad at. Model bias propagates into whatever is trained in the model. Plausible is not feasible. Multi-view consistency remains weak, which is why VLA-MBPO needs a dedicated mechanism for it (Zhang et al. 2026, arXiv 2603.20607). And it is compute-heavy: generating video is the most expensive operation in this book, which matters when RL needs millions of steps.

In code

The crate's world is one dimension wide, which is what makes the failure easy to see. The true task, the model and the logged data fit in a few lines.

rust/ch08-simulators/src/lib.rsThe pushing task, a learned world model, and the logged data it learns from
/// A robot pushes a block towards a target mark near the table's edge. `theta` in [0, 1] is how
/// hard it pushes. Past the edge the block falls off.
pub const TARGET: f64 = 0.6;
pub const EDGE: f64 = 0.8;

/// How far the block really slides: faster than proportional, because friction matters less at speed.
pub fn true_distance(theta: f64) -> f64 {
    0.3 * theta + 0.9 * theta * theta
}

/// Reward for a slide of `d`: 1 on the mark, falling off linearly, and -1 if the block falls.
pub fn reward(d: f64) -> f64 {
    if d > EDGE { -1.0 } else { 1.0 - (d - TARGET).abs() / TARGET }
}

pub fn true_reward(theta: f64) -> f64 {
    reward(true_distance(theta))
}

/// A learned world model: a straight-line fit of distance against push strength, plus the edge
/// if the data ever showed the block falling.
#[derive(Clone, Copy, Debug, PartialEq)]
pub struct WorldModel {
    pub intercept: f64,
    pub slope: f64,
    /// Push strength beyond which the model predicts a fall; `None` if it never saw one.
    pub edge_theta: Option<f64>,
}

impl WorldModel {
    pub fn distance(&self, theta: f64) -> f64 {
        self.intercept + self.slope * theta
    }
    /// The reward the policy sees when it trains inside the model.
    pub fn reward(&self, theta: f64) -> f64 {
        match self.edge_theta {
            Some(e) if theta > e => -1.0,
            _ => reward(self.distance(theta).min(EDGE)),
        }
    }
}

/// Logged pushes the world model is trained on: `n` evenly spaced strengths from `lo` to `hi`.
pub fn collect(lo: f64, hi: f64, n: usize) -> Vec<(f64, f64)> {
    (0..n)
        .map(|i| {
            let th = if n == 1 { lo } else { lo + (hi - lo) * i as f64 / (n - 1) as f64 };
            (th, true_distance(th))
        })
        .collect()
}

Exercises

The stubs are in rust/ch08-simulators/src/exercises.rs.

Exercise 8.1

Fit the world model

Implement fit_world_model. Then explain why the fitted line is accurate on the logged pushes and too low beyond them, and what kind of model would extrapolate better here. Would it help in general?

Check your answer from the rust/ folder:

cargo test -p ch08-simulators --test exercises ex8_1
Hint 1

Ordinary least squares: slope = Σ(x−x̄)(y−ȳ) / Σ(x−x̄)², intercept = ȳ − slope·x̄.

Hint 2

The edge needs both kinds of evidence: a push that stayed on and a push that fell.

Exercise 8.2

Train in imagination

Implement optimize, a greedy search over push strength inside the model. The tests check that it finds a push the model loves and the world punishes. Change the trust bound and the logged range, and find the smallest logged range for which the unconstrained policy succeeds.

Check your answer from the rust/ folder:

cargo test -p ch08-simulators --test exercises ex8_2
Hint 1

Clamp each candidate to [0, trust] before evaluating it.

Hint 2

Only move when a candidate is strictly better than staying.

Exercise 8.3

Check executability

Implement violations. Then propose one more check a real executability reward should include for a robot arm, and say what data you would need to calibrate it.

Check your answer from the rust/ folder:

cargo test -p ch08-simulators --test exercises ex8_3
Hint 1

Velocities are differences of positions divided by dt; accelerations are differences of velocities divided by dt.

Hint 2

Count speed and acceleration violations separately, then add them.

What we are not sure about

How much synthetic data is worth. The most-quoted numbers, NVIDIA's 40 percent improvement and the survey's "40 percent synthetic matches 100 percent real", are respectively a vendor claim on its own model and a survey without named sources. Controlled, independent studies of the exchange rate between synthetic and real data are still missing.

Whether evaluation in a world model can be trusted. The Interactive World Simulator reports strong correlation with real-world success on its tasks; WorldArena shows that visual quality is a poor proxy for usefulness. Which of these generalises is not yet known.

Whether RL in imagination scales. Every imagination-RL result in this chapter is on a modest set of tasks. The literature names the obstacles: model bias, compounding error, sparse rewards and multi-view inconsistency. None has been shown to disappear with scale.

Further reading

References

25 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.

  1. NVIDIA, A. Ali, J. Bai and 86 others World Simulation with Video Foundation Models for Physical AI. arXiv 2511.00062 v2, 2025-10-28.
  2. J. Bruce, M. Dennis, A. Edwards and 22 others Genie: Generative Interactive Environments. arXiv 2402.15391 v1, 2024-02-23.
  3. D. Hafner, T. Lillicrap, I. Fischer and 4 others Learning Latent Dynamics for Planning from Pixels. arXiv 1811.04551 v5, 2018-11-12.
  4. D. Hafner, T. Lillicrap, J. Ba, M. Norouzi Dream to Control: Learning Behaviors by Latent Imagination. arXiv 1912.01603 v3, 2019-12-03.
  5. D. Hafner, J. Pasukonis, J. Ba, T. Lillicrap Mastering Diverse Domains through World Models. arXiv 2301.04104 v2, 2023-01-10.
  6. J. Jang, S. Ye, Z. Lin and 25 others DreamGen: Unlocking Generalization in Robot Learning through Video World Models. arXiv 2505.12705 v2, 2025-05-19.
  7. NVIDIA Newsroom NVIDIA Announces Isaac GR00T N1, the World's First Open Humanoid Robot Foundation Model, and Simulation Frameworks to Speed Robot Development. nvidianews.nvidia.com, 2025-03-18.
  8. GigaWorld Team, A. Ye, B. Wang and 22 others GigaWorld-0: World Models as Data Engine to Empower Embodied AI. arXiv 2511.19861 v2, 2025-11-25.
  9. GigaBrain Team, A. Ye, B. Wang and 24 others GigaBrain-0: A World Model-Powered Vision-Language-Action Model. arXiv 2510.19430 v3, 2025-10-22.
  10. NVIDIA, H. A. Alhaija, J. Alvarez and 37 others Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control. arXiv 2503.14492 v2, 2025-03-18.
  11. SVRC (Robotics Center of Silicon Valley) State of Robotics 2026. roboticscenter.ai, 2026-03.
  12. X. Liu, Z. Bai, H. Ci and 2 others World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy. arXiv 2602.06508 v2, 2026-02-06.
  13. Z. Zhang, H. Ren, Y. Sun and 6 others Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models. arXiv 2603.20607 v1, 2026-03-21.
  14. X. Cheng, W. Yuan, Z. Mu and 5 others Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization. arXiv 2605.26282 v1, 2026-05-25.
  15. Z. Jiang, K. Liu, Y. Qin and 6 others World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation. arXiv 2509.19080 v2, 2025-09-23.
  16. F. Zhu, Z. Yan, Z. Hong and 3 others WMPO: World Model-based Policy Optimization for Vision-Language-Action Models. arXiv 2511.09515 v1, 2025-11-12.
  17. E. Y. Jia, W. Yuan, T. Shi and 3 others DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models. arXiv 2603.16860 v1, 2026-03-17.
  18. L. Yang, Z. Jiang, C. Sheng, Z. Tang WAM-OPD: On-Policy Distillation for World Action Models. arXiv 2608.22364 v1, 2026-08-23.
  19. Y. Wang, R. Syed, F. Wu and 7 others Interactive World Simulator for Robot Policy Training and Evaluation. arXiv 2603.08546 v1, 2026-03-09.
  20. S. Gao, W. Liang, K. Zheng and 27 others DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos. arXiv 2602.06949 v1, 2026-02-06.
  21. Y. Shang, Z. Li, Y. Ma and 18 others WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models. arXiv 2602.08971 v2, 2026-02-09.
  22. M. Assran, A. Bardes, D. Fan and 26 others V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. arXiv 2506.09985 v1, 2025-06-11.
  23. M. J. Kim, Y. Gao, T. Lin and 8 others Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning. arXiv 2601.16163 v1, 2026-01-22.
  24. R. Wang, Q. Liu, Y. Deng and 3 others EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards. arXiv 2603.17808 v2, 2026-03-18.
  25. Z. Zhang, Z. Li, B. Rahmati and 11 others Do World Action Models Generalize Better than VLAs? A Robustness Study. arXiv 2603.22078 v5, 2026-03-23.