Ch.06Part II · The seven familiesWorld-action models
World-action models
Couple future-state prediction with action generation in one model or one training loop.
- Math
- L2
- Sources
- 34
- Figures
- 7
- Exercises
- 4
- Last checked
- 26 Sept 2026
Equations with a walkthrough
21 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
DreamZero
Starts from a 14B Wan video diffusion model and denoises video and action tokens together in one transformer. Reports real-time closed-loop control at 7 Hz after model and system optimisations.
World Action Models are Zero-shot Policies. First arXiv version 2026-02-17. arXiv 2602.15922
* placement argued either way; select the model to see why.
Why this chapter exists
A reactive policy maps what it sees to what it does. It never asks what happens next. For most short tabletop tasks that is enough, and chapter 5 showed how far it goes. It fails in a characteristic way: an action that looks right now makes a later subgoal unreachable. A gripper closes too early, a puck is pushed too hard, an arm swings into something the camera could not see. The WAM survey lists long-horizon, delayed-effect, occluded and contact-rich tasks as the places where this bites (Lu et al. 2026, arXiv 2609.16074).
World-action models (WAMs) answer with prediction. They predict how the world will change alongside the actions that change it, in one model or in one training loop. The idea is not new. What changed in 2025 and 2026 is that open video generation models became good enough to supply the prediction, so a robotics lab can fine-tune one instead of paying to pretrain a world model from scratch (NVIDIA Technical Blog 2026).
After this chapter you will be able to:
- place a new WAM on the grid of figure 6.1 from its abstract, and say which of the four components it trains and which it runs;
- explain with the simulation in figure 6.3 why imagining consequences helps where reactive control fails, and what that imagination costs;
- read a WAM paper's latency and robustness claims critically, knowing which comparisons are still missing.
My take: WAMs will become the second major recipe for robot foundation models, alongside VLM-based VLAs. The open questions are which formulation of them wins, and which parts of the model architecture and pipeline actually matter. It is likely that the winner is neither pure VLA nor pure WAM, but a hybrid of both.
The mechanism in one line
In the notation of chapter 4, a reactive vision-language-action model predicts an action chunk of steps from the observation history and the instruction :
A world-action model predicts the future observations as well (Lu et al. 2026, arXiv 2609.16074):
Read the second line term by term. The inputs are unchanged. The hat marks a prediction. The only new thing is the extra output : the model has to say what it expects to see, not only what to do. That extra output can be used in three ways. It can be training signal that shapes the model's internal representation and is then thrown away. It can be a plan that the actions are read from. Or it can be a forecast used to check candidate actions before any of them is sent to the motors. Most of this chapter is about which of the three a model uses, and what each one costs.
Four components
The 2026 surveys compare architectures by breaking the generative process into four conditional distributions (Lu et al. 2026, arXiv 2609.16074; Zhu et al. 2025, arXiv 2504.02792):
| Component | Question it answers | Distribution |
|---|---|---|
| Policy | What should I do? | |
| Visual planning | What should the future look like? | |
| Forward dynamics | What happens if I do this? | |
| Inverse dynamics | Which actions produce this future? |
From these, two pathways generate actions. The first models futures and actions jointly:
The second plans, then acts: generate a visual plan, then recover the actions with an inverse-dynamics model (IDM):
The factorisation matters because inverse dynamics is often the easier problem. If you know the outcome you want, inferring the action that produced it is usually simpler than predicting the action from an instruction alone (NVIDIA Technical Blog 2026). The hard part, turning words into the right visual change, moves into the video model, which has already seen a great deal of it.
Plan, then act (LingBot-VA, mimic-video, VPP). First imagine the future, then infer the actions that produce it. The action head solves the easier inverse problem. At inference the steps run in the order shown.
The same decomposition reads naturally as code. The chapter's Rust crate states the four components as traits, and plan-then-act becomes a two-line function over any planner and inverse model that fit them:
/// The four components the 2026 surveys use to compare world-action models.
/// `Obs` is what the robot sees, `Act` what it sends to the motors, `Instr` the task.
pub trait Policy<Obs, Act, Instr> {
/// p(a_{t:t+H} | o_{<t}, l): actions straight from observations (a reactive VLA).
fn act(&self, history: &[Obs], instruction: &Instr) -> Vec<Act>;
}
pub trait ForwardDynamics<Obs, Act> {
/// p(o_{t:t+H} | o_{<t}, a_{t:t+H}): what happens if I do this?
fn imagine(&self, history: &[Obs], actions: &[Act]) -> Vec<Obs>;
}
pub trait InverseDynamics<Obs, Act> {
/// p(a_{t:t+H} | o_{<t}, o_{t:t+H}): which actions produce this future?
fn infer(&self, history: &[Obs], future: &[Obs]) -> Vec<Act>;
}
pub trait VisualPlanner<Obs, Instr> {
/// p(o_{t:t+H} | o_{<t}, l): what should the future look like?
fn plan(&self, history: &[Obs], instruction: &Instr) -> Vec<Obs>;
}/// Plan-then-act (the inverse-dynamics paradigm): imagine a good future, then recover the
/// actions that produce it. Any planner and inverse model that fit the traits will do.
pub fn plan_then_act<O: Clone, A, I>(
planner: &impl VisualPlanner<O, I>,
idm: &impl InverseDynamics<O, A>,
history: &[O],
instruction: &I,
) -> Vec<A> {
let imagined_future = planner.plan(history, instruction); // costly: video generation
idm.infer(history, &imagined_future) // cheap: a small action head
}Why prediction helps: imagine, then act
The simulation below is a deliberately small world. A puck slides on a table with very little friction: each step it keeps 97 percent of its velocity. The goal sits close to the right-hand edge, and an obstacle blocks the straight line to it. The controller sends an acceleration ten times a second.
The reactive controller pushes toward the goal and nothing else. Run it: it hits the obstacle after 15 steps. Nothing in its input tells it about the collision to come, or about the momentum that would carry it past the goal and off the edge even without the obstacle.
The imagine-then-act controller does what a WAM used as a planner does. Before acting it proposes 48 candidate action chunks, imagines each one 30 steps ahead with its world model, scores the imagined futures (end near the goal, slowly, touching nothing), and executes the first 5 steps of the best one. Then it looks again and repeats. It reaches the goal in 43 steps, having imagined 13,230 steps along the way: about 300 imagined steps for every step it actually takes.
- Outcome
- Control steps
- 00.0 s of robot time
- Imagined steps so far
- 00 per control step
Three things are worth trying.
Shorten the horizon. At 5 steps the planner cannot see far enough to route around the obstacle and stop in time; it creeps or stalls. Around 10 to 20 steps it does best. Longer is not automatically better: with a fixed number of random candidates, a 40-step search space is sparser, and the planner does worse. Imagination has to be spent where it pays.
Make the world model wrong. Set the friction error to minus 0.05, so the model imagines the puck stopping sooner than it really does. With 5 steps executed per replan, the planner still succeeds, because every half second it looks at the real puck and corrects. Now raise the steps executed before replanning to 30 and run a few seeds. The puck slides off the table. The table below counts outcomes over 20 random seeds, from cargo run -p ch06-world-action --example seeds:
| Steps executed before replanning | Perfect model | Friction error minus 0.05 | Friction error minus 0.10 |
|---|---|---|---|
| 5 | 20 of 20 reached | 20 of 20 reached | 17 of 20 reached |
| 15 | 20 of 20 reached | 18 of 20 reached | 0 of 20 reached |
| 30 | 18 of 20 reached | 1 of 20 reached | 0 of 20 reached |
That is compounding error in miniature. An imagined future drifts further from reality the further ahead it reaches, and the longer a plan runs without checking the world, the more of that drift reaches the motors. Real WAMs face the same choice. Most execute a fixed number of predicted actions after each inference; recent work argues the robot should instead execute longer while its imagined future still matches what it observes, and replan early when it does not (Wang et al. 2026, arXiv 2605.06222). LingBot-VA builds the same idea in from the start, with closed-loop rollouts that keep taking in real observations (Li et al. 2026, arXiv 2601.21998).
Try the third controller. It never imagines anything at test time. It was trained, once, from six episodes of the imagine-then-act controller, and at test time it copies what that teacher did in the most similar remembered states. It reaches the goal with zero imagined steps. This is the question Fast-WAM asks of real models, in miniature: is the value of imagination in running it, or in having learned from it (Yuan et al. 2026, arXiv 2603.16666)? The toy cannot answer that for large models. It only shows that the answer is not obvious: the student works near the situations it was trained on and has no way to check itself anywhere else.
/// Score an imagined future: end near the goal, slowly, and never hit anything on the way.
pub fn imagined_cost(world: &World, future: &[State]) -> f32 {
let mut cost = 0.0;
for (t, s) in future.iter().enumerate() {
match world.status(s) {
Outcome::Collided | Outcome::FellOff => return 1_000.0 + (future.len() - t) as f32,
_ => {}
}
cost += 0.05 * (s.x - world.goal.0).hypot(s.y - world.goal.1);
}
let end = future.last().copied().unwrap_or_default();
cost + 3.0 * (end.x - world.goal.0).hypot(end.y - world.goal.1) + 1.5 * end.speed()
}Two ways to couple prediction and action
Figure 6.1 sorts models on two axes. This section takes the rows; the next takes the columns.
Joint prediction: futures and actions together
In a joint model, one generative process produces both the future and the actions, so each shapes the other. The early versions trained a transformer to predict future frames and action chunks together; the modern versions start from a large pretrained video generator (Lu et al. 2026, arXiv 2609.16074).
DreamZero is the scaled-up case. It starts from Wan 2.1, a 14-billion-parameter image-to-video diffusion model (Team Wan et al. 2025, arXiv 2503.20314), and turns it into a joint world-action model: a single diffusion transformer denoises video tokens and action tokens together, with no separate inverse-dynamics module (Ye et al. 2026, arXiv 2602.15922; NVIDIA Technical Blog 2026). Its authors report over twice the generalisation to new tasks and environments of state-of-the-art VLAs in real-robot tests, and real-time control at 7 Hz after model and system optimisationsvendor claim.
Cosmos Policy takes a different route to the same end. Instead of adding action tokens and an action head, it writes the action into something the video model already knows how to generate: an extra latent frame. Future states and values (expected cumulative reward) are encoded as latent frames too, which lets the one model act directly or plan, by proposing actions and scoring their predicted futures (Kim et al. 2026, arXiv 2601.16163). Its authors report 98.5 percent average success on LIBERO and 67.1 percent on RoboCasavendor claim.
Other joint designs fill in the space between. UWM denoises video and actions with independent noise levels, so one trained model supports several inference modes (Zhu et al. 2025, arXiv 2504.02792). PAD was an early joint denoiser of images and actions (Guo et al. 2024, arXiv 2411.18179). WorldVLA puts text, images and actions into one autoregressive token space (Cen et al. 2025, arXiv 2506.21539). Motus adds structured joint attention with separate noise schedules for video and action (Bi et al. 2025, arXiv 2512.13030), and MotuBrain extends it into a three-stream model that serves as policy, world model, video generator and inverse dynamics at once (Motubrain Team et al. 2026, arXiv 2604.27792).
Inverse dynamics: imagine the future, then infer the action
The second row generates a visual plan first and reads the actions off it. The recipe goes back at least to 2023, when UniPi used a text-conditioned video generator as a high-level plan, but the video models of that era had to be trained from scratch at a cost few robotics labs could pay (NVIDIA Technical Blog 2026).
LingBot-VA is the modern form. It fine-tunes Wan 2.2-5B through 16,000 hours of cross-embodiment robot pretraining, predicts frames and actions in a shared latent space, and runs action prediction and motor execution in parallel to keep up with control (Li et al. 2026, arXiv 2601.21998; NVIDIA Technical Blog 2026). Its July 2026 successor argues that video models built for digital content are the wrong starting point, and pretrains a video-action model for embodiment from the ground up (Zhang et al. 2026, arXiv 2607.08639).
Several models skip the final RGB video and hand intermediate video-model features to the action decoder. mimic-video pairs an internet-scale video model with a flow-matching decoder that acts as an inverse-dynamics model on video latents (Pai et al. 2025, arXiv 2512.15692); Video Prediction Policy conditions a diffusion policy head on the representations of a text-guided video predictor (Hu et al. 2024, arXiv 2412.14803). Seer builds the inverse-dynamics interface into a single transformer by letting action tokens attend only to the current and predicted future states (Tian et al. 2024, arXiv 2412.15109), and Act2Goal predicts goal states rather than only the next frame (Zhou et al. 2025, arXiv 2512.23541). The survey also reads π0.7, which chapter 5 files as a reactive VLA, as a goal-conditioned WAM, because it predicts actions toward a goal image (Lu et al. 2026, arXiv 2609.16074; Physical Intelligence et al. 2026, arXiv 2604.15483). The boundary between the two families is thin, and chapter 5 says so too.
A third option: video only in training
The most consequential result of 2026 may be the one that questions the premise. Fast-WAM keeps video co-training while the model learns, then skips future prediction at test time. Across controlled variants it stays competitive with imagine-then-execute versions, while removing video co-training costs much more; it runs with 190 ms latency, which its authors put at over four times faster than imagine-then-execute WAMsvendor claim. Their reading: the main value of video prediction may lie in better representations learned during training, not in futures generated at test time (Yuan et al. 2026, arXiv 2603.16666).
Others reach the same design from different directions. GigaWorld-Policy predicts actions first and video second, with a causal mask so video tokens never influence actions, which makes video generation optional at inference (Ye et al. 2026, arXiv 2603.17240); version 0.5 reports 85 ms per inference on a single RTX 4090vendor claim (GigaWorld Team et al. 2026, arXiv 2607.13960). ImageWAM swaps the video generator for an image-editing model and never decodes the target image at all, reporting a quarter of the latency of video-based WAMsvendor claim (Zhang et al. 2026, arXiv 2606.19531). DreamWAM supervises the future as motion, geometry and semantics as well as RGB during training, and switches those branches off for deployment (Yuan et al. 2026, arXiv 2608.04996).
The debate is live. A privileged-foresight study argues that joint training does more than regularise: seeing the future during training imposes a correction on action generation that current-only policies capture only partially (Fang et al. 2026, arXiv 2604.25859). A September 2026 paper questions whether WAMs should be tied to large video-generation backbones at all (Fang et al. 2026, arXiv 2609.27455).
One backbone or two
The columns of figure 6.1 are about plumbing. End-to-end models process language, observations and actions in one backbone; DreamZero and Cosmos Policy are monolithic in this sense, and gain strong coupling between video and action at the price of one set of weights serving dense visual tokens and sparse action targets. Dual systems separate a world module from an action module and connect them through an interface (Lu et al. 2026, arXiv 2609.16074).
Two interfaces dominate. The first is Mixture-of-Transformers routing: separate weights per modality, fused through shared self-attention in every layer, as in LingBot-VA and Fast-WAM. The second is cross-attention, where an action expert reads the world module's features, as in mimic-video and, on the VLA side, π0 and GR00T. The NVIDIA post calls Mixture-of-Transformers the current default in both VLAs and recent WAMs, and expects it to become the dominant WAM architecture as a practical compromise between modularity and coupling (NVIDIA Technical Blog 2026).
The price of imagination
Generating video to decide on an action is expensive. The cost has three sources: the number of denoising steps, the number of future frames generated, and the size of the video backbone. A reactive VLA pays none of them.
Measured numbers exist, but they are not comparable with each other: each is the authors' own, on their own hardware and task. Read them as orders of magnitude.
| System | Reported speed | Conditions |
|---|---|---|
| WAMs that generate video at test time | 590 to 800 ms per action chunk vendor claim | joint prediction and inverse dynamics modes, as measured in the Fast-WAM study (NVIDIA Technical Blog 2026) |
| π0.5 (reactive VLA) | about 190 ms per chunk vendor claim | same comparison |
| Fast-WAM | 190 ms vendor claim | no future generation at test time (Yuan et al. 2026, arXiv 2603.16666) |
| DreamZero | 7 Hz closed loop vendor claim | 14B model after model and system optimisations (Ye et al. 2026, arXiv 2602.15922) |
| MotuBrain | up to 11 Hz vendor claim | step reduction, compilation, FP8, DiT caching, action-only inference (Motubrain Team et al. 2026, arXiv 2604.27792) |
| GigaWorld-Policy-0.5 | 85 ms vendor claim | action-only inference, one RTX 4090 (GigaWorld Team et al. 2026, arXiv 2607.13960) |
Training costs more too. A rough lower-bound estimate in the NVIDIA post puts DreamZero's action tuning at about 9 zettaFLOPs, and a full Wan-scale video pretraining plus action tuning at about 51, against about 6.9 for a small VLA trained from scratch (NVIDIA Technical Blog 2026). A strong video prior can reduce the robot data needed; in practice it often trades robot-data efficiency for compute.
The planner in figure 6.3 makes the same trade visible: its "imagined futures per replan" and "horizon" sliders are the test-time compute budget. Test-time scaling work on real WAMs asks when that extra budget is worth spending, ranking sampled rollouts by the geometric consistency of their predicted futures and spending the extra compute only when it is likely to help (Zhao et al. 2026, arXiv 2607.17454).
What the evidence says so far
Standard simulation suites no longer separate the best models: reported LIBERO success rates for recent WAMs sit near 98 percent (Kim et al. 2026, arXiv 2601.16163; Yuan et al. 2026, arXiv 2608.04996). The differences show under perturbation and on real robots.
The most direct comparison so far is a March 2026 robustness study that runs recent WAMs and VLAs on LIBERO-Plus and RoboTwin 2.0-Plus under visual and language perturbations. WAMs come out strong: LingBot-VA reaches 74.2 percent on RoboTwin 2.0-Plus and Cosmos Policy 82.2 percent on LIBERO-Plus. π0.5 matches them on some tasks, but only with extensive training on diverse robot data and varied objectives (Zhang et al. 2026, arXiv 2603.22078). The study is in simulation, and it is one study.
RoboArena runs pairwise, double-blind comparisons of policies on real DROID robots at many institutions (Atreya et al. 2025, arXiv 2506.18123; Khazatsky et al. 2024, arXiv 2403.12945). In the April 2026 snapshot DreamZero led π0.5 by 128 points, trained only on DROID data without a large cross-embodiment stage (NVIDIA Technical Blog 2026).
The makers' own numbers point the same way, and should be read as claims. DreamZero reports that video-only demonstrations from other robots or from humans, just 10 to 20 minutes of them, improve performance on unseen tasks by over 42 percent relative, and that it adapts to a new embodiment from 30 minutes of play datavendor claim (Ye et al. 2026, arXiv 2602.15922). GigaWorld-Policy reports a 95 percent improvement over π0.5 on RoboTwin 2.0vendor claim (Ye et al. 2026, arXiv 2603.17240). NVIDIA previewed GR00T N2 in March 2026 as built on DreamZero research and succeeding at new tasks in new environments more than twice as often as leading VLAs, with availability slated for the end of the yearvendor claim (NVIDIA Newsroom 2026).
Where world-action models are used
Manipulation is the home ground: nearly every model in figure 6.1 is evaluated on arm and bimanual tasks. The idea is spreading. DreamTrajectory brings trajectory-guided action generation with world-model alignment to mobile manipulation (Yang et al. 2026, arXiv 2608.01381), and Motus2 turns a world model into a self-evolving system for dexterous manipulation, with one set of weights exposing several control interfaces and a closed loop for improving its own policy (Bi et al. 2026, arXiv 2608.30237). Chapter 16 covers navigation and driving, where world models are used mostly for look-ahead planning and scenario generation. Chapter 8 covers the neighbouring family, where the world model trains or evaluates a separate policy instead of acting itself.
Build it yourself in Rust
The crate rust/ch06-world-action contains everything the simulation runs. Its imperfect world model has a single, controllable error:
/// A learned forward model is never exact. Here its only error is a friction bias:
/// `damping_error < 0` imagines more friction than there is, so imagined pucks stop early.
#[derive(Clone, Copy, Debug, Default)]
pub struct ForwardModel {
pub damping_error: f32,
}
impl ForwardModel {
pub fn perfect() -> Self {
ForwardModel { damping_error: 0.0 }
}
/// Roll an action chunk forward in imagination, starting from `s`.
pub fn rollout(&self, s: &State, actions: &[Action]) -> Vec<State> {
let damping = (DAMPING + self.damping_error).clamp(0.0, 1.0);
let mut out = Vec::with_capacity(actions.len());
let mut cur = *s;
for a in actions {
cur = step_with_damping(damping, &cur, a);
out.push(cur);
}
out
}
}The planner is random shooting with a receding horizon, the simplest version of the propose, imagine, choose loop. Besides random chunks it always considers coasting, braking, and a proposal from a simple approach controller, each judged by imagination like the rest:
/// Imagine-then-act: sample candidate action chunks, imagine each with the forward model,
/// execute the start of the best one, repeat (receding horizon).
pub struct ShootingPlanner {
pub model: ForwardModel,
pub config: PlannerConfig,
rng: Rng,
queue: Vec<Action>,
best: Vec<Action>,
calls: u64,
/// The candidates imagined at the last replan and the index of the chosen one.
pub last_imagined: Vec<Vec<State>>,
pub last_choice: usize,
}
impl ShootingPlanner {
pub fn new(model: ForwardModel, config: PlannerConfig, seed: u64) -> Self {
ShootingPlanner {
model,
config,
rng: Rng::new(seed),
queue: Vec::new(),
best: Vec::new(),
calls: 0,
last_imagined: Vec::new(),
last_choice: 0,
}
}
fn random_chunk(&mut self) -> Vec<Action> {
let seg = self.config.horizon.div_ceil(self.config.knots);
let mut chunk = Vec::with_capacity(self.config.horizon);
for _ in 0..self.config.knots {
let a = Action { ax: self.rng.uniform(-A_MAX, A_MAX), ay: self.rng.uniform(-A_MAX, A_MAX) };
chunk.extend(std::iter::repeat_n(a, seg));
}
chunk.truncate(self.config.horizon);
chunk
}
fn replan(&mut self, world: &World, s: &State) {
let h = self.config.horizon;
let mut candidates: Vec<Vec<Action>> = Vec::with_capacity(self.config.samples + 1);
// Warm start: the rest of the previous best chunk, padded by holding its last action.
if self.best.len() > self.config.execute {
let mut warm: Vec<Action> = self.best[self.config.execute..].to_vec();
let last = *warm.last().unwrap();
warm.resize(h, last);
candidates.push(warm);
}
// Two defaults every shooting planner keeps: coast, and brake against the current motion.
candidates.push(vec![Action::default(); h]);
let v = s.speed().max(1e-6);
let brake = Action { ax: -A_MAX * s.vx / v, ay: -A_MAX * s.vy / v };
candidates.push(vec![brake; (v / (A_MAX * DT)).ceil().min(h as f32) as usize]);
candidates.last_mut().unwrap().resize(h, Action::default());
// A proposal from a simple approach controller, played out in imagination and then
// judged like every other candidate: a policy proposes, the world model checks.
let damping = (DAMPING + self.model.damping_error).clamp(0.0, 1.0);
let (mut pd, mut cur) = (Vec::with_capacity(h), *s);
for _ in 0..h {
let a = Action { ax: 2.0 * (world.goal.0 - cur.x) - 2.5 * cur.vx, ay: 2.0 * (world.goal.1 - cur.y) - 2.5 * cur.vy }.clamped();
cur = step_with_damping(damping, &cur, &a);
pd.push(a);
}
self.calls += h as u64;
candidates.push(pd);
while candidates.len() < self.config.samples {
let c = self.random_chunk();
candidates.push(c);
}
let mut best_cost = f32::INFINITY;
self.last_imagined.clear();
for (i, c) in candidates.iter().enumerate() {
let future = self.model.rollout(s, c);
self.calls += c.len() as u64;
let cost = imagined_cost(world, &future);
if cost < best_cost {
best_cost = cost;
self.last_choice = i;
}
self.last_imagined.push(future);
}
self.best = candidates.swap_remove(self.last_choice);
self.queue = self.best.iter().take(self.config.execute).rev().copied().collect();
}
}
impl Controller for ShootingPlanner {
fn control(&mut self, world: &World, s: &State) -> Action {
if self.queue.is_empty() {
self.replan(world, s);
}
self.queue.pop().unwrap_or_default()
}
fn model_calls(&self) -> u64 {
self.calls
}
}The distilled student keeps what the planner did and drops the planner:
/// Fast-WAM's question in miniature: keep imagination for training, drop it at test time.
/// The student memorises what the imagine-then-act teacher did and, at test time, copies the
/// action taken in the most similar remembered states. It never calls the world model.
pub struct DistilledPolicy {
memory: Vec<(State, Action)>,
}
pub fn distill(mut teacher: ShootingPlanner, starts: &[State], max_steps: usize) -> DistilledPolicy {
let world = World::default();
let mut memory = Vec::new();
for start in starts {
let mut s = *start;
for _ in 0..max_steps {
let a = teacher.control(&world, &s);
memory.push((s, a));
s = world.step(&s, &a);
if world.status(&s) != Outcome::Running {
break;
}
}
teacher.queue.clear();
teacher.best.clear();
}
DistilledPolicy { memory }
}
impl Controller for DistilledPolicy {
fn control(&mut self, _world: &World, s: &State) -> Action {
// Inverse-distance weighting over the three nearest remembered states.
let dist = |m: &State| {
((m.x - s.x).powi(2) + (m.y - s.y).powi(2) + 0.25 * ((m.vx - s.vx).powi(2) + (m.vy - s.vy).powi(2))).sqrt()
};
let mut nearest: Vec<(f32, Action)> = self.memory.iter().map(|(m, a)| (dist(m), *a)).collect();
nearest.sort_by(|a, b| a.0.total_cmp(&b.0));
let (mut wsum, mut ax, mut ay) = (0.0, 0.0, 0.0);
for (d, a) in nearest.iter().take(3) {
let w = 1.0 / (d + 1e-3);
wsum += w;
ax += w * a.ax;
ay += w * a.ay;
}
Action { ax: ax / wsum, ay: ay / wsum }.clamped()
}
}Run the comparisons from the rust/ folder with cargo run -p ch06-world-action --example summary and cargo run -p ch06-world-action --example seeds.
Exercises
The stubs are in rust/ch06-world-action/src/exercises.rs. Reference solutions are behind a feature flag, cargo test -p ch06-world-action --features solutions, if you want to check your reasoning afterwards.
Exercise 6.1
Recover the action
Implement inverse_dynamics(prev, next), which returns the action that turned one real state into the next. This is the inverse-dynamics component of figure 6.2 in its simplest form: exact, because the toy's physics is known. In a WAM the same function is a learned network, and it has to cope with imagined states that no real action could produce.
Check your answer from the rust/ folder:
cargo test -p ch06-world-action --test exercises ex6_1Hint 1
The velocity update is linear in the action. Solve it for the action.
Hint 2
Position does not appear in the answer: the velocity change already carries all the information.
Exercise 6.2
Watch imagination drift
Implement drift: execute the same action chunk in the real world and in a model's imagination, return the distance between the two pucks after every step, and the first step at which it exceeds a tolerance. Then use it to answer a question the chapter only gestured at: with a friction error of minus 0.03, how many steps ahead can the planner trust its imagination if the goal is 0.4 units wide?
Check your answer from the rust/ folder:
cargo test -p ch06-world-action --test exercises ex6_2Hint 1
Roll the model forward once with ForwardModel::rollout; step the real world yourself, one action at a time.
Hint 2
position returns the first index that satisfies a predicate.
Exercise 6.3
Choose by imagining
Implement choose_chunk, the core of the planner: imagine every candidate action chunk and return the index of the one whose future scores best. The tests check that it brakes before the table edge and goes around the obstacle rather than through it.
Check your answer from the rust/ folder:
cargo test -p ch06-world-action --test exercises ex6_3Hint 1
imagined_cost already knows what a good future is. Your job is only to imagine each candidate and compare.
Hint 2
Keep the lower index on ties by comparing with a strict less-than.
Exercise 6.4
The price of imagination
Implement sustainable_rate_hz: given the candidates imagined per replan, the horizon, the cost of one imagined step, a fixed overhead and the steps executed per replan, return the highest control rate the planner can sustain. Then find the smallest number of candidates and horizon that still reach the goal in figure 6.3, and compute the rate they allow at 0.1 ms per imagined step.
Check your answer from the rust/ folder:
cargo test -p ch06-world-action --test exercises ex6_4Hint
Work out the milliseconds per replan first, then how many control steps each replan pays for.
What we are not sure about
Whether imagination is needed at test time. Fast-WAM, GigaWorld-Policy and ImageWAM suggest that on today's benchmarks most of the benefit survives without generating futures during deployment. A privileged-foresight study argues that something is lost, and adaptive-execution work shows that imagined futures are useful as a check on reality. The evidence on each side is mostly simulated, and none of it is a matched comparison on real robots.
Whether WAMs are more robust than VLAs. One robustness study, in simulation, says WAMs reach strong robustness with less diverse training data. One real-robot leaderboard, on one platform, puts a WAM first. Neither is the independent, matched, real-world comparison the question needs.
Whether the taxonomy will hold. The family is barely a year old as a named category. Figure 6.1 follows the WAM survey; the NVIDIA post uses three different axes, including a "representation-only" paradigm that the 2x2 has no cell for; a second survey organises the same models in two other views (Shen et al. 2026, arXiv 2606.20781). Several placements are argued either way. Expect this chapter to be redrawn; it is dated for that reason.
Whether the toy transfers. The simulation in this chapter shows mechanisms, not the performance of any system. A puck on a table has four state variables; a robot's world has millions. The lessons about horizon, drift and replanning are qualitative.
Further reading
- The WAM survey this chapter follows, especially sections II, III and VI (Lu et al. 2026, arXiv 2609.16074).
- The NVIDIA post for a practitioner's map of the field and its costs (NVIDIA Technical Blog 2026).
- Fast-WAM, the key paper above (Yuan et al. 2026, arXiv 2603.16666).
- DreamZero and Cosmos Policy, the two joint-prediction designs of figure 6.5 (Ye et al. 2026, arXiv 2602.15922; Kim et al. 2026, arXiv 2601.16163).
- The robustness study, the closest thing yet to a matched comparison (Zhang et al. 2026, arXiv 2603.22078).
- A second survey with a different taxonomy (Shen et al. 2026, arXiv 2606.20781), and the concise tutorial from world models to WAMs (Zhang et al. 2026, arXiv 2607.00836).
References
34 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.
- NVIDIA Technical Blog Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models. developer.nvidia.com, 2026-06-15.
- C. Zhu, R. Yu, S. Feng and 3 others Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets. arXiv 2504.02792 v3, 2025-04-03.
- R. Wang, Y. Zhang, J. Lin and 4 others When to Trust Imagination: Adaptive Action Execution for World Action Models. arXiv 2605.06222 v2, 2026-05-07.
- L. Li, Q. Zhang, Y. Luo and 9 others Causal World Modeling for Robot Control. arXiv 2601.21998 v2, 2026-01-29.
- T. Yuan, Z. Dong, Y. Liu, H. Zhao Fast-WAM: Do World Action Models Need Test-time Future Imagination?. arXiv 2603.16666 v2, 2026-03-17.
- Team Wan, A. Wang, B. Ai and 59 others Wan: Open and Advanced Large-Scale Video Generative Models. arXiv 2503.20314 v2, 2025-03-26.
- S. Ye, Y. Ge, K. Zheng and 33 others World Action Models are Zero-shot Policies. arXiv 2602.15922 v1, 2026-02-17.
- M. J. Kim, Y. Gao, T. Lin and 8 others Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning. arXiv 2601.16163 v1, 2026-01-22.
- Y. Guo, Y. Hu, J. Zhang and 4 others Prediction with Action: Visual Policy Learning via Joint Denoising Process. arXiv 2411.18179 v1, 2024-11-27.
- J. Cen, C. Yu, H. Yuan and 9 others WorldVLA: Towards Autoregressive Action World Model. arXiv 2506.21539 v1, 2025-06-26.
- H. Bi, H. Tan, S. Xie and 13 others Motus: A Unified Latent Action World Model. arXiv 2512.13030 v2, 2025-12-15.
- Motubrain Team, C. Xiang, F. Bao and 17 others Motubrain: An Advanced World Action Model for Robot Control. arXiv 2604.27792 v5, 2026-04-30.
- Q. Zhang, L. Li, L. Zhang and 26 others Native Video-Action Pretraining for Generalizable Robot Control. arXiv 2607.08639 v2, 2026-07-09.
- J. Pai, L. Achenbach, V. Montesinos and 3 others mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs. arXiv 2512.15692 v2, 2025-12-17.
- Y. Hu, Y. Guo, P. Wang and 6 others Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations. arXiv 2412.14803 v2, 2024-12-19.
- Y. Tian, S. Yang, J. Zeng and 4 others Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation. arXiv 2412.15109 v1, 2024-12-19.
- P. Zhou, L. Chen, S. Chen and 5 others Act2Goal: From World Model To General Goal-conditioned Policy. arXiv 2512.23541 v1, 2025-12-29.
- Physical Intelligence, B. Ai, A. Amin and 85 others π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities. arXiv 2604.15483 v2, 2026-04-16.
- A. Ye, B. Wang, C. Ni and 21 others GigaWorld-Policy: An Efficient Action-Centered World--Action Model. arXiv 2603.17240 v2, 2026-03-18.
- GigaWorld Team, A. Ye, A. Ma and 26 others GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch. arXiv 2607.13960 v3, 2026-07-15.
- Y. Zhang, W. Zhang, Z. Qi and 7 others ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?. arXiv 2606.19531 v1, 2026-06-17.
- S. Yuan, W. Zhao, X. Shi and 6 others DreamWAM: Beyond RGB Future Prediction for World Action Models. arXiv 2608.04996 v1, 2026-08-05.
- P. Fang, H. Chen, X. Cai Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models. arXiv 2604.25859 v2, 2026-04-28.
- X. Fang, B. Duan, H. Wu and 2 others Latent evolving World Action Model. arXiv 2609.27455 v2, 2026-09-23.
- Z. Zhao, M. Cho, H. shen and 4 others Test-Time Scaling for World Action Models via Zero-Shot Geometric Evaluation. arXiv 2607.17454 v1, 2026-07-20.
- Z. Zhang, Z. Li, B. Rahmati and 11 others Do World Action Models Generalize Better than VLAs? A Robustness Study. arXiv 2603.22078 v5, 2026-03-23.
- P. Atreya, K. Pertsch, T. Lee and 29 others RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies. arXiv 2506.18123 v2, 2025-06-22.
- A. Khazatsky, K. Pertsch, S. Nair and 98 others DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset. arXiv 2403.12945 v2, 2024-03-19.
- NVIDIA Newsroom NVIDIA and Global Robotics Leaders Take Physical AI to the Real World (GTC 2026, GR00T N2 preview). nvidianews.nvidia.com, 2026-03-16.
- Z. Yang, W. Zhang, X. Chen and 9 others DreamTrajectory: Trajectory-Guided Action Generation with World Model Alignment for Mobile Manipulation. arXiv 2608.01381 v2, 2026-08-02.
- H. Bi, Z. Zhou, Y. Tang and 16 others Motus2: A Self-Evolving General World Model for Dexterous Manipulation. arXiv 2608.30237 v2, 2026-08-31.
- Q. Shen, S. Zhang, Y. Liao and 5 others World Action Models: A Survey. arXiv 2606.20781 v1, 2026-06-18.
- X. Zhang, X. Zeng, W. Zhang From World Models to World Action Models: A Concise Tutorial for Robotics. arXiv 2607.00836 v8, 2026-07-01.