Ch.04Part I · Foundations
The reader's toolkit
The one chapter with real notation: policies, action chunks, world models, latent actions and the four-component decomposition.
- Math
- L2
- Sources
- 17
- Figures
- 4
- Exercises
- 3
- Last checked
- 3 Oct 2026
Equations with a walkthrough
5 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
| Symbol | Read it as | Example |
|---|---|---|
| the observation at time | camera images plus joint angles | |
| everything observed before | the last few frames | |
| the instruction | "put the cup in the sink" | |
| the action at time | target joint positions for one control step | |
| an action chunk of steps | the next half-second of motion | |
| a latent state: a compact summary of the world | a vector inside the model | |
| a latent action: an action-like code inferred from video | what changed between two frames | |
| a hat marks a prediction | , the predicted next frame |
Why this chapter exists
This is the one chapter with real notation. Everything later refers back to it: when chapter 6 says a model "predicts futures and actions jointly" or chapter 7 says it "infers a latent action", the symbols and pictures for those phrases are here. It is kept short and operational. Every equation sits next to a picture or a simulation of the same idea, and nothing requires more than undergraduate probability.
The notation is borrowed from the WAM survey, so that a reader who finishes this book can read the 2026 papers directly (Lu et al. 2026, arXiv 2609.16074).
After this chapter you will be able to:
- read a policy, a world model and a latent action written in the survey's notation, and draw each;
- explain why action chunking helps, using the simulation in figure 4.2, including a 2026 result that overturns the usual explanation;
- explain the difference between emitting actions as tokens and generating them by denoising, and why it matters when there are two right answers.
To address these challenges, we develop a simple yet novel algorithm, Action Chunking with Transformers (ACT), which learns a generative model over action sequences.
The setting
A robot acts in a world it only partly observes. The survey formalises this as a partially observable decision process; for this book the vocabulary matters more than the theory (Lu et al. 2026, arXiv 2609.16074). Figure 4.1 is the vocabulary: every symbol the rest of the book uses.
A policy, and why it predicts chunks
A policy maps what the robot has seen and been told to what it should do. Robot foundation models almost always predict a chunk of actions at once:
Read it left to right: the next actions are a function of the observation history and the instruction. The robot executes the chunk open loop, without looking again, then calls the model for the next one. ACT introduced chunking for robot imitation (Zhao et al. 2023, arXiv 2304.13705), and nearly every model in this book uses it.
Chunking has an obvious cost: during the chunk the robot is blind, so a long chunk reacts late to anything unexpected. Its obvious benefit is amortising a slow model over many control steps, which chapter 2 showed. Figure 4.2 shows both at once. A robot tracks a moving target; each inference takes a few control steps; each chunk extends the model's prediction, which gets slightly worse the further ahead it reaches.
Mean error against chunk length at this latency. The vertical line marks your current setting.
With five steps of latency and the robot stopping to think, the error is lowest around a dozen actions per chunk; at one action per chunk the robot spends most of its time frozen, and at forty it drifts. Thinking while moving cuts the best error by more than half and shifts the sweet spot to shorter chunks. Real-time chunking brings exactly this to flow-based models by generating the next chunk during execution while freezing the actions already committed (Black et al. 2025, arXiv 2506.07339).
A careful 2026 study adds a twist. Testing the usual explanations for why chunking helps imitation, temporal consistency, a shorter effective horizon and better representations, it finds that none of them accounts for the gains. What does is greater non-Markovian expressivity and less compounding error, much of which a policy that acts on observations a few steps old also captures, plus an "implicit ensembling" benefit specific to chunks (Lazzati et al. 2026, arXiv 2608.02547). The toy above shows the deployment side of chunking; that study is the one to read on the learning side.
/// The target the robot must follow.
pub fn target(k: usize) -> f64 {
(0.08 * k as f64).sin()
}
/// How far a prediction j steps ahead is off: learned predictions get worse the further out they reach.
pub const PREDICTION_DRIFT: f64 = 0.004;
/// Follow the target with chunks of `chunk` predicted positions, executed open loop.
/// Inference takes `latency_steps` control steps. Synchronous: the robot holds still while the
/// model thinks. Asynchronous: it keeps executing the previous chunk meanwhile.
/// Returns the mean distance between robot and target.
pub fn tracking_error(chunk: usize, latency_steps: usize, steps: usize, asynchronous: bool) -> f64 {
let chunk = chunk.max(1);
let (mut k, mut pos, mut err) = (0usize, 0.0f64, 0.0f64);
let (mut plan, mut next): (Vec<f64>, usize) = (Vec::new(), 0);
while k < steps {
let t0 = k; // the model sees the world now ...
for _ in 0..latency_steps {
if k >= steps { break; }
if asynchronous && next < plan.len() { pos = plan[next]; next += 1; }
err += (pos - target(k)).abs();
k += 1;
}
// ... and its plan covers the steps after it finishes thinking.
plan = (0..chunk).map(|i| target(t0 + latency_steps + i) + PREDICTION_DRIFT * (latency_steps + i) as f64).collect();
next = 0;
let run = if asynchronous { chunk.saturating_sub(latency_steps).max(1) } else { chunk };
for _ in 0..run {
if k >= steps { break; }
pos = plan[next];
next += 1;
err += (pos - target(k)).abs();
k += 1;
}
}
err / steps as f64
}Two ways to emit actions
A model has to turn its internal representation into numbers a motor can use. There are two families of answer.
Tokens. Discretise each action dimension into bins and let a language-model backbone predict the bin index, as it predicts words. RT-2 discretises each continuous dimension uniformly into 256 bins (Brohan et al. 2023, arXiv 2307.15818); OpenVLA does the same (Kim et al. 2024, arXiv 2406.09246). With bins over a range , a value is replaced by its bin's centre, so the rounding error is at most half a bin:
Tokens fit the backbone's native training objective and make the model easy to fine-tune, at the price of quantisation and of predicting many tokens one by one. FAST compresses an action chunk with the discrete cosine transform before tokenising, so the model emits far fewer tokens (Pertsch et al. 2025, arXiv 2501.09747). In 2026 the tokenizer itself became a research subject: ActionCodec derives design principles from what helps the model learn rather than only reconstruction error (Dong et al. 2026, arXiv 2602.15397), and ActionPiece argues that a good tokenizer must preserve which actions are physically close to which, not only reconstruct each one (Lian et al. 2026, arXiv 2609.18487). A survey organises vision-language-action models entirely by the kind of action token they use (Zhong et al. 2025, arXiv 2507.01925).
Denoising. Start from noise and refine it into an action chunk over steps. Diffusion Policy introduced this for robots, representing the policy as a conditional denoising process (Chi et al. 2023, arXiv 2303.04137); π0 uses flow matching, a close relative, in a separate action expert attached to the VLM (Black et al. 2024, arXiv 2410.24164). Operationally: noise goes in, an action chunk comes out, in steps. Flow matching trains a network to predict, at each point on a straight path from noise to a clean action ,
and sampling follows the predicted direction in small steps, . The survey writes the same interpolation as its equation 9 (Lu et al. 2026, arXiv 2609.16074).
Why bother? Because real tasks often have more than one right answer. To get around an obstacle you can go left or right; going straight is the one wrong choice. A model trained to minimise squared error predicts the average of the demonstrations, which is straight into the obstacle. Denoising with enough steps keeps the two answers apart. Figure 4.3 makes the point with the exact flow for a two-way choice.
With one step every sample lands near zero: the average, the action that hits the obstacle. With a handful of steps the samples split toward the two modes, and with many they settle on them. This is why few-step generation, which chapter 15 needs for speed, must be engineered carefully rather than simply turned down.
/// A one-dimensional action distribution with two modes: steer left (-1) or right (+1) around an
/// obstacle, never straight ahead. Averaging the two is the one thing a policy must not do.
#[derive(Clone, Copy, Debug)]
pub struct TwoModes {
pub mu: f64,
pub sigma: f64,
}
impl TwoModes {
/// Exact velocity of the straight-line (rectified) flow from N(0, 1) noise to this mixture:
/// the expected value of (action - noise) given the current point x at time t. A flow-matching
/// action head learns to approximate exactly this field; here we can write it down.
pub fn velocity(&self, x: f64, t: f64) -> f64 {
let (mut num, mut den) = (0.0, 0.0);
for mu in [-self.mu, self.mu] {
let var = (1.0 - t).powi(2) + t * t * self.sigma * self.sigma;
let w = (-(x - t * mu).powi(2) / (2.0 * var)).exp() / var.sqrt();
let v = mu + (t * self.sigma * self.sigma - (1.0 - t)) / var * (x - t * mu);
num += w * v;
den += w;
}
num / den
}
}A world model
A world model predicts how the world changes when the robot acts:
Read it as: given a summary of the world now and an action, predict the summary after the action. Explicit world models decode the prediction back to pixels, so you can watch what they imagine; implicit ones keep it in latent space and are trained to be useful rather than to look right (Lu et al. 2026, arXiv 2609.16074). PlaNet and Dreamer are the classic latent world models of chapter 3 (Hafner et al. 2018, arXiv 1811.04551; Hafner et al. 2019, arXiv 1912.01603); the video generators of chapter 6 are explicit ones at enormous scale; the JEPA models of chapter 7 are implicit ones at scale.
A latent action
Most video has no action labels. A latent action recovers an action-like signal from the video itself: infer a code from two consecutive frames, then train a forward model to predict the second frame from the first and the code (Ye et al. 2024, arXiv 2410.11758):
Read it as: whatever must have happened between these two frames is, by definition, the latent action. The code is only useful if it is small, so that it cannot simply copy the next frame. Trained on human video, it gives a robot-independent description of motion that a small robot-specific decoder can later map to motor commands. That is how chapter 7's models use the internet's largest data layer.
Four components, two paradigms, two architectures
The 2026 surveys compare architectures with four conditional distributions, and every family in Part II is described as which of the four it trains and which it runs at inference (Lu et al. 2026, arXiv 2609.16074; Zhu et al. 2025, arXiv 2504.02792):
Plan, then act (LingBot-VA, mimic-video, VPP). First imagine the future, then infer the actions that produce it. The action head solves the easier inverse problem. At inference the steps run in the order shown.
Two transition paradigms combine them. Joint prediction generates observations and actions together; inverse dynamics imagines the future first and then infers the actions that produce it. Two architecture patterns carry them. End-to-end models use one backbone for everything; dual systems separate a world module from an action module and connect them by cross-attention or by mixture-of-transformers routing, where each modality keeps its own weights but they share attention. Chapter 6 fills in the resulting two-by-two with eighteen models.
Exercises
The stubs are in rust/ch04-toolkit/src/exercises.rs.
Exercise 4.1
Tokens
Implement encode and decode for uniform binning, and check the half-bin error bound for 256 bins. Then use figure 4.3 to find how many bins it takes before the tokenised trajectory is visually indistinguishable from the smooth one.
Check your answer from the rust/ folder:
cargo test -p ch04-toolkit --test exercises ex4_1Hint 1
Clamp first, then scale the value to [0, bins) and truncate; clamp the index to bins - 1 for the top edge.
Hint 2
The decoded value is the bin's centre, half a bin above its lower edge.
Exercise 4.2
Denoising
Implement sample: integrate the flow's velocity with K Euler steps. With 64 steps, noise slightly right of zero should end near the right-hand mode.
Check your answer from the rust/ folder:
cargo test -p ch04-toolkit --test exercises ex4_2Hint
Time runs from 0 to 1 in K equal steps; use the time at the start of each step.
Exercise 4.3
The regression trap
Implement fraction_near_modes and measure it for one step and for 32. Then explain, in two sentences, why a policy trained with mean squared error drives into the obstacle and a denoising policy does not.
Check your answer from the rust/ folder:
cargo test -p ch04-toolkit --test exercises ex4_3What we are not sure about
Notation is not universal. This chapter follows the WAM survey. The π-series and GR00T papers write the same objects with different symbols, and some papers call a chunk a "trajectory" or a "plan". Appendix B will give the crosswalk.
Why chunking works is still being argued. The 2026 study cited above overturns the textbook explanations for imitation learning, but it is one study, and the deployment benefit shown in figure 4.2 is a separate effect from the learning benefit it measures.
Tokens versus denoising is not settled. Better tokenizers keep closing the gap in precision, and hybrid models train with tokens and act with a denoising head. The two-way choice in figure 4.3 is the clearest case for denoising; on many real tasks the difference is smaller.
Further reading
- The WAM survey, sections II and III, whose notation this chapter uses (Lu et al. 2026, arXiv 2609.16074).
- The concise tutorial from world models to world-action models (Zhang et al. 2026, arXiv 2607.00836).
- ACT for chunking, and the 2026 study of why it works (Zhao et al. 2023, arXiv 2304.13705; Lazzati et al. 2026, arXiv 2608.02547).
- Diffusion Policy and π0 for denoising action heads (Chi et al. 2023, arXiv 2303.04137; Black et al. 2024, arXiv 2410.24164).
- FAST and the action-tokenization survey for the token side (Pertsch et al. 2025, arXiv 2501.09747; Zhong et al. 2025, arXiv 2507.01925).
References
17 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.
- T. Z. Zhao, V. Kumar, S. Levine, C. Finn Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv 2304.13705 v1, 2023-04-23.
- K. Black, M. Y. Galliker, S. Levine Real-Time Execution of Action Chunking Flow Policies. arXiv 2506.07339 v2, 2025-06-09.
- F. Lazzati, K. Stachowicz, W. Chen and 3 others Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control?. arXiv 2608.02547 v1, 2026-08-03.
- A. Brohan, N. Brown, J. Carbajal and 51 others RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv 2307.15818 v1, 2023-07-28.
- M. J. Kim, K. Pertsch, S. Karamcheti and 15 others OpenVLA: An Open-Source Vision-Language-Action Model. arXiv 2406.09246 v3, 2024-06-13.
- K. Pertsch, K. Stachowicz, B. Ichter and 6 others FAST: Efficient Action Tokenization for Vision-Language-Action Models. arXiv 2501.09747 v1, 2025-01-16.
- Z. Dong, Y. Liu, S. Zhang and 8 others ActionCodec: What Makes for Good Action Tokenizers. arXiv 2602.15397 v1, 2026-02-17.
- S. Lian, B. Yu, Z. Shen and 5 others ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models. arXiv 2609.18487 v1, 2026-09-16.
- Y. Zhong, F. Bai, S. Cai and 11 others A Survey on Vision-Language-Action Models: An Action Tokenization Perspective. arXiv 2507.01925 v1, 2025-07-02.
- C. Chi, Z. Xu, S. Feng and 5 others Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. arXiv 2303.04137 v5, 2023-03-07.
- K. Black, N. Brown, D. Driess and 21 others π0: A Vision-Language-Action Flow Model for General Robot Control. arXiv 2410.24164 v4, 2024-10-31.
- D. Hafner, T. Lillicrap, I. Fischer and 4 others Learning Latent Dynamics for Planning from Pixels. arXiv 1811.04551 v5, 2018-11-12.
- D. Hafner, T. Lillicrap, J. Ba, M. Norouzi Dream to Control: Learning Behaviors by Latent Imagination. arXiv 1912.01603 v3, 2019-12-03.
- S. Ye, J. Jang, B. Jeon and 13 others Latent Action Pretraining from Videos. arXiv 2410.11758 v2, 2024-10-15.
- C. Zhu, R. Yu, S. Feng and 3 others Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets. arXiv 2504.02792 v3, 2025-04-03.
- X. Zhang, X. Zeng, W. Zhang From World Models to World Action Models: A Concise Tutorial for Robotics. arXiv 2607.00836 v8, 2026-07-01.