Skip to content

Ch.04Part I · Foundations

The reader's toolkit

The one chapter with real notation: policies, action chunks, world models, latent actions and the four-component decomposition.

Math
L2

Equations with a walkthrough

Sources
17

5 papers from 2026

Figures
4

Most are interactive

Exercises
3

Rust, with tests

Last checked
3 Oct 2026

The field moves monthly

Papers cited, by half-year of first arXiv version

SymbolRead it asExample
oto_tthe observation at time ttcamera images plus joint angles
o<to_{<t}everything observed before ttthe last few frames
ℓ\ellthe instruction"put the cup in the sink"
ata_tthe action at time tttarget joint positions for one control step
at:t+Ha_{t:t+H}an action chunk of HH stepsthe next half-second of motion
ztz_ta latent state: a compact summary of the worlda vector inside the model
utu_ta latent action: an action-like code inferred from videowhat changed between two frames
 ^\hat{\ }a hat marks a predictiono^t+1\hat o_{t+1}, the predicted next frame
Figure 4.1Eight symbols carry the whole book: what the robot sees, is told and does, now and in the future, and the compact codes in between. The notation card. Every later chapter uses these symbols with these meanings. Source: notation of Lu et al. 2026 (arXiv 2609.16074), section II.

Why this chapter exists

This is the one chapter with real notation. Everything later refers back to it: when chapter 6 says a model "predicts futures and actions jointly" or chapter 7 says it "infers a latent action", the symbols and pictures for those phrases are here. It is kept short and operational. Every equation sits next to a picture or a simulation of the same idea, and nothing requires more than undergraduate probability.

The notation is borrowed from the WAM survey, so that a reader who finishes this book can read the 2026 papers directly (Lu et al. 2026, arXiv 2609.16074).

After this chapter you will be able to:

  • read a policy, a world model and a latent action written in the survey's notation, and draw each;
  • explain why action chunking helps, using the simulation in figure 4.2, including a 2026 result that overturns the usual explanation;
  • explain the difference between emitting actions as tokens and generating them by denoising, and why it matters when there are two right answers.

To address these challenges, we develop a simple yet novel algorithm, Action Chunking with Transformers (ACT), which learns a generative model over action sequences.

Tony Z. Zhao, Vikash Kumar, Sergey Levine and Chelsea Finn, authors of the ACT paper. Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (arXiv 2304.13705), 23 April 2023.

The setting

A robot acts in a world it only partly observes. The survey formalises this as a partially observable decision process; for this book the vocabulary matters more than the theory (Lu et al. 2026, arXiv 2609.16074). Figure 4.1 is the vocabulary: every symbol the rest of the book uses.

A policy, and why it predicts chunks

A policy maps what the robot has seen and been told to what it should do. Robot foundation models almost always predict a chunk of HH actions at once:

at:t+H=f(o<t,ℓ)a_{t:t+H} = f(o_{<t}, \ell)

Read it left to right: the next HH actions are a function of the observation history and the instruction. The robot executes the chunk open loop, without looking again, then calls the model for the next one. ACT introduced chunking for robot imitation (Zhao et al. 2023, arXiv 2304.13705), and nearly every model in this book uses it.

Chunking has an obvious cost: during the chunk the robot is blind, so a long chunk reacts late to anything unexpected. Its obvious benefit is amortising a slow model over many control steps, which chapter 2 showed. Figure 4.2 shows both at once. A robot tracks a moving target; each inference takes a few control steps; each chunk extends the model's prediction, which gets slightly worse the further ahead it reaches.

Mean error against chunk length at this latency. The vertical line marks your current setting.

Figure 4.2With a slow model there is a best chunk length: short chunks waste time thinking, long ones drift; thinking while moving beats both. Top: the target (dashed) and the robot (solid) over 240 steps; shaded steps are inference. Bottom: mean tracking error against chunk length at the current latency, for both execution modes. Change the chunk length and latency.toy simulation Source: rust/ch04-toolkit, tracking_error; the same model runs in the page.

With five steps of latency and the robot stopping to think, the error is lowest around a dozen actions per chunk; at one action per chunk the robot spends most of its time frozen, and at forty it drifts. Thinking while moving cuts the best error by more than half and shifts the sweet spot to shorter chunks. Real-time chunking brings exactly this to flow-based models by generating the next chunk during execution while freezing the actions already committed (Black et al. 2025, arXiv 2506.07339).

A careful 2026 study adds a twist. Testing the usual explanations for why chunking helps imitation, temporal consistency, a shorter effective horizon and better representations, it finds that none of them accounts for the gains. What does is greater non-Markovian expressivity and less compounding error, much of which a policy that acts on observations a few steps old also captures, plus an "implicit ensembling" benefit specific to chunks (Lazzati et al. 2026, arXiv 2608.02547). The toy above shows the deployment side of chunking; that study is the one to read on the learning side.

rust/ch04-toolkit/src/lib.rsThe chunking model behind figure 4.2
/// The target the robot must follow.
pub fn target(k: usize) -> f64 {
    (0.08 * k as f64).sin()
}

/// How far a prediction j steps ahead is off: learned predictions get worse the further out they reach.
pub const PREDICTION_DRIFT: f64 = 0.004;

/// Follow the target with chunks of `chunk` predicted positions, executed open loop.
/// Inference takes `latency_steps` control steps. Synchronous: the robot holds still while the
/// model thinks. Asynchronous: it keeps executing the previous chunk meanwhile.
/// Returns the mean distance between robot and target.
pub fn tracking_error(chunk: usize, latency_steps: usize, steps: usize, asynchronous: bool) -> f64 {
    let chunk = chunk.max(1);
    let (mut k, mut pos, mut err) = (0usize, 0.0f64, 0.0f64);
    let (mut plan, mut next): (Vec<f64>, usize) = (Vec::new(), 0);
    while k < steps {
        let t0 = k; // the model sees the world now ...
        for _ in 0..latency_steps {
            if k >= steps { break; }
            if asynchronous && next < plan.len() { pos = plan[next]; next += 1; }
            err += (pos - target(k)).abs();
            k += 1;
        }
        // ... and its plan covers the steps after it finishes thinking.
        plan = (0..chunk).map(|i| target(t0 + latency_steps + i) + PREDICTION_DRIFT * (latency_steps + i) as f64).collect();
        next = 0;
        let run = if asynchronous { chunk.saturating_sub(latency_steps).max(1) } else { chunk };
        for _ in 0..run {
            if k >= steps { break; }
            pos = plan[next];
            next += 1;
            err += (pos - target(k)).abs();
            k += 1;
        }
    }
    err / steps as f64
}

Two ways to emit actions

A model has to turn its internal representation into numbers a motor can use. There are two families of answer.

Tokens. Discretise each action dimension into bins and let a language-model backbone predict the bin index, as it predicts words. RT-2 discretises each continuous dimension uniformly into 256 bins (Brohan et al. 2023, arXiv 2307.15818); OpenVLA does the same (Kim et al. 2024, arXiv 2406.09246). With NN bins over a range [lo,hi][lo, hi], a value is replaced by its bin's centre, so the rounding error is at most half a bin:

∣a−decode(encode(a))∣≤hi−lo2N\left| a - \mathrm{decode}(\mathrm{encode}(a)) \right| \le \frac{hi - lo}{2N}

Tokens fit the backbone's native training objective and make the model easy to fine-tune, at the price of quantisation and of predicting many tokens one by one. FAST compresses an action chunk with the discrete cosine transform before tokenising, so the model emits far fewer tokens (Pertsch et al. 2025, arXiv 2501.09747). In 2026 the tokenizer itself became a research subject: ActionCodec derives design principles from what helps the model learn rather than only reconstruction error (Dong et al. 2026, arXiv 2602.15397), and ActionPiece argues that a good tokenizer must preserve which actions are physically close to which, not only reconstruct each one (Lian et al. 2026, arXiv 2609.18487). A survey organises vision-language-action models entirely by the kind of action token they use (Zhong et al. 2025, arXiv 2507.01925).

Denoising. Start from noise and refine it into an action chunk over KK steps. Diffusion Policy introduced this for robots, representing the policy as a conditional denoising process (Chi et al. 2023, arXiv 2303.04137); π0 uses flow matching, a close relative, in a separate action expert attached to the VLM (Black et al. 2024, arXiv 2410.24164). Operationally: noise goes in, an action chunk comes out, in KK steps. Flow matching trains a network vv to predict, at each point xτx_\tau on a straight path from noise ϵ\epsilon to a clean action aa,

xτ=(1−τ) ϵ+τ a,v(xτ,τ)≈a−ϵx_\tau = (1 - \tau)\,\epsilon + \tau\, a, \qquad v(x_\tau, \tau) \approx a - \epsilon

and sampling follows the predicted direction in KK small steps, x←x+1K v(x,τ)x \leftarrow x + \tfrac{1}{K}\, v(x, \tau). The survey writes the same interpolation as its equation 9 (Lu et al. 2026, arXiv 2609.16074).

Why bother? Because real tasks often have more than one right answer. To get around an obstacle you can go left or right; going straight is the one wrong choice. A model trained to minimise squared error predicts the average of the demonstrations, which is straight into the obstacle. Denoising with enough steps keeps the two answers apart. Figure 4.3 makes the point with the exact flow for a two-way choice.

Tokens: round to the nearest bin
Denoising: flow from noise in K steps
Figure 4.3Tokens trade precision for compatibility with language models; denoising keeps several valid actions apart, if given enough steps. Left: a smooth action trajectory (dashed) and its tokenised version (solid); change the number of bins. Right: 600 actions sampled from noise by the exact flow to a two-way choice at -1 and +1; change the number of denoising steps K and watch one step collapse onto the average.toy simulation Source: rust/ch04-toolkit (Binning, TwoModes); the velocity field is computed in closed form, not learned.

With one step every sample lands near zero: the average, the action that hits the obstacle. With a handful of steps the samples split toward the two modes, and with many they settle on them. This is why few-step generation, which chapter 15 needs for speed, must be engineered carefully rather than simply turned down.

rust/ch04-toolkit/src/lib.rsThe exact velocity field for a two-way choice
/// A one-dimensional action distribution with two modes: steer left (-1) or right (+1) around an
/// obstacle, never straight ahead. Averaging the two is the one thing a policy must not do.
#[derive(Clone, Copy, Debug)]
pub struct TwoModes {
    pub mu: f64,
    pub sigma: f64,
}

impl TwoModes {
    /// Exact velocity of the straight-line (rectified) flow from N(0, 1) noise to this mixture:
    /// the expected value of (action - noise) given the current point x at time t. A flow-matching
    /// action head learns to approximate exactly this field; here we can write it down.
    pub fn velocity(&self, x: f64, t: f64) -> f64 {
        let (mut num, mut den) = (0.0, 0.0);
        for mu in [-self.mu, self.mu] {
            let var = (1.0 - t).powi(2) + t * t * self.sigma * self.sigma;
            let w = (-(x - t * mu).powi(2) / (2.0 * var)).exp() / var.sqrt();
            let v = mu + (t * self.sigma * self.sigma - (1.0 - t)) / var * (x - t * mu);
            num += w * v;
            den += w;
        }
        num / den
    }
}

A world model

A world model predicts how the world changes when the robot acts:

z^t+1=f(zt,at)\hat z_{t+1} = f(z_t, a_t)

Read it as: given a summary of the world now and an action, predict the summary after the action. Explicit world models decode the prediction back to pixels, so you can watch what they imagine; implicit ones keep it in latent space and are trained to be useful rather than to look right (Lu et al. 2026, arXiv 2609.16074). PlaNet and Dreamer are the classic latent world models of chapter 3 (Hafner et al. 2018, arXiv 1811.04551; Hafner et al. 2019, arXiv 1912.01603); the video generators of chapter 6 are explicit ones at enormous scale; the JEPA models of chapter 7 are implicit ones at scale.

A latent action

Most video has no action labels. A latent action recovers an action-like signal from the video itself: infer a code utu_t from two consecutive frames, then train a forward model to predict the second frame from the first and the code (Ye et al. 2024, arXiv 2410.11758):

ut∼q(ut∣zt,zt+1),z^t+1=F(zt,ut)u_t \sim q(u_t \mid z_t, z_{t+1}), \qquad \hat z_{t+1} = F(z_t, u_t)

Read it as: whatever must have happened between these two frames is, by definition, the latent action. The code is only useful if it is small, so that it cannot simply copy the next frame. Trained on human video, it gives a robot-independent description of motion that a small robot-specific decoder can later map to motor commands. That is how chapter 7's models use the internet's largest data layer.

Four components, two paradigms, two architectures

The 2026 surveys compare architectures with four conditional distributions, and every family in Part II is described as which of the four it trains and which it runs at inference (Lu et al. 2026, arXiv 2609.16074; Zhu et al. 2025, arXiv 2504.02792):

policy: p(at:t+H∣o<t,ℓ)forward dynamics: p(ot:t+H∣o<t,at:t+H)inverse dynamics: p(at:t+H∣o<t,ot:t+H)visual planning: p(ot:t+H∣o<t,ℓ)\begin{aligned} \text{policy:}\ & p(a_{t:t+H} \mid o_{<t}, \ell) & \text{forward dynamics:}\ & p(o_{t:t+H} \mid o_{<t}, a_{t:t+H}) \\ \text{inverse dynamics:}\ & p(a_{t:t+H} \mid o_{<t}, o_{t:t+H}) & \text{visual planning:}\ & p(o_{t:t+H} \mid o_{<t}, \ell) \end{aligned}
Past observationsand instructiono<t, ℓFuture observationso t:t+HActionsa t:t+HPolicyp(a | o<t, ℓ)Visual planningp(o | o<t, ℓ)Forward dynamicsp(o | o<t, a)Inverse dynamicsp(a | o<t, o)

Plan, then act (LingBot-VA, mimic-video, VPP). First imagine the future, then infer the actions that produce it. The action head solves the easier inverse problem. At inference the steps run in the order shown.

Figure 4.4Four arrows between past, future and action describe every family: which ones a model learns, and which it still runs at inference. The four components as arrows. Choose a style and a phase; inference steps light up in the order they run. Part II chapters each fill in this template for their family. Source: after Lu et al. 2026 (arXiv 2609.16074), equations 10 to 15.

Two transition paradigms combine them. Joint prediction generates observations and actions together; inverse dynamics imagines the future first and then infers the actions that produce it. Two architecture patterns carry them. End-to-end models use one backbone for everything; dual systems separate a world module from an action module and connect them by cross-attention or by mixture-of-transformers routing, where each modality keeps its own weights but they share attention. Chapter 6 fills in the resulting two-by-two with eighteen models.

Exercises

The stubs are in rust/ch04-toolkit/src/exercises.rs.

Exercise 4.1

Tokens

Implement encode and decode for uniform binning, and check the half-bin error bound for 256 bins. Then use figure 4.3 to find how many bins it takes before the tokenised trajectory is visually indistinguishable from the smooth one.

Check your answer from the rust/ folder:

cargo test -p ch04-toolkit --test exercises ex4_1
Hint 1

Clamp first, then scale the value to [0, bins) and truncate; clamp the index to bins - 1 for the top edge.

Hint 2

The decoded value is the bin's centre, half a bin above its lower edge.

Exercise 4.2

Denoising

Implement sample: integrate the flow's velocity with K Euler steps. With 64 steps, noise slightly right of zero should end near the right-hand mode.

Check your answer from the rust/ folder:

cargo test -p ch04-toolkit --test exercises ex4_2
Hint

Time runs from 0 to 1 in K equal steps; use the time at the start of each step.

Exercise 4.3

The regression trap

Implement fraction_near_modes and measure it for one step and for 32. Then explain, in two sentences, why a policy trained with mean squared error drives into the obstacle and a denoising policy does not.

Check your answer from the rust/ folder:

cargo test -p ch04-toolkit --test exercises ex4_3

What we are not sure about

Notation is not universal. This chapter follows the WAM survey. The π-series and GR00T papers write the same objects with different symbols, and some papers call a chunk a "trajectory" or a "plan". Appendix B will give the crosswalk.

Why chunking works is still being argued. The 2026 study cited above overturns the textbook explanations for imitation learning, but it is one study, and the deployment benefit shown in figure 4.2 is a separate effect from the learning benefit it measures.

Tokens versus denoising is not settled. Better tokenizers keep closing the gap in precision, and hybrid models train with tokens and act with a denoising head. The two-way choice in figure 4.3 is the clearest case for denoising; on many real tasks the difference is smaller.

Further reading

References

17 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.

  1. Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.
  2. T. Z. Zhao, V. Kumar, S. Levine, C. Finn Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv 2304.13705 v1, 2023-04-23.
  3. K. Black, M. Y. Galliker, S. Levine Real-Time Execution of Action Chunking Flow Policies. arXiv 2506.07339 v2, 2025-06-09.
  4. F. Lazzati, K. Stachowicz, W. Chen and 3 others Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control?. arXiv 2608.02547 v1, 2026-08-03.
  5. A. Brohan, N. Brown, J. Carbajal and 51 others RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv 2307.15818 v1, 2023-07-28.
  6. M. J. Kim, K. Pertsch, S. Karamcheti and 15 others OpenVLA: An Open-Source Vision-Language-Action Model. arXiv 2406.09246 v3, 2024-06-13.
  7. K. Pertsch, K. Stachowicz, B. Ichter and 6 others FAST: Efficient Action Tokenization for Vision-Language-Action Models. arXiv 2501.09747 v1, 2025-01-16.
  8. Z. Dong, Y. Liu, S. Zhang and 8 others ActionCodec: What Makes for Good Action Tokenizers. arXiv 2602.15397 v1, 2026-02-17.
  9. S. Lian, B. Yu, Z. Shen and 5 others ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models. arXiv 2609.18487 v1, 2026-09-16.
  10. Y. Zhong, F. Bai, S. Cai and 11 others A Survey on Vision-Language-Action Models: An Action Tokenization Perspective. arXiv 2507.01925 v1, 2025-07-02.
  11. C. Chi, Z. Xu, S. Feng and 5 others Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. arXiv 2303.04137 v5, 2023-03-07.
  12. K. Black, N. Brown, D. Driess and 21 others π0: A Vision-Language-Action Flow Model for General Robot Control. arXiv 2410.24164 v4, 2024-10-31.
  13. D. Hafner, T. Lillicrap, I. Fischer and 4 others Learning Latent Dynamics for Planning from Pixels. arXiv 1811.04551 v5, 2018-11-12.
  14. D. Hafner, T. Lillicrap, J. Ba, M. Norouzi Dream to Control: Learning Behaviors by Latent Imagination. arXiv 1912.01603 v3, 2019-12-03.
  15. S. Ye, J. Jang, B. Jeon and 13 others Latent Action Pretraining from Videos. arXiv 2410.11758 v2, 2024-10-15.
  16. C. Zhu, R. Yu, S. Feng and 3 others Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets. arXiv 2504.02792 v3, 2025-04-03.
  17. X. Zhang, X. Zeng, W. Zhang From World Models to World Action Models: A Concise Tutorial for Robotics. arXiv 2607.00836 v8, 2026-07-01.