Skip to content

Ch.07Part II · The seven familiesLatent prediction

Latent prediction and latent action models

Predict in representation space rather than pixel space, or infer latent actions so that action-free video becomes training signal.

Math
L2

Equations with a walkthrough

Sources
16

7 papers from 2026

Figures
5

Most are interactive

Exercises
3

Rust, with tests

Last checked
3 Oct 2026

The field moves monthly

Papers cited, by half-year of first arXiv version

Pixel branchnext framepredictionwhat the loss seessquared error per pixelFeature branchblock position012243648target 18prediction 19The background never reaches this branch: the encoder keeps only where the block is.
Pixel loss caused by the action error
61%average over 200 frames; this frame’s loss 0.067
Feature loss caused by the action error
100%squared position error 1.00
Figure 7.1A pixel loss spends most of its signal on things no model can predict; a loss on features hears mainly what the action changed. A robot pushes a block along a strip of 48 pixels while the background flickers unpredictably. The same wrong prediction is scored in pixels (top) and in features (bottom). Raise the flicker and watch where the pixel loss comes from.toy simulation Source: this book's toy, in rust/ch07-latent; the argument follows VLA-JEPA (arXiv 2602.10098), JEPA-VLA (2602.11832) and DINO-WM (2411.04983).

Why this chapter exists

Chapter 6 showed what a robot gains by predicting the future, and what it pays: generating video is expensive, and most of what a video model predicts is irrelevant to the robot. The exact texture of a tablecloth, the flicker of a screen, a shadow moving across the counter all cost model capacity and compute, and none of them change what the gripper should do.

This chapter covers two related ideas that answer that cost. The first is to predict in representation space: forecast the features of the next frame, not its pixels, so the model is free to ignore whatever it cannot or need not predict. The second is to infer actions that were never recorded. A model watching two consecutive frames of a person stirring a pot can compress "what happened" into a small latent action, and that turns the largest layer of chapter 2's data pyramid, video with no actions at all, into something a policy can learn from.

After this chapter you will be able to:

  • write down the difference between a pixel loss and a latent loss, explain what each term does, and use figure 7.1 to say why the latent one learns faster from noisy footage;
  • explain with figure 7.2 how a latent action model reads an action off two frames, and why a human hand and a robot gripper can share one;
  • judge where a latent model's claims are strong (cheap inference, learning from human video) and where they are still weak (per-robot decoding, collapse, real-robot comparisons).

To this end, we present DINO World Model (DINO-WM), a new method to model visual dynamics without reconstructing the visual world.

Gaoyue Zhou, Hengkai Pan, Yann LeCun and Lerrel Pinto, authors of DINO-WM. DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning (arXiv 2411.04983), 7 November 2024.

Predicting in the right space

Two losses

In chapter 4's notation, a model that predicts the next observation in pixels is trained to make its prediction o^t+1\hat o_{t+1} match the real frame:

Lpixel=∥o^t+1−ot+1∥2\mathcal{L}_{\text{pixel}} = \big\lVert \hat o_{t+1} - o_{t+1} \big\rVert^2

Every pixel counts equally. If the background flickers in a way no model could foresee, that flicker shows up in the loss on every training step, and it is the same size whether the model got the action right or wrong.

A joint-embedding predictive architecture, or JEPA, moves the comparison into feature space (Assran et al. 2025, arXiv 2506.09985; Balestriero et al. 2025, arXiv 2511.08544):

Llatent=∥ gϕ(fθ(ot), at)−fˉ (ot+1) ∥2\mathcal{L}_{\text{latent}} = \big\lVert\, g_\phi\big(f_\theta(o_t),\ a_t\big) - \bar f\,(o_{t+1}) \,\big\rVert^2

Read it term by term. fθf_\theta is an encoder that turns the current frame into a feature vector. gϕg_\phi is a predictor that takes those features and the action ata_t and guesses the next frame's features. fˉ\bar f is a target encoder: it turns the real next frame into the features the guess is compared with. In most JEPAs it is a slowly updated copy of fθf_\theta that receives no gradient. The loss never touches a pixel. Because the encoder is learned too, it is free to drop anything that cannot be predicted, which is the property JEPA-VLA's authors single out: predictive video embeddings are "adept at flexibly discarding unpredictable environment factors" (Miao et al. 2026, arXiv 2602.11832).

Figure 7.1 makes the difference concrete with an idealised encoder that keeps only where the block is. With no flicker, the whole pixel loss is about the action. At a flicker of 0.3, only about 61 percent of it is, and at 0.5 only about 36 percent: the rest is noise that pushes the model around on every step without teaching it anything. The feature loss stays entirely about the action. Exercise 7.2 reproduces the numbers.

The catch: collapse

The latent loss has an obvious cheat. If the encoder outputs the same vector for every frame, the predictor can match it perfectly and the loss is zero. This is representation collapse, and every JEPA needs a guard against it. The classic guards are the slowly updated target encoder and a stop-gradient. LeJEPA replaces these heuristics with a single regulariser, SIGReg, that pushes the embeddings towards an isotropic Gaussian distribution, which its authors prove is the best target for downstream prediction (Balestriero et al. 2025, arXiv 2511.08544). Exercise 7.1 builds a detector for the same failure in a latent action codebook.

What the family has shown

The lineage runs from general video features to robot control. DINO-WM predicts future DINOv2 patch features from offline trajectories and plans by optimising actions towards a goal image's features, solving tasks zero-shot in six environments (Zhou et al. 2024, arXiv 2411.04983). V-JEPA 2 was pretrained on over one million hours of internet video. Its action-conditioned variant, V-JEPA 2-AC, was post-trained on less than 62 hours of unlabelled robot video from DROID, then deployed zero-shot on Franka arms in two labs, picking and placing objects by planning towards image goals (Assran et al. 2025, arXiv 2506.09985). SkyJEPA took the idea to quadrotors, combining a JEPA dynamics model with sampling-based control on embedded hardware (Rao et al. 2026, arXiv 2606.23444).

In 2026 the idea reached VLAs and world-action models. VLA-JEPA pretrains with "leakage-free state prediction": future frames only produce targets, never inputs, so the model cannot learn to copy the answer (Sun et al. 2026, arXiv 2602.10098). JEPA-VLA adds V-JEPA 2 embeddings to existing VLAs and reports gains on LIBERO, LIBERO-Plus, RoboTwin 2.0 and real robots (Miao et al. 2026, arXiv 2602.11832). LaWAM predicts latent visual subgoals instead of future video. Its authors report 98.6 percent success on LIBERO, 187 ms per action chunk, and up to 24 times lower latency than pixel-space world-action models (Chen et al. 2026, arXiv 2606.15768). Two different August 2026 papers are both called JEPA-WAM. One couples latent transition prediction with action generation in a pretrained V-JEPA space; it reports 79.2 percent on LIBERO-Plus without large-scale robot pretraining and 86.3 percent when built on π0.5 (Lin et al. 2026, arXiv 2608.09381). The other encodes generated goal images with a frozen V-JEPA 2.1 encoder to improve instruction following (Liu et al. 2026, arXiv 2609.20277). Cite them by arXiv number.

Inferring actions that were never recorded

A latent action

Video of people doing things is abundant; robot actions are not. A latent action model bridges the two by learning, from video alone, a small code for "what changed between these two frames". It has two parts trained together: an inverse model EE that looks at both frames and outputs a code, and a forward model DD that must reconstruct the second frame from the first plus that code:

zt=q(E(ot, ot+1)),LLAM=∥D(ot, zt)−ot+1∥2z_t = q\big(E(o_t,\ o_{t+1})\big), \qquad \mathcal{L}_{\text{LAM}} = \big\lVert D(o_t,\ z_t) - o_{t+1} \big\rVert^2

The trick is the bottleneck qq. In LAPA it is a VQ-VAE codebook: the code is snapped to the nearest of a small set of learned vectors (Ye et al. 2024, arXiv 2410.11758). Because ztz_t is tiny and the forward model already sees oto_t, the cheapest thing for ztz_t to carry is the change itself, the motion, not the appearance. Genie used this to make an 11-billion-parameter model of interactive environments from unlabelled internet video (Bruce et al. 2024, arXiv 2402.15391); robotics borrowed it to make video look like demonstration data.

frame 0frame 1latent action modelwhat changed between frames?codebook: 9 latent actionsforward modelLatent actions read from each videohuman videorobot videoSame task, different bodies:the same sequence of latent actions.dashed: predicted fromframe 0 and thelatent action
Figure 7.2A latent action is the model's own word for 'what changed'; if two bodies do the same thing, they get the same words. A hand and a gripper each slide a cup right, lift it, pause and bring it back. The latent action model reads one code off each pair of frames; the forward model uses the code to predict the next frame. Switch videos and compare the two rows of codes.toy simulation Source: this book's toy, after the VQ-VAE latent actions of LAPA (arXiv 2410.11758) and Genie (2402.15391); codebook and quantizer in rust/ch07-latent.

The toy idealises the hard part: here the latent action model is told where the hand is. A real one must find the hand, ignore the camera shake and not smuggle appearance into the code. VLA-JEPA's authors argue that many latent-action objectives fail exactly here, staying "anchored to pixel variation rather than action-relevant state transitions" and vulnerable to "appearance bias, nuisance motion, and information leakage" (Sun et al. 2026, arXiv 2602.10098). ViPRA adds perceptual losses and optical-flow consistency to keep its latent actions physically grounded (Routray et al. 2025, arXiv 2511.07732). DreamDojo uses continuous rather than discrete latent actions as unified proxy actions across 44,000 hours of egocentric human video (Gao et al. 2026, arXiv 2602.06949).

From human video to one robot

1. Learn latent actionsselect for detailsconsumesinternet and human video, no actions2. Pretrain on latent actionsselect for detailsconsumesthe same video, now labelled with latent actions3. Decode for one robotselect for detailsconsumesa few hundred action-labelled robot demonstrations

Decode for one robot

A latent action means 'move right' in the camera's view, not 'turn joint 3 by 0.1 rad'. A small decoder, trained per robot, learns that mapping.

LAPA fine-tunes on small-scale robot data; ViPRA's chunked flow-matching decoder needs 100 to 200 teleoperated demonstrations and controls at up to 22 Hz.

Toy model of stage 3: fitting a two-by-two linear decoder from labelled pairs. Error of the fitted decoder, averaged over 40 random draws, against the number of labelled pairs.

Figure 7.3Latent actions move the expensive, robot-specific step to the end of the pipeline, where it needs hundreds of demonstrations instead of millions. Three stages, each consuming a different layer of the data pyramid. Select a stage for what it does and which models do it. Below, a toy of the last stage: how many labelled pairs a small decoder needs, at three levels of label noise.toy simulation Source: stages after LAPA (arXiv 2410.11758) and ViPRA (2511.07732); decoder fit in rust/ch07-latent.

A latent action says "the hand moved right" in the camera's view. A robot needs "joint 3 by 0.1 radians". The last stage learns that mapping from a small set of real robot demonstrations, labelled with both:

a^t=hψ(zt),ψ fitted on a few hundred (zt, at) pairs\hat a_t = h_\psi(z_t), \qquad \psi \text{ fitted on a few hundred } (z_t,\ a_t) \text{ pairs}

That is a far smaller problem than learning control from scratch, and the published numbers reflect it. ViPRA's decoder needs 100 to 200 teleoperated demonstrations and controls at up to 22 Hz; its authors report a 16 percent gain on the SIMPLER benchmark and 13 percent on real-world tasks over strong baselines (Routray et al. 2025, arXiv 2511.07732). LAPA reports outperforming a state-of-the-art VLA trained with action labels on real-world language-conditioned tasks, and positive transfer when pretraining on human manipulation video only (Ye et al. 2024, arXiv 2410.11758). Exercise 7.3 fits the toy decoder.

Latent tokens for a whole body

The same idea works one level down. GR00T N1.7 can drive a Unitree G1 humanoid through SONIC: the VLA "predicts compact latent action tokens that a learned whole-body controller decodes into full-body joint commands", including legs, arms and hands (NVIDIA 2026). SONIC itself is a motion-tracking foundation model scaled from 1.2 to 42 million parameters on more than 100 million frames from 700 hours of motion capture; a unified token space lets the same policy take commands from VR teleoperation or from a VLA (Luo et al. 2025, arXiv 2511.07820).

images, instructionVLAGR00T N1.7compact latent action tokenswhole-bodycontrollerSONIC: learnedmotion trackingjoint commands: legsThe VLA does not need to know how to balance or place a foot. It says what the body should do ina few tokens; the controller, trained on motion capture, knows how a body does it.
Figure 7.4The VLA chooses what the body should do in a few tokens; a controller trained on human motion decides how. The SONIC path in GR00T N1.7: latent action tokens out of the VLA, joint commands for legs, arms and hands out of the whole-body controller.schematic, not measured Source: NVIDIA Isaac-GR00T repository (GR00T N1.7 README) and Luo et al. 2025 (arXiv 2511.07820).

Here the latent space is not discovered from video but designed as an interface. It does the same job: it separates the part that needs language and vision from the part that needs physics, so each can be trained on the data that suits it.

The family on the four-component template

Past observationsand instructiono<t, ℓFuture featuresz t:t+HActionsa t:t+HPolicyp(a | o<t, ℓ)Visual planningp(o | o<t, ℓ)Forward dynamicsp(o | o<t, a)Inverse dynamicsp(a | o<t, o)

Latent-action policy (LAPA, ViPRA). Training first runs inverse dynamics on unlabelled video to read off latent actions, then teaches a policy to predict them. At inference only the policy runs, followed by a small per-robot decoder.

Figure 7.5Latent models use the same four components as world-action models, but predict features instead of pixels, which changes what each component costs. Chapter 4's template filled in for four styles of latent model. 'Future features' replaces future observations: nothing is rendered. Select a style; switch between training and inference. Source: placements from each paper's abstract: LAPA (arXiv 2410.11758), ViPRA (2511.07732), V-JEPA 2 (2506.09985), DINO-WM (2411.04983), VLA-JEPA (2602.10098), JEPA-WAM (2608.09381), LaWAM (2606.15768).

Why it exists, good at, bad at

Why it exists. Pixel prediction wastes capacity on appearance and can take shortcuts through it; latent targets force the model to learn the transitions that matter for control. Latent actions unlock video with no actions, the largest layer of the data pyramid.

What it is good at. Inference is cheaper than video generation, since nothing is rendered: LaWAM's reported 24-fold latency advantage over pixel-space WAMs is the clearest number (Chen et al. 2026, arXiv 2606.15768). Latent actions are the most direct bridge from human video to robot control, and they let one policy serve several bodies through per-robot decoders.

What it is bad at. Latent targets are hard to inspect. A video prediction can be watched; a feature vector cannot, so debugging a latent model means training probes. Collapse needs active prevention. And the last step, from latent action to real command, is still trained per embodiment on real robot data.

In code

The crate's strip world is small enough to read in one go: a block, a flickering background, a best-possible pixel predictor and an idealised encoder.

rust/ch07-latent/src/lib.rsThe strip world behind figure 7.1
/// A one-dimensional camera image of `WIDTH` pixels: a bright block of `BLOCK` pixels that the
/// robot pushes, over a background that flickers unpredictably (lighting, shadows, a screen).
pub const WIDTH: usize = 48;
pub const BLOCK: f64 = 6.0;

/// Deterministic noise in [-1, 1] for pixel `i` of frame `seed`.
pub fn noise(seed: u64, i: usize) -> f64 {
    let mut x = seed.wrapping_mul(0x9E37_79B9_7F4A_7C15) ^ (i as u64).wrapping_mul(0xBF58_476D_1CE4_E5B9);
    x ^= x >> 31;
    x = x.wrapping_mul(0x94D0_49BB_1331_11EB);
    x ^= x >> 29;
    (x % 20_001) as f64 / 10_000.0 - 1.0
}

/// Render the strip with the block's left edge at `pos` (sub-pixel positions blend edge pixels).
pub fn render(pos: f64, flicker: f64, seed: u64) -> Vec<f64> {
    (0..WIDTH)
        .map(|i| {
            let (l, r) = (i as f64, i as f64 + 1.0);
            let cover = (r.min(pos + BLOCK) - l.max(pos)).clamp(0.0, 1.0);
            cover + (1.0 - cover) * flicker * noise(seed, i)
        })
        .collect()
}

/// The best a pixel predictor can do: move the block, and predict the background's average (zero),
/// because the flicker in the next frame cannot be known in advance.
pub fn predict_pixels(pos: f64, action: f64) -> Vec<f64> {
    render(pos + action, 0.0, 0)
}

/// An idealised feature encoder: where the block is. A JEPA-style encoder trained on footage like
/// this learns something close to it, because the flicker cannot be predicted and so is not worth
/// encoding. Here it is a matched filter: the window of `BLOCK` pixels with the most brightness.
pub fn encode(frame: &[f64]) -> f64 {
    let w = BLOCK as usize;
    (0..=WIDTH - w)
        .map(|s| (s, frame[s..s + w].iter().sum::<f64>()))
        .fold((0, f64::MIN), |best, cur| if cur.1 > best.1 { cur } else { best })
        .0 as f64
}

/// Mean squared error between two frames.
pub fn pixel_loss(pred: &[f64], actual: &[f64]) -> f64 {
    pred.iter().zip(actual).map(|(p, a)| (p - a).powi(2)).sum::<f64>() / pred.len() as f64
}

The latent action side is a codebook and a displacement. Everything hard about real latent action models is hidden in displacement, which here is given rather than learned.

rust/ch07-latent/src/lib.rsA ring codebook and the idealised latent action
/// A codebook of latent actions, as in a VQ-VAE latent action model: a handful of vectors, each
/// standing for one kind of change between two frames. Here: "stay", plus `k` directions.
#[derive(Clone, Debug)]
pub struct Codebook {
    pub codes: Vec<[f64; 2]>,
}

impl Codebook {
    pub fn ring(k: usize, step: f64) -> Codebook {
        let mut codes = vec![[0.0, 0.0]];
        codes.extend((0..k).map(|j| {
            let a = std::f64::consts::TAU * j as f64 / k as f64;
            [step * a.cos(), step * a.sin()]
        }));
        Codebook { codes }
    }
}

/// The idealised latent-action encoder: how far the tracked hand or gripper moved between two
/// frames. A learned latent action model has to discover this from pixels alone.
pub fn displacement(before: [f64; 2], after: [f64; 2]) -> [f64; 2] {
    [after[0] - before[0], after[1] - before[1]]
}

Exercises

The stubs are in rust/ch07-latent/src/exercises.rs.

Exercise 7.1

Quantize, and catch a collapse

Implement quantize and usage. Usage is a cheap health check for a latent action model: if most transitions land on one code, the codebook has collapsed. What usage would you expect from a healthy model trained on varied manipulation video, and what would you check first if it fell to one code?

Check your answer from the rust/ folder:

cargo test -p ch07-latent --test exercises ex7_1
Hint 1

Nearest code by squared Euclidean distance; ties go to the lower index.

Hint 2

Mark each code that receives at least one vector, then divide by the codebook size.

Exercise 7.2

What the pixel loss hears

Implement signal_share: the fraction of a pixel loss caused by a wrong action. Average it over 40 seeds at several flicker levels and compare with figure 7.1. Then write two sentences on what this means for the gradient a pixel-predicting model receives on noisy footage.

Check your answer from the rust/ folder:

cargo test -p ch07-latent --test exercises ex7_2
Hint 1

Render the actual next frame with flicker; the predictions never contain flicker.

Hint 2

Compute the loss with the right action and with the wrong one, then take the difference.

Exercise 7.3

Fit the per-robot decoder

Implement fit_decoder by least squares, returning None when the data cannot determine the map. The tests include demonstrations that move in only one direction. Explain why a real decoder trained on demonstrations of a single task can fail in the same way on a new task.

Check your answer from the rust/ folder:

cargo test -p ch07-latent --test exercises ex7_3
Hint 1

Accumulate the sums of z zᵀ and a zᵀ, then M = (Σ a zᵀ)(Σ z zᵀ)⁻¹.

Hint 2

If the 2 × 2 matrix Σ z zᵀ is singular, the latent actions all lie on one line.

What we are not sure about

Whether latent prediction beats video prediction at matched compute. The strongest comparisons are against older baselines or on simulated benchmarks. Real-robot comparisons between a latent world-action model and a video-generating one, with the same data and compute, are still rare.

What latent actions actually encode. A learned latent action can carry camera motion, lighting or the identity of the person, not just the movement. Leakage-free objectives and flow-consistency losses are attempts to stop this, and there is no standard test that checks it.

How far the human-to-robot transfer goes. The large human-video results (DreamDojo's 44,000 hours, Dyna-2's million hours) come from the organisations that built the models, measured on their own robots. Whether the same curves hold for other bodies, especially non-humanoid ones, is open.

Whether the decoder stays small. The per-robot step is cheap in published work, at 100 to 200 demonstrations for ViPRA. It is not known whether that holds for dexterous hands or whole-body control, where the mapping from a latent action to joints is far less direct.

Further reading

References

16 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.

  1. M. Assran, A. Bardes, D. Fan and 26 others V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. arXiv 2506.09985 v1, 2025-06-11.
  2. R. Balestriero, Y. LeCun LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics. arXiv 2511.08544 v3, 2025-11-11.
  3. S. Miao, N. Feng, J. Wu and 4 others JEPA-VLA: Video Predictive Embedding is Needed for VLA Models. arXiv 2602.11832 v1, 2026-02-12.
  4. G. Zhou, H. Pan, Y. LeCun, L. Pinto DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning. arXiv 2411.04983 v2, 2024-11-07.
  5. P. Rao, W. Zhang, R. Balestriero and 2 others SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors. arXiv 2606.23444 v2, 2026-06-22.
  6. J. Sun, W. Zhang, Z. Qi and 6 others VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model. arXiv 2602.10098 v2, 2026-02-10.
  7. J. Chen, K. Wang, K. Chen and 9 others LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies. arXiv 2606.15768 v1, 2026-06-14.
  8. Y. Lin, J. He, S. Bao and 6 others JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling. arXiv 2608.09381 v1, 2026-08-10.
  9. T. Liu, J. Zhu, T. Su and 4 others JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations. arXiv 2609.20277 v1, 2026-08-04.
  10. S. Ye, J. Jang, B. Jeon and 13 others Latent Action Pretraining from Videos. arXiv 2410.11758 v2, 2024-10-15.
  11. J. Bruce, M. Dennis, A. Edwards and 22 others Genie: Generative Interactive Environments. arXiv 2402.15391 v1, 2024-02-23.
  12. S. Routray, H. Pan, U. Jain and 2 others ViPRA: Video Prediction for Robot Actions. arXiv 2511.07732 v2, 2025-11-11.
  13. S. Gao, W. Liang, K. Zheng and 27 others DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos. arXiv 2602.06949 v1, 2026-02-06.
  14. NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
  15. Z. Luo, Y. Yuan, T. Wang and 25 others SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control. Science Robotics 11 (117), eaed4592 (2026). arXiv 2511.07820 v4, 2025-11-11.
  16. Dyna Robotics Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models. dyna.co, 2026-08.