Ch.07Part II · The seven familiesLatent prediction
Latent prediction and latent action models
Predict in representation space rather than pixel space, or infer latent actions so that action-free video becomes training signal.
- Math
- L2
- Sources
- 16
- Figures
- 5
- Exercises
- 3
- Last checked
- 3 Oct 2026
Equations with a walkthrough
7 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
- Pixel loss caused by the action error
- 61%average over 200 frames; this frame’s loss 0.067
- Feature loss caused by the action error
- 100%squared position error 1.00
Why this chapter exists
Chapter 6 showed what a robot gains by predicting the future, and what it pays: generating video is expensive, and most of what a video model predicts is irrelevant to the robot. The exact texture of a tablecloth, the flicker of a screen, a shadow moving across the counter all cost model capacity and compute, and none of them change what the gripper should do.
This chapter covers two related ideas that answer that cost. The first is to predict in representation space: forecast the features of the next frame, not its pixels, so the model is free to ignore whatever it cannot or need not predict. The second is to infer actions that were never recorded. A model watching two consecutive frames of a person stirring a pot can compress "what happened" into a small latent action, and that turns the largest layer of chapter 2's data pyramid, video with no actions at all, into something a policy can learn from.
After this chapter you will be able to:
- write down the difference between a pixel loss and a latent loss, explain what each term does, and use figure 7.1 to say why the latent one learns faster from noisy footage;
- explain with figure 7.2 how a latent action model reads an action off two frames, and why a human hand and a robot gripper can share one;
- judge where a latent model's claims are strong (cheap inference, learning from human video) and where they are still weak (per-robot decoding, collapse, real-robot comparisons).
To this end, we present DINO World Model (DINO-WM), a new method to model visual dynamics without reconstructing the visual world.
Predicting in the right space
Two losses
In chapter 4's notation, a model that predicts the next observation in pixels is trained to make its prediction match the real frame:
Every pixel counts equally. If the background flickers in a way no model could foresee, that flicker shows up in the loss on every training step, and it is the same size whether the model got the action right or wrong.
A joint-embedding predictive architecture, or JEPA, moves the comparison into feature space (Assran et al. 2025, arXiv 2506.09985; Balestriero et al. 2025, arXiv 2511.08544):
Read it term by term. is an encoder that turns the current frame into a feature vector. is a predictor that takes those features and the action and guesses the next frame's features. is a target encoder: it turns the real next frame into the features the guess is compared with. In most JEPAs it is a slowly updated copy of that receives no gradient. The loss never touches a pixel. Because the encoder is learned too, it is free to drop anything that cannot be predicted, which is the property JEPA-VLA's authors single out: predictive video embeddings are "adept at flexibly discarding unpredictable environment factors" (Miao et al. 2026, arXiv 2602.11832).
Figure 7.1 makes the difference concrete with an idealised encoder that keeps only where the block is. With no flicker, the whole pixel loss is about the action. At a flicker of 0.3, only about 61 percent of it is, and at 0.5 only about 36 percent: the rest is noise that pushes the model around on every step without teaching it anything. The feature loss stays entirely about the action. Exercise 7.2 reproduces the numbers.
The catch: collapse
The latent loss has an obvious cheat. If the encoder outputs the same vector for every frame, the predictor can match it perfectly and the loss is zero. This is representation collapse, and every JEPA needs a guard against it. The classic guards are the slowly updated target encoder and a stop-gradient. LeJEPA replaces these heuristics with a single regulariser, SIGReg, that pushes the embeddings towards an isotropic Gaussian distribution, which its authors prove is the best target for downstream prediction (Balestriero et al. 2025, arXiv 2511.08544). Exercise 7.1 builds a detector for the same failure in a latent action codebook.
What the family has shown
The lineage runs from general video features to robot control. DINO-WM predicts future DINOv2 patch features from offline trajectories and plans by optimising actions towards a goal image's features, solving tasks zero-shot in six environments (Zhou et al. 2024, arXiv 2411.04983). V-JEPA 2 was pretrained on over one million hours of internet video. Its action-conditioned variant, V-JEPA 2-AC, was post-trained on less than 62 hours of unlabelled robot video from DROID, then deployed zero-shot on Franka arms in two labs, picking and placing objects by planning towards image goals (Assran et al. 2025, arXiv 2506.09985). SkyJEPA took the idea to quadrotors, combining a JEPA dynamics model with sampling-based control on embedded hardware (Rao et al. 2026, arXiv 2606.23444).
In 2026 the idea reached VLAs and world-action models. VLA-JEPA pretrains with "leakage-free state prediction": future frames only produce targets, never inputs, so the model cannot learn to copy the answer (Sun et al. 2026, arXiv 2602.10098). JEPA-VLA adds V-JEPA 2 embeddings to existing VLAs and reports gains on LIBERO, LIBERO-Plus, RoboTwin 2.0 and real robots (Miao et al. 2026, arXiv 2602.11832). LaWAM predicts latent visual subgoals instead of future video. Its authors report 98.6 percent success on LIBERO, 187 ms per action chunk, and up to 24 times lower latency than pixel-space world-action models (Chen et al. 2026, arXiv 2606.15768). Two different August 2026 papers are both called JEPA-WAM. One couples latent transition prediction with action generation in a pretrained V-JEPA space; it reports 79.2 percent on LIBERO-Plus without large-scale robot pretraining and 86.3 percent when built on π0.5 (Lin et al. 2026, arXiv 2608.09381). The other encodes generated goal images with a frozen V-JEPA 2.1 encoder to improve instruction following (Liu et al. 2026, arXiv 2609.20277). Cite them by arXiv number.
Inferring actions that were never recorded
A latent action
Video of people doing things is abundant; robot actions are not. A latent action model bridges the two by learning, from video alone, a small code for "what changed between these two frames". It has two parts trained together: an inverse model that looks at both frames and outputs a code, and a forward model that must reconstruct the second frame from the first plus that code:
The trick is the bottleneck . In LAPA it is a VQ-VAE codebook: the code is snapped to the nearest of a small set of learned vectors (Ye et al. 2024, arXiv 2410.11758). Because is tiny and the forward model already sees , the cheapest thing for to carry is the change itself, the motion, not the appearance. Genie used this to make an 11-billion-parameter model of interactive environments from unlabelled internet video (Bruce et al. 2024, arXiv 2402.15391); robotics borrowed it to make video look like demonstration data.
The toy idealises the hard part: here the latent action model is told where the hand is. A real one must find the hand, ignore the camera shake and not smuggle appearance into the code. VLA-JEPA's authors argue that many latent-action objectives fail exactly here, staying "anchored to pixel variation rather than action-relevant state transitions" and vulnerable to "appearance bias, nuisance motion, and information leakage" (Sun et al. 2026, arXiv 2602.10098). ViPRA adds perceptual losses and optical-flow consistency to keep its latent actions physically grounded (Routray et al. 2025, arXiv 2511.07732). DreamDojo uses continuous rather than discrete latent actions as unified proxy actions across 44,000 hours of egocentric human video (Gao et al. 2026, arXiv 2602.06949).
From human video to one robot
Decode for one robot
A latent action means 'move right' in the camera's view, not 'turn joint 3 by 0.1 rad'. A small decoder, trained per robot, learns that mapping.
LAPA fine-tunes on small-scale robot data; ViPRA's chunked flow-matching decoder needs 100 to 200 teleoperated demonstrations and controls at up to 22 Hz.
Toy model of stage 3: fitting a two-by-two linear decoder from labelled pairs. Error of the fitted decoder, averaged over 40 random draws, against the number of labelled pairs.
A latent action says "the hand moved right" in the camera's view. A robot needs "joint 3 by 0.1 radians". The last stage learns that mapping from a small set of real robot demonstrations, labelled with both:
That is a far smaller problem than learning control from scratch, and the published numbers reflect it. ViPRA's decoder needs 100 to 200 teleoperated demonstrations and controls at up to 22 Hz; its authors report a 16 percent gain on the SIMPLER benchmark and 13 percent on real-world tasks over strong baselines (Routray et al. 2025, arXiv 2511.07732). LAPA reports outperforming a state-of-the-art VLA trained with action labels on real-world language-conditioned tasks, and positive transfer when pretraining on human manipulation video only (Ye et al. 2024, arXiv 2410.11758). Exercise 7.3 fits the toy decoder.
Latent tokens for a whole body
The same idea works one level down. GR00T N1.7 can drive a Unitree G1 humanoid through SONIC: the VLA "predicts compact latent action tokens that a learned whole-body controller decodes into full-body joint commands", including legs, arms and hands (NVIDIA 2026). SONIC itself is a motion-tracking foundation model scaled from 1.2 to 42 million parameters on more than 100 million frames from 700 hours of motion capture; a unified token space lets the same policy take commands from VR teleoperation or from a VLA (Luo et al. 2025, arXiv 2511.07820).
Here the latent space is not discovered from video but designed as an interface. It does the same job: it separates the part that needs language and vision from the part that needs physics, so each can be trained on the data that suits it.
The family on the four-component template
Latent-action policy (LAPA, ViPRA). Training first runs inverse dynamics on unlabelled video to read off latent actions, then teaches a policy to predict them. At inference only the policy runs, followed by a small per-robot decoder.
Why it exists, good at, bad at
Why it exists. Pixel prediction wastes capacity on appearance and can take shortcuts through it; latent targets force the model to learn the transitions that matter for control. Latent actions unlock video with no actions, the largest layer of the data pyramid.
What it is good at. Inference is cheaper than video generation, since nothing is rendered: LaWAM's reported 24-fold latency advantage over pixel-space WAMs is the clearest number (Chen et al. 2026, arXiv 2606.15768). Latent actions are the most direct bridge from human video to robot control, and they let one policy serve several bodies through per-robot decoders.
What it is bad at. Latent targets are hard to inspect. A video prediction can be watched; a feature vector cannot, so debugging a latent model means training probes. Collapse needs active prevention. And the last step, from latent action to real command, is still trained per embodiment on real robot data.
In code
The crate's strip world is small enough to read in one go: a block, a flickering background, a best-possible pixel predictor and an idealised encoder.
/// A one-dimensional camera image of `WIDTH` pixels: a bright block of `BLOCK` pixels that the
/// robot pushes, over a background that flickers unpredictably (lighting, shadows, a screen).
pub const WIDTH: usize = 48;
pub const BLOCK: f64 = 6.0;
/// Deterministic noise in [-1, 1] for pixel `i` of frame `seed`.
pub fn noise(seed: u64, i: usize) -> f64 {
let mut x = seed.wrapping_mul(0x9E37_79B9_7F4A_7C15) ^ (i as u64).wrapping_mul(0xBF58_476D_1CE4_E5B9);
x ^= x >> 31;
x = x.wrapping_mul(0x94D0_49BB_1331_11EB);
x ^= x >> 29;
(x % 20_001) as f64 / 10_000.0 - 1.0
}
/// Render the strip with the block's left edge at `pos` (sub-pixel positions blend edge pixels).
pub fn render(pos: f64, flicker: f64, seed: u64) -> Vec<f64> {
(0..WIDTH)
.map(|i| {
let (l, r) = (i as f64, i as f64 + 1.0);
let cover = (r.min(pos + BLOCK) - l.max(pos)).clamp(0.0, 1.0);
cover + (1.0 - cover) * flicker * noise(seed, i)
})
.collect()
}
/// The best a pixel predictor can do: move the block, and predict the background's average (zero),
/// because the flicker in the next frame cannot be known in advance.
pub fn predict_pixels(pos: f64, action: f64) -> Vec<f64> {
render(pos + action, 0.0, 0)
}
/// An idealised feature encoder: where the block is. A JEPA-style encoder trained on footage like
/// this learns something close to it, because the flicker cannot be predicted and so is not worth
/// encoding. Here it is a matched filter: the window of `BLOCK` pixels with the most brightness.
pub fn encode(frame: &[f64]) -> f64 {
let w = BLOCK as usize;
(0..=WIDTH - w)
.map(|s| (s, frame[s..s + w].iter().sum::<f64>()))
.fold((0, f64::MIN), |best, cur| if cur.1 > best.1 { cur } else { best })
.0 as f64
}
/// Mean squared error between two frames.
pub fn pixel_loss(pred: &[f64], actual: &[f64]) -> f64 {
pred.iter().zip(actual).map(|(p, a)| (p - a).powi(2)).sum::<f64>() / pred.len() as f64
}The latent action side is a codebook and a displacement. Everything hard about real latent action models is hidden in displacement, which here is given rather than learned.
/// A codebook of latent actions, as in a VQ-VAE latent action model: a handful of vectors, each
/// standing for one kind of change between two frames. Here: "stay", plus `k` directions.
#[derive(Clone, Debug)]
pub struct Codebook {
pub codes: Vec<[f64; 2]>,
}
impl Codebook {
pub fn ring(k: usize, step: f64) -> Codebook {
let mut codes = vec![[0.0, 0.0]];
codes.extend((0..k).map(|j| {
let a = std::f64::consts::TAU * j as f64 / k as f64;
[step * a.cos(), step * a.sin()]
}));
Codebook { codes }
}
}
/// The idealised latent-action encoder: how far the tracked hand or gripper moved between two
/// frames. A learned latent action model has to discover this from pixels alone.
pub fn displacement(before: [f64; 2], after: [f64; 2]) -> [f64; 2] {
[after[0] - before[0], after[1] - before[1]]
}Exercises
The stubs are in rust/ch07-latent/src/exercises.rs.
Exercise 7.1
Quantize, and catch a collapse
Implement quantize and usage. Usage is a cheap health check for a latent action model: if most transitions land on one code, the codebook has collapsed. What usage would you expect from a healthy model trained on varied manipulation video, and what would you check first if it fell to one code?
Check your answer from the rust/ folder:
cargo test -p ch07-latent --test exercises ex7_1Hint 1
Nearest code by squared Euclidean distance; ties go to the lower index.
Hint 2
Mark each code that receives at least one vector, then divide by the codebook size.
Exercise 7.2
What the pixel loss hears
Implement signal_share: the fraction of a pixel loss caused by a wrong action. Average it over 40 seeds at several flicker levels and compare with figure 7.1. Then write two sentences on what this means for the gradient a pixel-predicting model receives on noisy footage.
Check your answer from the rust/ folder:
cargo test -p ch07-latent --test exercises ex7_2Hint 1
Render the actual next frame with flicker; the predictions never contain flicker.
Hint 2
Compute the loss with the right action and with the wrong one, then take the difference.
Exercise 7.3
Fit the per-robot decoder
Implement fit_decoder by least squares, returning None when the data cannot determine the map. The tests include demonstrations that move in only one direction. Explain why a real decoder trained on demonstrations of a single task can fail in the same way on a new task.
Check your answer from the rust/ folder:
cargo test -p ch07-latent --test exercises ex7_3Hint 1
Accumulate the sums of z zᵀ and a zᵀ, then M = (Σ a zᵀ)(Σ z zᵀ)⁻¹.
Hint 2
If the 2 × 2 matrix Σ z zᵀ is singular, the latent actions all lie on one line.
What we are not sure about
Whether latent prediction beats video prediction at matched compute. The strongest comparisons are against older baselines or on simulated benchmarks. Real-robot comparisons between a latent world-action model and a video-generating one, with the same data and compute, are still rare.
What latent actions actually encode. A learned latent action can carry camera motion, lighting or the identity of the person, not just the movement. Leakage-free objectives and flow-consistency losses are attempts to stop this, and there is no standard test that checks it.
How far the human-to-robot transfer goes. The large human-video results (DreamDojo's 44,000 hours, Dyna-2's million hours) come from the organisations that built the models, measured on their own robots. Whether the same curves hold for other bodies, especially non-humanoid ones, is open.
Whether the decoder stays small. The per-robot step is cheap in published work, at 100 to 200 demonstrations for ViPRA. It is not known whether that holds for dexterous hands or whole-body control, where the mapping from a latent action to joints is far less direct.
Further reading
- V-JEPA 2, for the clearest demonstration of planning in feature space on real robots (Assran et al. 2025, arXiv 2506.09985).
- DINO-WM, a compact and readable latent world model (Zhou et al. 2024, arXiv 2411.04983).
- LAPA, the paper that made latent actions a VLA pretraining recipe (Ye et al. 2024, arXiv 2410.11758).
- VLA-JEPA, for the argument against pixel-anchored latent actions (Sun et al. 2026, arXiv 2602.10098).
- LaWAM, for latent subgoals in a world-action model (Chen et al. 2026, arXiv 2606.15768).
- LeJEPA, for a principled answer to collapse (Balestriero et al. 2025, arXiv 2511.08544).
- SONIC, for latent tokens as an interface to a whole-body controller (Luo et al. 2025, arXiv 2511.07820).
References
16 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- M. Assran, A. Bardes, D. Fan and 26 others V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. arXiv 2506.09985 v1, 2025-06-11.
- R. Balestriero, Y. LeCun LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics. arXiv 2511.08544 v3, 2025-11-11.
- S. Miao, N. Feng, J. Wu and 4 others JEPA-VLA: Video Predictive Embedding is Needed for VLA Models. arXiv 2602.11832 v1, 2026-02-12.
- G. Zhou, H. Pan, Y. LeCun, L. Pinto DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning. arXiv 2411.04983 v2, 2024-11-07.
- P. Rao, W. Zhang, R. Balestriero and 2 others SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors. arXiv 2606.23444 v2, 2026-06-22.
- J. Sun, W. Zhang, Z. Qi and 6 others VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model. arXiv 2602.10098 v2, 2026-02-10.
- J. Chen, K. Wang, K. Chen and 9 others LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies. arXiv 2606.15768 v1, 2026-06-14.
- Y. Lin, J. He, S. Bao and 6 others JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling. arXiv 2608.09381 v1, 2026-08-10.
- T. Liu, J. Zhu, T. Su and 4 others JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations. arXiv 2609.20277 v1, 2026-08-04.
- S. Ye, J. Jang, B. Jeon and 13 others Latent Action Pretraining from Videos. arXiv 2410.11758 v2, 2024-10-15.
- J. Bruce, M. Dennis, A. Edwards and 22 others Genie: Generative Interactive Environments. arXiv 2402.15391 v1, 2024-02-23.
- S. Routray, H. Pan, U. Jain and 2 others ViPRA: Video Prediction for Robot Actions. arXiv 2511.07732 v2, 2025-11-11.
- S. Gao, W. Liang, K. Zheng and 27 others DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos. arXiv 2602.06949 v1, 2026-02-06.
- NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
- Z. Luo, Y. Yuan, T. Wang and 25 others SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control. Science Robotics 11 (117), eaed4592 (2026). arXiv 2511.07820 v4, 2025-11-11.
- Dyna Robotics Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models. dyna.co, 2026-08.