Ch.02Part I · Foundations
Why now: the convergence
Explain the timing as six ingredients that matured within roughly the same 24 months.
- Math
- L0
- Sources
- 26
- Figures
- 5
- Exercises
- 3
- Last checked
- 3 Oct 2026
No equations
6 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
State of Robotics 2026 2026-03
Teleoperation data at $118 an hour, down from about $340 in early 2024; fine-tuning on 200 to 500 demonstrations beats training from scratch on 1,000 or more (industry survey).
Why this chapter exists
Every idea in this book is older than the models that made it famous. Learning robot policies from demonstrations, predicting video, discretising actions into tokens, using a world model to plan: each has a paper from the 2010s or earlier. So why did robot foundation models arrive in 2024 to 2026 and not five years before?
The honest answer is that no one thing changed. Six ingredients matured within roughly the same 24 months: vision-language models good enough to be a backbone, video models good enough to be a dynamics prior, robot datasets large enough to pool, human video usable as robot data, inference fast enough for a robot, and economics that made fine-tuning cheaper than starting over. Each was necessary, none was sufficient. This chapter tells the story of their convergence, because that story predicts what the next convergence will need.
After this chapter you will be able to:
- name the six ingredients and say which family of model depends most on each;
- explain, with the simulator in figure 2.3, how action chunking turns a slow model into a fast controller;
- tell which of the commonly quoted numbers about this shift are measured, and which come from vendors or industry surveys.
To make progress toward general-purpose robots, researchers need platforms that are both capable and broadly accessible.
1. Vision-language models became good enough to be a backbone
The first generalist robot policies borrowed their understanding of the world from vision-language models (VLMs). The recipe, which chapter 5 explains in full, is to take a pretrained VLM and teach it to emit actions. It only works if the VLM already knows what a mug is, where the handle sits and what "put it in the sink" means.
By 2024 open VLMs crossed that bar. PaliGemma, a 3-billion-parameter model built for transfer, became the backbone of π0 and π0.5 (Beyer et al. 2024, arXiv 2407.07726; Physical Intelligence et al. 2025, arXiv 2504.16054). Qwen-VL and its successors became backbones for others (Bai et al. 2023, arXiv 2308.12966). In 2026 π0.7 moved to Gemma 3 (Gemma Team et al. 2025, arXiv 2503.19786; Physical Intelligence et al. 2026, arXiv 2604.15483), and NVIDIA's GR00T N1.7 to Cosmos-Reason2, a VLM trained for physical reasoning and released in December 2025 (NVIDIA 2026). A backbone is now a choice, not a research project.
2. Video models became good enough to be a dynamics prior
A video generator that has watched enough of the world has learned, implicitly, how things move: what happens when a cup is tipped, how cloth falls. Chapter 6 showed what that buys a robot policy. The ingredient arrived in early 2025, when large open video models appeared. NVIDIA's Cosmos world foundation models were trained on about 100 million clips drawn from roughly 20 million hours of raw video (NVIDIA et al. 2025, arXiv 2501.03575). Wan's 14-billion-parameter model was trained on billions of images and videos (Team Wan et al. 2025, arXiv 2503.20314). A Google DeepMind study then showed a frontier video model solving visual tasks it had never been trained for (Wiedemer et al. 2025, arXiv 2509.20328).
The important word is open. When the video backbone had to be pretrained from scratch, the approach was out of reach for most robotics labs; once open models existed, a lab could fine-tune one (NVIDIA Technical Blog 2026).
3. Robot datasets crossed the million-trajectory mark
Robots produce data slowly. A trajectory is one demonstration of one task: minutes of an operator's time, recorded on one robot in one room. For years every lab trained on its own few thousand.
Pooling changed the scale. Open X-Embodiment gathered over one million real robot trajectories from 22 robot types, contributed by 21 institutions, in standard formats (Open X-Embodiment Collaboration et al. 2023, arXiv 2310.08864). DROID collected 76,000 trajectories (350 hours) across 564 scenes in the wild (Khazatsky et al. 2024, arXiv 2403.12945), and BridgeData V2 60,096 trajectories on a low-cost robot (Walke et al. 2023, arXiv 2308.12952). A million trajectories is still tiny next to the web, which is why the next ingredient mattered.
Robot trajectories
| Example | Reported size | Source |
|---|---|---|
| Open X-Embodiment | 1M+ trajectories, 22 embodiments | O'Neill et al. 2023, arXiv 2310.08864 |
| DROID | 76k trajectories, 350 hours, 564 scenes | Khazatsky et al. 2024, arXiv 2403.12945 |
| BridgeData V2 | 60,096 trajectories, 24 environments | Walke et al. 2023, arXiv 2308.12952 |
Who can use it: every family for fine-tuning; reactive VLAs (chapter 5) and large behavior models (chapter 11) learn mostly from this layer.
4. Human video became usable as robot data
There is far more video of people doing things with their hands than of robots doing anything. The problem is that human video carries no robot actions. Two developments made it usable. Better capture added the missing signal: EgoDex recorded 829 hours of egocentric video with tracked 3D hand and finger poses (Hoque et al. 2025, arXiv 2505.11709), and the UMI hand-held gripper turned human demonstrations into robot-ready trajectories (Chi et al. 2024, arXiv 2402.10329). Better learning made the rest of it usable: latent actions, chapter 7's subject, infer an action-like signal from consecutive frames.
The scale then moved quickly. EgoScale trained on 20,854 hours of action-labelled egocentric video and reports a log-linear scaling law between hours of human data and prediction loss (Zheng et al. 2026, arXiv 2602.16710). DreamDojo pretrained a world model on 44,000 hours (Gao et al. 2026, arXiv 2602.06949). In August 2026 Dyna Robotics reported pretraining on more than one million hoursvendor claim (Dyna Robotics 2026). The direction is clear. Whether the scaling holds across labs and robots is not yet established, and chapter 13 looks at the evidence.
5. Inference became fast enough for a robot
A robot arm runs its controller tens to hundreds of times a second. A large model takes a tenth of a second or more to think. Two ideas closed the gap.
The first is the action chunk: predict the next several actions at once, then execute them while the model is idle. ACT introduced the idea for robot imitation in 2023 (Zhao et al. 2023, arXiv 2304.13705). One inference now pays for many motor commands. The second is thinking while moving: compute the next chunk while the current one executes, so the robot never stops. Real-time chunking made this work for flow-based VLAs without retraining (Black et al. 2025, arXiv 2506.07339).
The simulator below shows why both matter. Set a 200-millisecond model and a 50 Hz controller. With one action per inference the robot delivers under 5 commands a second and spends most of its time waiting. Raise the chunk length and the waiting is amortised; switch to thinking while moving and the robot runs at its controller's full rate as soon as a chunk lasts longer than an inference.
- Commands per second delivered
- 25.0controller runs at 50 Hz
- One chunk lasts
- 200 ms10 actions at 50 Hz
- Inference takes
- 200 msthe robot waits every time
Hardware and engineering did the rest. Quantised models, exported to inference engines and run on robot-mounted computers such as NVIDIA's Jetson Thor (NVIDIA Newsroom 2026), now reach control-relevant rates on the robot itself. GR00T N1.7 ships with ONNX and TensorRT export (NVIDIA Technical Blog 2026). An industry survey puts quantised VLAs at 10 to 25 Hz on consumer-grade GPUsindustry survey (SVRC (Robotics Center of Silicon Valley) 2026), and a 2026 study across edge accelerators finds that right-sized devices can meet control-rate constraints at lower cost and energy than flagship GPUs (Zhou et al. 2026, arXiv 2604.24447).
/// Synchronous execution: the robot stops while the model thinks, then runs the whole chunk.
/// Returns motor commands per second actually delivered.
pub fn sync_rate_hz(latency_ms: f32, chunk_len: usize, control_hz: f32) -> f32 {
let execute_ms = chunk_len as f32 * 1000.0 / control_hz;
chunk_len as f32 * 1000.0 / (latency_ms + execute_ms)
}
/// Asynchronous execution: the next chunk is computed while the current one runs. The robot never
/// stops as long as inference finishes before the steps it is executing run out.
pub fn async_keeps_up(latency_ms: f32, execute_steps: usize, control_hz: f32) -> bool {
latency_ms <= execute_steps as f32 * 1000.0 / control_hz
}6. The economics flipped
The last ingredient is money. For most of the 2010s a new robot task meant collecting thousands of demonstrations and training a policy from scratch. Two numbers from the State of Robotics 2026 report describe the change, and both should be read as industry-survey claims rather than measurements.
The first is the price of data. One hour of teleoperation data, captured, labelled and packaged, fell from about 136 by the end of 2025 and $118 by March 2026industry survey (SVRC (Robotics Center of Silicon Valley) 2026).
The second is how much data a new task needs. The same report finds that fine-tuning a pretrained VLA on 200 to 500 demonstrations now outperforms a task-specific policy trained from scratch on more than 1,000industry survey. Model makers report smaller numbers still, for their own models: LingBot-VLA used 130 post-training episodes per task in its evaluation (Wu et al. 2026, arXiv 2601.18692), and MotuBrain adapts to a new humanoid from 50 to 100 trajectoriesvendor claim (Motubrain Team et al. 2026, arXiv 2604.27792).
When fine-tuning a shared base beats building from scratch, the economically rational move is to build the base. That is the moment foundation models stop being a research bet and become a product strategy, which is how chapter 3's fourth phase begins.
What converged, and what it predicts
Laid side by side, the six tracks in figure 2.1 explain the timing better than any single milestone. The robot datasets of 2023 had nothing strong enough to learn from them until the VLM backbones of 2024. The video models of 2025 had no way to drive a robot until action chunking and fast inference made the output usable. Human video was a curiosity until the latent-action and hand-tracking work of 2025 and 2026 gave it a robot-shaped signal.
The pattern suggests what to watch. The next shift will need its own convergence: probably cheap, broad real-robot evaluation (chapter 15), safety cases for learned controllers (chapter 17), and memory for tasks longer than a few minutes (chapter 14). Each is still a single track.
The chapter's Rust crate also fits the one scaling curve in this chapter with a stated number. EgoScale reports average task completion rising from 0.30 after 1,000 hours of human pretraining to 0.71 after 20,000 (Zheng et al. 2026, arXiv 2602.16710). A log-linear curve through those two points is easy to draw, and exercise 2.3 shows how quickly it starts promising more than 100 percent task completion if you follow it beyond the data.
/// A log-linear scaling curve, y = a + b ln(x), the form EgoScale reports for human-video hours.
#[derive(Clone, Copy, Debug, PartialEq)]
pub struct LogLinear {
pub a: f64,
pub b: f64,
}
impl LogLinear {
pub fn predict(&self, x: f64) -> f64 {
self.a + self.b * x.ln()
}
}
/// EgoScale reports average task completion of 0.30 after 1k hours of egocentric pretraining
/// and 0.71 after 20k hours (Zheng et al. 2026, arXiv 2602.16710, section 5).
pub const EGOSCALE_POINTS: [(f64, f64); 2] = [(1_000.0, 0.30), (20_000.0, 0.71)];Exercises
The stubs are in rust/ch02-why-now/src/exercises.rs.
Exercise 2.1
How long must a chunk be?
Implement min_chunk_for_rate: the shortest chunk that, executed synchronously, delivers at least a target rate. Then use figure 2.3 to check your answer for a 200 ms model, a 50 Hz controller and a 25 Hz target.
Check your answer from the rust/ folder:
cargo test -p ch02-why-now --test exercises ex2_1Hint 1
Try chunk lengths in order and stop at the first that reaches the target rate.
Hint 2
No chunk length can deliver more commands per second than the controller itself runs.
Exercise 2.2
Fit the curve
Implement fit_log_linear, a least-squares fit of a straight line against the logarithm of the hours. With EgoScale's two points it must pass through both exactly; with four noisy points it must recover the slope.
Check your answer from the rust/ folder:
cargo test -p ch02-why-now --test exercises ex2_2Hint
Fit a straight line to (ln x, y) by ordinary least squares.
Exercise 2.3
Know when to stop believing it
Implement guarded_predict, which refuses to extrapolate beyond the data or to return a prediction outside its valid range. Then compute what the unguarded curve predicts at 400,000 hours, and write one sentence on why a vendor's scaling plot should always state the range it was measured over.
Check your answer from the rust/ folder:
cargo test -p ch02-why-now --test exercises ex2_3Hint
Check the range of the fitted x values first, then the range of the prediction.
What we are not sure about
Whether the scaling is real across labs. The clearest scaling evidence for human video comes from model makers measuring their own models on their own robots: EgoScale's log-linear curve, Dyna-2's million hours. Independent replication is not yet published as far as this book's sources show.
What the economic numbers measure. The data-cost and demonstrations figures come from one industry survey. It does not publish the underlying deployments, and "fine-tuning beats training from scratch" depends heavily on how close the new task is to the pretraining data.
Whether the order matters. This chapter tells the six tracks as converging; it is equally possible that one of them, cheap robot data or open video models, did most of the work and the others were already ready. The timeline cannot separate those readings.
Further reading
- Open X-Embodiment, the dataset that made pooling robot data normal (Open X-Embodiment Collaboration et al. 2023, arXiv 2310.08864).
- EgoScale, for the strongest public case for human video at scale (Zheng et al. 2026, arXiv 2602.16710).
- Real-time chunking, for how asynchronous execution works with flow-based policies (Black et al. 2025, arXiv 2506.07339).
- The NVIDIA post on world-action models, for why open video backbones changed what a lab can afford (NVIDIA Technical Blog 2026).
- The State of Robotics 2026 report, read as an industry survey (SVRC (Robotics Center of Silicon Valley) 2026).
References
26 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- L. Beyer, A. Steiner, A. S. Pinto and 32 others PaliGemma: A versatile 3B VLM for transfer. arXiv 2407.07726 v2, 2024-07-10.
- Physical Intelligence, K. Black, N. Brown and 33 others π0.5: a Vision-Language-Action Model with Open-World Generalization. arXiv 2504.16054 v1, 2025-04-22.
- J. Bai, S. Bai, S. Yang and 6 others Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv 2308.12966 v3, 2023-08-24.
- Gemma Team, A. Kamath, J. Ferret and 212 others Gemma 3 Technical Report. arXiv 2503.19786 v1, 2025-03-25.
- Physical Intelligence, B. Ai, A. Amin and 85 others π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities. arXiv 2604.15483 v2, 2026-04-16.
- NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
- NVIDIA, N. Agarwal, A. Ali and 75 others Cosmos World Foundation Model Platform for Physical AI. arXiv 2501.03575 v3, 2025-01-07.
- Team Wan, A. Wang, B. Ai and 59 others Wan: Open and Advanced Large-Scale Video Generative Models. arXiv 2503.20314 v2, 2025-03-26.
- T. Wiedemer, Y. Li, P. Vicol and 6 others Video models are zero-shot learners and reasoners. arXiv 2509.20328 v2, 2025-09-24.
- NVIDIA Technical Blog Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models. developer.nvidia.com, 2026-06-15.
- Open X-Embodiment Collaboration, A. O'Neill, A. Rehman and 291 others Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv 2310.08864 v9, 2023-10-13.
- A. Khazatsky, K. Pertsch, S. Nair and 98 others DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset. arXiv 2403.12945 v2, 2024-03-19.
- H. Walke, K. Black, A. Lee and 11 others BridgeData V2: A Dataset for Robot Learning at Scale. arXiv 2308.12952 v3, 2023-08-24.
- R. Hoque, P. Huang, D. J. Yoon and 2 others EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video. arXiv 2505.11709 v3, 2025-05-16.
- C. Chi, Z. Xu, C. Pan and 5 others Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. arXiv 2402.10329 v3, 2024-02-15.
- R. Zheng, D. Niu, Y. Xie and 12 others EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data. arXiv 2602.16710 v1, 2026-02-18.
- S. Gao, W. Liang, K. Zheng and 27 others DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos. arXiv 2602.06949 v1, 2026-02-06.
- Dyna Robotics Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models. dyna.co, 2026-08.
- T. Z. Zhao, V. Kumar, S. Levine, C. Finn Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv 2304.13705 v1, 2023-04-23.
- K. Black, M. Y. Galliker, S. Levine Real-Time Execution of Action Chunking Flow Policies. arXiv 2506.07339 v2, 2025-06-09.
- NVIDIA Newsroom NVIDIA Announces NVIDIA Isaac GR00T Reference Humanoid Robot for Academic Research. nvidianews.nvidia.com, 2026-05-31.
- NVIDIA Technical Blog Develop Humanoid Robot Policies End-to-End with NVIDIA Isaac GR00T. developer.nvidia.com, 2026-07-07.
- SVRC (Robotics Center of Silicon Valley) State of Robotics 2026. roboticscenter.ai, 2026-03.
- K. Zhou, Q. Chen, D. Peng and 3 others Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment. arXiv 2604.24447 v1, 2026-04-27.
- W. Wu, F. Lu, Y. Wang and 22 others A Pragmatic VLA Foundation Model. arXiv 2601.18692 v4, 2026-01-26.
- Motubrain Team, C. Xiang, F. Bao and 17 others Motubrain: An Advanced World Action Model for Robot Control. arXiv 2604.27792 v5, 2026-04-30.