Ch.03Part I · Foundations
A short history in five phases
Give the newcomer the arc, so every later chapter reads as a response to a prior limitation.
- Math
- L0
- Sources
- 31
- Figures
- 3
- Exercises
- 3
- Last checked
- 3 Oct 2026
No equations
7 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
WAM survey 2026-09
World-action models named and mapped as a family.
Why this chapter exists
A newcomer to robot foundation models meets a wall of names: RT-2, Octo, π0.5, GR00T, DreamZero. Read in isolation they look like a list of products. Read in order they are an argument, each generation fixing what the previous one could not do and breaking something it took for granted. This chapter gives the order, so that every later chapter reads as an answer to a question the field was already asking.
After this chapter you will be able to:
- tell the history of robot foundation models in five phases, and say what each fixed and what it broke;
- trace any current model back through the lineage of figure 3.3, distinguishing a team's next version from an idea borrowed from elsewhere;
- move between this book's phases and the five research phases of the 2026 comprehensive review.
Those speeds are very different and driven by very different realities. I think that many people get confused by that and make the mistake of jumping between those domains of reality, thinking all the speeds will be the same.
Brooks is writing about the speed of research ideas, of hype, and of deployment. Keep his distinction in mind: this chapter is a history of ideas, and chapter 16 is where deployment gets its own, slower, timeline.
Phase 1: modular pipelines (before 2017)
Classical robots were built from separate blocks: perception, state estimation, planning and control. Each block could be designed, validated and replaced on its own, and the representations between them could be inspected, which made the whole debuggable and, for industrial cells, certifiable (Lu et al. 2026, arXiv 2609.16074).
The approach works when the task, the dynamics and the environment can be specified precisely. It breaks in the open world. Hand-designed representations are brittle when conditions shift, an error at one block's interface propagates through the rest, and a model built for one task does not scale across scenes, robot bodies and instructions (Lu et al. 2026, arXiv 2609.16074). Everything that follows trades some of the modular pipeline's transparency for learning's breadth. Chapter 15 counts what was given up.
Phase 2: learning in the lab (2017 to 2023)
Two lines of learning research matured in parallel, both mostly in simulation and in single labs.
The first learned world models for control. PlaNet learned the dynamics of an environment from pixels and planned in a compact latent space (Hafner et al. 2018, arXiv 1811.04551). Dreamer learned behaviours purely by imagining trajectories inside such a model (Hafner et al. 2019, arXiv 1912.01603). TD-MPC paired short-horizon planning in a learned latent model with a learned value for the long run (Hansen et al. 2022, arXiv 2203.04955). DreamerV3 made the recipe general: one configuration across more than 150 tasks, and the first algorithm to collect diamonds in Minecraft from scratch (Hafner et al. 2023, arXiv 2301.04104).
The second made imitation work on real robots. Diffusion Policy represented a robot's policy as a conditional denoising process, which let it capture the several valid ways of doing a task instead of averaging them (Chi et al. 2023, arXiv 2303.04137). ACT predicted chunks of future actions with a transformer and learned fine bimanual skills on low-cost hardware (Zhao et al. 2023, arXiv 2304.13705). Chapter 4 explains both mechanisms; chapter 11 follows their descendants.
What it fixed: learning replaced hand design, and both lines worked on real problems. What it broke: each model was trained for one task or one robot, from data collected for it. Nothing transferred, and nothing understood an instruction.
Phase 3: language and vision enter (2022 to 2024)
The third phase borrowed understanding from models trained on the web.
RT-1 trained a robotics transformer on large, diverse real-world robot data and showed it absorbing that diversity as the model grew (Brohan et al. 2022, arXiv 2212.06817). RT-2 then took the step that named the field: it wrote robot actions as text tokens, so a vision-language model pretrained on the internet could be co-fine-tuned to output them, and web knowledge transferred to a robot arm (Brohan et al. 2023, arXiv 2307.15818). RoboFlamingo showed that an open vision-language model, lightly fine-tuned, made an effective imitator (Li et al. 2023, arXiv 2311.01378).
Data was pooled in the same years. Open X-Embodiment gathered over a million trajectories from 22 robot types (Open X-Embodiment Collaboration et al. 2023, arXiv 2310.08864), and two open generalist policies were trained on it: Octo, on 800,000 of those trajectories (Octo Model Team et al. 2024, arXiv 2405.12213), and OpenVLA, a 7-billion-parameter vision-language-action model trained on 970,000 demonstrations (Kim et al. 2024, arXiv 2406.09246).
What it fixed: robots could follow instructions about objects they had never been trained on, and the weights were open. What it broke: speed and precision. A large model emitting one discretised action token at a time was slow and coarse, which is what the next phase set out to solve.
Phase 4: generalist policies as products (2024 to 2025)
In the fourth phase companies began to sell the result.
Covariant announced RFM-1 in March 2024, an 8-billion-parameter model trained on text, images, video, robot actions and physical measurements, built on tens of millions of trajectories from its warehouse fleetvendor claim (Covariant 2024). Five months later Amazon licensed the technology and hired Covariant's founders (Covariant 2024), an early sign of what such models were worth.
Physical Intelligence's π0 added a flow-matching action expert to a pretrained vision-language model and trained it across many robot types, replacing slow discrete tokens with fast continuous action chunks (Black et al. 2024, arXiv 2410.24164). Figure's Helix split the work by speed for humanoids: a vision-language model running at 7 to 9 Hz for understanding, above a 200 Hz policy for control (Figure AI 2025). NVIDIA released GR00T N1 as an open dual-system foundation model for humanoids (NVIDIA et al. 2025, arXiv 2503.14734), and Google DeepMind launched Gemini Robotics, with physical actions as a new output of Gemini 2.0, alongside an embodied-reasoning model (Gemini Robotics Team et al. 2025, arXiv 2503.20020). π0.5 extended π0 to open-world generalisation (Physical Intelligence et al. 2025, arXiv 2504.16054), π*0.6 learned from its own experience with reinforcement learning (Physical Intelligence et al. 2025, arXiv 2511.14759), and Generalist AI's GEN-0 reported scaling laws from over 270,000 hours of real-world manipulation datavendor claim (Generalist AI 2025).
What it fixed: speed, dexterity and breadth, in products that customers could buy. What it broke: these models are reactive. They map what they see to what they do and do not model what happens next, which is where the fifth phase starts.
Phase 5: the world-action turn (2026)
In 2026 prediction of consequences moved into the policy. Cosmos Policy fine-tuned a video model to act and plan (Kim et al. 2026, arXiv 2601.16163); LingBot-VA learned frame prediction and action together (Li et al. 2026, arXiv 2601.21998); DreamZero denoised video and actions in one 14-billion-parameter model (Ye et al. 2026, arXiv 2602.15922); Fast-WAM asked whether the imagined video is needed at test time at all (Yuan et al. 2026, arXiv 2603.16666). NVIDIA previewed GR00T N2 in March as based on DreamZero (NVIDIA Newsroom 2026). The reactive line kept moving too: π0.7 became steerable through goal images, metadata and language (Physical Intelligence et al. 2026, arXiv 2604.15483). By September a survey named the family and mapped it (Lu et al. 2026, arXiv 2609.16074). Chapter 6 tells this phase in full.
What it fixed: reasoning about consequences, and learning from video without action labels. What it broke, so far: inference cost, and the simplicity of a model that only maps images to actions. Whether those are temporary is chapter 6's open question.
What each phase fixed and what it broke
| Phase | Fixed | Broke |
|---|---|---|
| 1. Modular pipelines | Transparent, certifiable, debuggable systems | Brittle outside the conditions they were designed for |
| 2. Learning in the lab | Behaviour learned rather than hand-designed | One model per task and per robot; no language |
| 3. Language and vision enter | Instructions and web knowledge transfer to robots; open weights | Slow, coarse discretised actions |
| 4. Policies as products | Fast continuous control, dexterity, breadth; commercial use | Reactive: no model of consequences |
| 5. The world-action turn | Predicting consequences; learning from action-free video | Inference cost; harder to inspect and certify |
The lineage
Figure 3.3 draws the same history as a family tree. Solid arrows join a team's versions; dashed arrows mark a model that builds on another's backbone, recipe or central idea. Select any model to light up everything it descends from. π0.7, for example, descends in a direct line from π*0.6, π0.5 and π0, and through π0 from ACT's chunking and Diffusion Policy's denoising.
π0.7 2026-04, phase 5: The world-action turn
Builds on π*0.6, π0.5, π0, ACT, Diffusion Policy.
The dashed edges are judgments, and one deserves a warning. DreamZero builds on Wan's video model directly; the edge from DreamerV3 records only that both belong to the world-model thread the WAM survey traces, not that DreamZero reuses Dreamer's code (Lu et al. 2026, arXiv 2609.16074).
The chapter's Rust crate stores the lineage as a small graph, which makes such judgments checkable:
/// A model, dated by its first public release (year, month).
#[derive(Clone, Debug)]
pub struct Model {
pub name: &'static str,
pub released: (u16, u8),
pub phase: Phase,
}
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum Link {
/// The same team's next version.
Successor,
/// Uses the earlier work as its backbone, its recipe or its central idea.
BuildsOn,
}
#[derive(Clone, Debug)]
pub struct Edge {
pub from: &'static str,
pub to: &'static str,
pub link: Link,
}
#[derive(Clone, Debug)]
pub struct Lineage {
pub models: Vec<Model>,
pub edges: Vec<Edge>,
}Writing that graph caught a real inconsistency, which exercise 3.2 asks you to reproduce. The first draft drew GR00T N2 as the successor of GR00T N1.7. But N2 was previewed in March 2026, a month before N1.7 shipped in April (NVIDIA Newsroom 2026; NVIDIA 2026). A model cannot descend from one released after it, so the graph now draws both as successors of N1.
This book's phases and the review's
The 2026 comprehensive review also tells the history in five phases (Psiris et al. 2026, arXiv 2604.15395). Its boundaries follow three axes: the main source of training data, how integrated the system is, and what the model outputs. This book keeps the idea of five phases but draws them around a different question, what each generation fixed and broke, and names them in its own words. The two readings line up like this:
| Review's phase | Years | Closest phase in this book |
|---|---|---|
| 1. Native language and vision models in robotic systems | 2018 to 2021 | 1 and 2: perception models inside modular pipelines; learning in the lab |
| 2. Grounded planning with vision-language representations | 2021 to 2022 | 3, early part |
| 3. Embodied vision-language-action policies | 2022 to 2023 | 3 |
| 4. Memory, autonomous task composition, web-to-robot transfer | 2023 to 2024 | 3 and 4 |
| 5. Multi-sensory generalisation and real-world deployment | 2024 to present | 4 and 5 |
The main difference is the last row. The review's fifth phase runs to the present; this book separates 2026 as the world-action turn, because prediction of consequences entering the policy changes what the model is, not only where it is deployed.
Exercises
The stubs are in rust/ch03-lineage/src/exercises.rs.
Exercise 3.1
Where did it come from?
Implement ancestors: every model a given model descends from, through any chain of edges. Then compare your answer for GR00T N2 with figure 3.3.
Check your answer from the rust/ folder:
cargo test -p ch03-lineage --test exercises ex3_1Hint 1
Walk backwards over edges from the model, keeping a set of models already found.
Hint 2
Both kinds of link count as descent.
Exercise 3.2
Nothing descends from the future
Implement inconsistent_edges: edges that point to an earlier model or name a model not in the list. Add the edge the first draft had, from GR00T N1.7 to GR00T N2, and confirm your function catches it.
Check your answer from the rust/ folder:
cargo test -p ch03-lineage --test exercises ex3_2Hint
Look up both ends of every edge; an edge whose end is missing is inconsistent too.
Exercise 3.3
The longest family line
Implement longest_successor_chain, the length of the longest chain of same-team versions. Which line is it, and what does that say about how long a single lab has stayed at the front?
Check your answer from the rust/ folder:
cargo test -p ch03-lineage --test exercises ex3_3What we are not sure about
Where the phase boundaries fall. They are this book's editorial choice, and real work overlaps: Diffusion Policy and ACT belong to the lab phase in spirit and to 2023 in date, the same year RT-2 opened the next phase. The review draws its boundaries differently, for stated reasons, and neither is the history.
Who was first. Several firsts are contested. NVIDIA called GR00T N1 the world's first open humanoid robot foundation model; Figure called Helix the first VLA to control a whole humanoid upper body at high rate; Covariant called RFM-1 the first time generative AI gave commercial robots a deeper understanding of language and the physical world. This book reports each claim and its claimant rather than adjudicating (NVIDIA Newsroom 2025; Figure AI 2025; Covariant 2024).
Whether the fifth phase is a phase. The world-action turn is less than a year old. If chapter 6's open question resolves toward "video only in training", 2026 may read in retrospect as a change in training recipe rather than in what a robot model is.
Further reading
- The 2026 comprehensive review's history and its five research phases (Psiris et al. 2026, arXiv 2604.15395).
- RT-2, the paper that named vision-language-action models (Brohan et al. 2023, arXiv 2307.15818).
- DreamerV3, the most general statement of the world-model thread (Hafner et al. 2023, arXiv 2301.04104).
- π0, the template for the policies of phase 4 (Black et al. 2024, arXiv 2410.24164).
- A 2025 anatomy of vision-language-action models, from modules to milestones (Xu et al. 2025, arXiv 2512.11362).
References
31 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.
- D. Hafner, T. Lillicrap, I. Fischer and 4 others Learning Latent Dynamics for Planning from Pixels. arXiv 1811.04551 v5, 2018-11-12.
- D. Hafner, T. Lillicrap, J. Ba, M. Norouzi Dream to Control: Learning Behaviors by Latent Imagination. arXiv 1912.01603 v3, 2019-12-03.
- N. Hansen, X. Wang, H. Su Temporal Difference Learning for Model Predictive Control. arXiv 2203.04955 v2, 2022-03-09.
- D. Hafner, J. Pasukonis, J. Ba, T. Lillicrap Mastering Diverse Domains through World Models. arXiv 2301.04104 v2, 2023-01-10.
- C. Chi, Z. Xu, S. Feng and 5 others Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. arXiv 2303.04137 v5, 2023-03-07.
- T. Z. Zhao, V. Kumar, S. Levine, C. Finn Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv 2304.13705 v1, 2023-04-23.
- A. Brohan, N. Brown, J. Carbajal and 48 others RT-1: Robotics Transformer for Real-World Control at Scale. arXiv 2212.06817 v2, 2022-12-13.
- A. Brohan, N. Brown, J. Carbajal and 51 others RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv 2307.15818 v1, 2023-07-28.
- X. Li, M. Liu, H. Zhang and 9 others Vision-Language Foundation Models as Effective Robot Imitators. arXiv 2311.01378 v3, 2023-11-02.
- Open X-Embodiment Collaboration, A. O'Neill, A. Rehman and 291 others Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv 2310.08864 v9, 2023-10-13.
- Octo Model Team, D. Ghosh, H. Walke and 16 others Octo: An Open-Source Generalist Robot Policy. arXiv 2405.12213 v2, 2024-05-20.
- M. J. Kim, K. Pertsch, S. Karamcheti and 15 others OpenVLA: An Open-Source Vision-Language-Action Model. arXiv 2406.09246 v3, 2024-06-13.
- Covariant Covariant introduces RFM-1 to give robots the human-like ability to reason. covariant.ai, 2024-03-11.
- K. Black, N. Brown, D. Driess and 21 others π0: A Vision-Language-Action Flow Model for General Robot Control. arXiv 2410.24164 v4, 2024-10-31.
- Figure AI Helix: A Vision-Language-Action Model for Generalist Humanoid Control. figure.ai, 2025-02-20.
- NVIDIA, J. Bjorck, F. Castañeda and 39 others GR00T N1: An Open Foundation Model for Generalist Humanoid Robots. arXiv 2503.14734 v2, 2025-03-18.
- Gemini Robotics Team, S. Abeyruwan, J. Ainslie and 115 others Gemini Robotics: Bringing AI into the Physical World. arXiv 2503.20020 v1, 2025-03-25.
- Physical Intelligence, K. Black, N. Brown and 33 others π0.5: a Vision-Language-Action Model with Open-World Generalization. arXiv 2504.16054 v1, 2025-04-22.
- Physical Intelligence, A. Amin, R. Aniceto and 53 others π*0.6: a VLA That Learns From Experience. arXiv 2511.14759 v2, 2025-11-18.
- Generalist AI GEN-0: Embodied Foundation Models That Scale with Physical Interaction. generalistai.com, 2025-11-04.
- M. J. Kim, Y. Gao, T. Lin and 8 others Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning. arXiv 2601.16163 v1, 2026-01-22.
- L. Li, Q. Zhang, Y. Luo and 9 others Causal World Modeling for Robot Control. arXiv 2601.21998 v2, 2026-01-29.
- S. Ye, Y. Ge, K. Zheng and 33 others World Action Models are Zero-shot Policies. arXiv 2602.15922 v1, 2026-02-17.
- T. Yuan, Z. Dong, Y. Liu, H. Zhao Fast-WAM: Do World Action Models Need Test-time Future Imagination?. arXiv 2603.16666 v2, 2026-03-17.
- NVIDIA Newsroom NVIDIA and Global Robotics Leaders Take Physical AI to the Real World (GTC 2026, GR00T N2 preview). nvidianews.nvidia.com, 2026-03-16.
- Physical Intelligence, B. Ai, A. Amin and 85 others π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities. arXiv 2604.15483 v2, 2026-04-16.
- NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
- A. Psiris, V. Argyriou, E. K. Markakis and 5 others Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions. Transactions on Machine Learning Research (TMLR), 07/2026. arXiv 2604.15395 v3, 2026-04-16.
- NVIDIA Newsroom NVIDIA Announces Isaac GR00T N1, the World's First Open Humanoid Robot Foundation Model, and Simulation Frameworks to Speed Robot Development. nvidianews.nvidia.com, 2025-03-18.
- C. Xu, S. Zhang, Y. Liu and 11 others An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges. arXiv 2512.11362 v3, 2025-12-12.