Ch.01Part I · Foundations
What is a robot foundation model
Define the term precisely enough to use and loosely enough to survive new releases, and contrast it with task-specific policies, modular pipelines and plain vision-language models.
- Math
- L1
- Sources
- 13
- Figures
- 4
- Exercises
- 3
- Last checked
- 3 Oct 2026
One boxed equation per idea
5 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
- Reactive VLAs
- World-action models
- Latent prediction
- World simulators
- Geometry-first
- Embodied reasoning
- Large behavior models
Why this chapter exists
"Robot foundation model" is used for almost anything a robotics company ships in 2026. Vendors apply it to vision-language-action policies, to video generators, to reasoning models that never touch a motor, and sometimes to a fine-tuned version of last year's product. A term that covers everything explains nothing, so this chapter pins it down: precisely enough to be useful, loosely enough to survive the next twelve months of releases.
The definition matters for the rest of the book in a practical way. The seven families of Part II are seven answers to two questions about a model: what does it predict, and in what space? If you can answer those from an abstract, you can place a new model within a minute and know which chapter explains how it works and where it breaks.
After this chapter you will be able to:
- say what makes a model a robot foundation model, and name three things that are useful to roboticists but are not;
- state the minimum interface every such model shares, and the one optional output that separates reactive models from world-action models;
- place a new model on the family grid from its abstract, and recognise the cases that sit on a boundary.
To be useful and helpful to people, AI models for robotics need three principal qualities: they have to be general, meaning they’re able to adapt to different situations; they have to be interactive, meaning they can understand and respond quickly to instructions or changes in their environment; and they have to be dexterous, meaning they can do the kinds of things people generally can do with their hands and fingers, like carefully manipulate objects.
A working definition
Foundation models, in the sense the robotics surveys borrow from language and vision, are large networks trained on massive, heterogeneous data that are then adapted to many downstream tasks (Psiris et al. 2026, arXiv 2604.15395). A 2023 survey put the robotics version of the question directly: how can existing foundation models from language and vision be used for general-purpose robots, and what would a robotics-specific foundation model look like (Hu et al. 2023, arXiv 2312.08782)? Three years later the second question has dozens of answers.
This book uses a definition with three tests. A robot foundation model is a model that:
- produces something a robot can act on: an action chunk, a prediction of the future that is used for control, or a plan that a controller below it executes;
- was pretrained broadly: across many tasks, many robot bodies (embodiments), or web-scale video and images, rather than on one task in one lab;
- adapts cheaply: by prompting, by conditioning on a goal, or by fine-tuning on a few hundred demonstrations rather than retraining.
The tests are deliberately about function, not architecture or size. They admit SmolVLA, a policy designed to be trained on a single GPU (Shukor et al. 2025, arXiv 2506.01844), and exclude the largest chat model that cannot move an arm.
The minimum interface
Every model in the book, whatever its internals, presents the same interface to the robot: observations and language in, an action chunk out. Observations are recent camera frames plus the robot's own joint state; the instruction is a sentence; the output is actions to execute at a fixed control rate. One inference produces many motor commands, which is what makes slow models usable on fast robots (chapter 2 shows the arithmetic).
The one optional output is the future. A world-action model also returns what it expects to see. That single addition is the line between chapter 5's reactive models and chapter 6's world-action models, and it is why the grid has a column for it.
In code the interface is one required method and one optional one. The chapter's Rust crate states it as a trait, and every later chapter's code implements some version of it:
/// What the robot senses at one instant: camera frames and its own joint state.
#[derive(Clone, Debug, Default)]
pub struct Observation {
pub frames: Vec<Vec<u8>>,
pub proprio: Vec<f32>,
}
/// H actions to execute at a fixed control rate. One inference, many motor commands.
#[derive(Clone, Debug, PartialEq)]
pub struct ActionChunk {
pub rate_hz: f32,
pub actions: Vec<Vec<f32>>,
}
/// The minimum interface of a robot foundation model: observations and language in,
/// an action chunk out. Predicting the future is optional, which is exactly the line
/// between reactive models and world-action models.
pub trait RobotFoundationModel {
fn act(&self, history: &[Observation], instruction: &str) -> ActionChunk;
/// Imagined future observations, if the model makes them.
fn imagine(&self, _history: &[Observation], _instruction: &str) -> Option<Vec<Observation>> {
None
}
}What the definition leaves out
Three neighbours are close enough to confuse and different enough to keep out.
A task-specific policy can be excellent: a diffusion policy trained on two hundred demonstrations of one task can be more reliable at that task than any generalist. It fails the second test. It has no breadth to adapt from, and a new task means new data and new training.
A classical modular pipeline splits perception, state estimation, planning and control into separate blocks, each designed and verified on its own. Its strengths, debuggability and certifiability, are exactly what learned models gave up (chapter 3 tells the story, chapter 15 counts the cost). It fails the first test as a whole: no single learned model produces the action.
A plain vision-language model answers questions about images. It has the breadth and adapts by prompting, but produces nothing a robot acts on. Fine-tuned to output actions it becomes the backbone of a chapter 5 model; trained for spatial reasoning and success detection it becomes a chapter 10 model. Off the shelf it is neither.
For each system, decide whether the book counts it as a robot foundation model. Then check.
- π0.5 a VLA pretrained on many robots and tasks
- DreamZero a world-action model on a video backbone
- Cosmos-Predict2.5 a video world model used to generate robot training data
- Gemini Robotics-ER 1.6 an embodied reasoning model that plans and judges success
- An open-vocabulary object detector finds objects in images from text queries
- A chat model writing code that calls hand-written skills an LLM planner
- A Diffusion Policy trained for one task on one robot a strong imitation policy
- TRI's Large Behavior Model a diffusion policy pretrained on many dexterous tasks
The industrial survey of March 2026 shows how large the category already is under a definition like this one: it evaluates 324 manipulation-capable robot foundation models against 149 readiness criteria (Kube et al. 2026, arXiv 2603.06749). Its verdict is the useful corrective to vendor language. Industrial maturity is limited and uneven, and even the highest-rated models satisfy only a fraction of the criteria.
Two questions that sort everything
The family grid of figure 1.1 asks two questions of every model.
What does it predict? Actions only, as in a reactive vision-language-action model (VLA); actions together with predicted futures, as in a world-action model (WAM); futures only, as in a world model used to plan or to generate data; or a plan in language, as in an embodied reasoning model.
In what space? Raw pixels or the latents of a video model; learned feature latents that are never decoded back to pixels (the JEPA approach of chapter 7); explicit geometry such as point clouds, Gaussians or occupancy grids (chapter 9); language or action tokens; or continuous actions with no language backbone at all (chapter 11's large behavior models).
The answers decide almost everything that matters in practice: which data a model can learn from, what it costs to run, and how it fails. A model that predicts in pixels can learn from internet video; a model that predicts only actions needs action-labelled robot data; a model that predicts in geometry needs depth or point clouds. Chapter 13's data pyramid makes this precise.
Which column of the family grid does each model belong in? Read the hint as you would read an abstract.
- OpenVLA predicts discretised action tokens from an image and an instruction
- Cosmos Policy encodes actions, future states and values as latent video frames
- PointWorld predicts how a point cloud moves under a robot's action
- Gemini Robotics-ER 2 tracks task progress and orchestrates steps for a robot
- Fast-WAM trained to predict video and actions, runs action-only at test time
The chapter's crate turns the two questions into a function. A model card records what a model outputs and in what space; placing it on the grid is then mechanical, and exercise 1.1 asks you to write it:
/// Down the family grid: the space a model predicts in.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum Space {
PixelsOrVideoLatents,
FeatureLatents,
ExplicitGeometry,
LanguageOrActionTokens,
ContinuousNoLanguage,
}
/// Across the family grid: what a model predicts.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum Predicts {
ActionsOnly,
ActionsPlusFutures,
FuturesOnly,
PlanInLanguage,
}
/// The facts about a model that decide where it goes, read off its paper or model card.
#[derive(Clone, Debug)]
pub struct ModelCard {
pub name: &'static str,
pub outputs_actions: bool,
pub outputs_futures: bool,
pub outputs_language_plan: bool,
pub space: Space,
/// Pretrained across many tasks, embodiments or web-scale data rather than one task.
pub broad_pretraining: bool,
/// Adapted to new tasks by prompting or light fine-tuning.
pub adapts_cheaply: bool,
/// Trained on embodied data (robot or egocentric), not only web text and images.
pub embodied_training: bool,
}Three boundary cases
The grid is approximate on purpose. Three models show how to handle a model that fits two cells.
π0.7 conditions on goal images as well as language: you can show it what the end of the task should look like (Physical Intelligence et al. 2026, arXiv 2604.15483). The WAM survey reads it as a goal-conditioned world-action model for exactly that reason (Lu et al. 2026, arXiv 2609.16074). This book places it with the reactive models of chapter 5, because at inference it never generates a future: the goal image is an input, not a prediction. The rule the book uses for boundary cases is to place a model by what it predicts at inference, and to say so when training tells a different story.
Octo is an open generalist policy trained on 800,000 trajectories from Open X-Embodiment and instructed by language commands or goal images (Octo Model Team et al. 2024, arXiv 2405.12213). It could sit with the VLAs of chapter 5 or the large behavior models of chapter 11. Its language encoder is small and its strength is the scaled imitation recipe, so chapter 11 discusses it as the bridge between the two.
GR00T N1.7 is a VLA by the grid's first reading, built on a Cosmos-Reason2 backbone with a diffusion transformer action head. For whole-body humanoid control it emits compact latent action tokens that a separate learned controller expands into joint commands (NVIDIA 2026). That path is a latent-action interface, chapter 7's territory. One product, two families, depending on the robot it drives.
How the families map to the surveys
The seven families are this book's synthesis, not a published taxonomy, and a reader who knows the surveys should be able to move between them. The WAM survey of September 2026 builds world-action models from four research threads: model-based reinforcement learning and world models, video generative models, latent action pretraining, and vision-language-action models (Lu et al. 2026, arXiv 2609.16074). The April 2026 comprehensive review organises the wider field by model type (language, vision, vision-language and vision-language-action models), architecture, learning paradigm and application (Psiris et al. 2026, arXiv 2604.15395). A June 2026 survey of world-action models offers two further views of the same models (Shen et al. 2026, arXiv 2606.20781).
| Book family | Chapter | Closest thread in the WAM survey | Closest model type in the 2026 review |
|---|---|---|---|
| Reactive VLAs | 5 | Vision-language-action models | VLAs |
| World-action models | 6 | All four threads combined | VLAs with predictive components |
| Latent prediction and latent actions | 7 | Latent action pretraining; world models | Vision foundation models plus VLAs |
| World models as simulators | 8 | World models; video generative models | Vision foundation models |
| Geometry-first models | 9 | World models (structured, 3D) | Vision foundation models |
| Embodied reasoning | 10 | Not a WAM thread | Vision-language models |
| Large behavior models | 11 | Not a WAM thread | Learning paradigms (imitation) |
Exercises
The stubs are in rust/ch01-interface/src/exercises.rs; reference solutions compile with --features solutions.
Exercise 1.1
Place a model
Implement place_on_grid: return the grid cell for a model card, or None for a model that predicts nothing the grid covers. The tests place π0.5, DreamZero, V-JEPA 2-AC, PointWorld, Gemini Robotics-ER 1.6 and TRI's large behavior model, and expect a perception-only model to have no cell.
Check your answer from the rust/ folder:
cargo test -p ch01-interface --test exercises ex1_1Hint 1
The column depends only on the three output flags; the row is the card's space.
Hint 2
Order the checks so that a model with both actions and futures is not caught by the actions-only branch.
Exercise 1.2
Apply the definition
Implement is_in_scope with the book's three tests plus the requirement of embodied training. The tests expect eight systems in and three out: a perception-only model, a general language model planning over hand-written skills, and a single-task diffusion policy.
Check your answer from the rust/ folder:
cargo test -p ch01-interface --test exercises ex1_2Hint 1
Write the three tests from 'A working definition' as three booleans and combine them.
Hint 2
Embodied training is what separates an embodied reasoning model from a general chat model writing plans.
Exercise 1.3
Check a chunk before it runs
Implement chunk_duration_ms: reject empty chunks, chunks whose actions have different or zero dimension, and non-positive rates; otherwise return how long the chunk takes to execute. A 25-action chunk at 50 Hz lasts half a second, which is the window chapter 2 uses to hide inference latency.
Check your answer from the rust/ folder:
cargo test -p ch01-interface --test exercises ex1_3What we are not sure about
Whether the definition will hold. Vendors use the term loosely, and this book's three tests exclude some products marketed as robot foundation models and include some research systems that are not marketed that way. The test most likely to need revision is the third: as models get better at prompting and in-context learning (GEN-1.5 claims to learn a new task from a single demonstration (Generalist AI 2026)), "adapts cheaply" may stop distinguishing anything.
Whether seven is the right number. The families are this book's synthesis. They merge distinctions the surveys keep apart and split others the surveys merge. Embodied reasoning and large behavior models are the most likely to be folded into neighbouring chapters if the field consolidates.
Where the boundary cases go. π0.7, Octo and GR00T N1.7 each fit two cells, and the rule "place by what is predicted at inference" is a choice. A reader who places by training objective will move several models, Fast-WAM among them, into different chapters.
Further reading
- The April 2026 comprehensive review, for the broadest map of the field and its five research phases (Psiris et al. 2026, arXiv 2604.15395).
- The 2023 survey and meta-analysis that first asked what a robotics-specific foundation model would look like (Hu et al. 2023, arXiv 2312.08782).
- The industrial readiness survey, for how far the category still is from the factory floor (Kube et al. 2026, arXiv 2603.06749).
- A 2025 anatomy of vision-language-action models, from modules to milestones (Xu et al. 2025, arXiv 2512.11362).
- Two earlier surveys of foundation models for manipulation and robot learning (Li et al. 2024, arXiv 2404.18201; Xiao et al. 2023, arXiv 2311.14379).
References
13 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- A. Psiris, V. Argyriou, E. K. Markakis and 5 others Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions. Transactions on Machine Learning Research (TMLR), 07/2026. arXiv 2604.15395 v3, 2026-04-16.
- Y. Hu, Q. Xie, V. Jain and 20 others Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis. arXiv 2312.08782 v3, 2023-12-14.
- M. Shukor, D. Aubakirova, F. Capuano and 11 others SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics. arXiv 2506.01844 v1, 2025-06-02.
- D. Kube, S. Hadwiger, T. Meisen Robotic Foundation Models for Industrial Control: A Comprehensive Survey and Readiness Assessment Framework. arXiv 2603.06749 v1, 2026-03-06.
- Physical Intelligence, B. Ai, A. Amin and 85 others π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities. arXiv 2604.15483 v2, 2026-04-16.
- Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.
- Octo Model Team, D. Ghosh, H. Walke and 16 others Octo: An Open-Source Generalist Robot Policy. arXiv 2405.12213 v2, 2024-05-20.
- NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
- Q. Shen, S. Zhang, Y. Liao and 5 others World Action Models: A Survey. arXiv 2606.20781 v1, 2026-06-18.
- Generalist AI GEN-1.5: Embodied Foundation Models are One-Shot Learners. generalistai.com, 2026-08-19.
- C. Xu, S. Zhang, Y. Liu and 11 others An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges. arXiv 2512.11362 v3, 2025-12-12.
- D. Li, Y. Jin, Y. Sun and 11 others What Foundation Models can Bring for Robot Learning in Manipulation : A Survey. arXiv 2404.18201 v7, 2024-04-28.
- X. Xiao, J. Liu, Z. Wang and 5 others Robot Learning in the Era of Foundation Models: A Survey. arXiv 2311.14379 v1, 2023-11-24.