Shannon Press · 2026 · 19 chapters
A robot that has watched the internet still has to learn to move.
This handbook explains the models that let robots learn from web video, from people and from their own imagination: how each one works, what it is good and bad at, and what remains unsolved. It is organised by mechanism, not by product name, so it survives the next release cycle.
Nineteen chapters move from what a robot foundation model is, through the seven families that make up the field, to data, post-training, deployment and the problems still open, with every claim sourced and every idea made into something you can play with.
The method
Three things in every chapter
FFigures
Every chapter opens with a figure and explains each idea with an interactive diagram or a live simulation you can push until it breaks.
SSources
Every claim is cited inline, every quote is verbatim and linked, and every number a vendor reports about its own model is tagged as a claim.
RRust
Every chapter has a crate with three exercises, reference solutions and tests. Code shown on the page is read from the crate, so it compiles.
The map
Seven families on one grid
Every model in the book answers two questions: what does it predict, and in what space? Where it lands on this grid decides what data it can learn from and what it costs to run.
- Reactive VLAs
- World-action models
- Latent prediction
- World simulators
- Geometry-first
- Embodied reasoning
- Large behavior models
Reading paths
Three ways to read it
Executive path
about 25 pages
Chapter openers and all Part II figures, then chapters 12, 17 and 19.
Engineer path
the whole book
Everything, in order. The Rust exercises assume this path.
Researcher path
about 90 pages
Chapter 4, all of Part II, Part IV and the sources.
Contents
Four parts, nineteen chapters
Foundations
- CH.01What is a robot foundation modelDefine the term precisely enough to use and loosely enough to survive new releases, and contrast it with task-specific policies, modular pipelines and plain vision-language models.
- CH.02Why now: the convergenceExplain the timing as six ingredients that matured within roughly the same 24 months.
- CH.03A short history in five phasesGive the newcomer the arc, so every later chapter reads as a response to a prior limitation.
- CH.04The reader's toolkitThe one chapter with real notation: policies, action chunks, world models, latent actions and the four-component decomposition.
The seven families
- CH.05Reactive vision-language-action modelsA pretrained VLM backbone plus an action head: observations and language in, an action chunk out, no modelling of the future.
- CH.06World-action modelsCouple future-state prediction with action generation in one model or one training loop.
- CH.07Latent prediction and latent action modelsPredict in representation space rather than pixel space, or infer latent actions so that action-free video becomes training signal.
- CH.08World models as simulators, data engines and RL environmentsThe world model is not the policy: it generates futures, synthetic trajectories or imagined environments that other models train or plan in.
- CH.09Geometry-first and structured world modelsPredict in an explicit spatial or physical representation (points, Gaussians, particles, occupancy, touch) instead of RGB.
- CH.10Embodied reasoning modelsVLM-level models for spatial understanding, task decomposition and success detection that sit above a controller from another family.
- CH.11Large behavior models and scaled imitationImitation learning at scale without a language backbone: Diffusion Policy, ACT and their descendants.
- CH.12Choosing between familiesThe practitioner chapter: a comparison matrix, a decision flow and three worked scenarios.
Cross-cutting engineering
- CH.13Data: the pyramid, scaling and provenanceEvery family's strengths trace to which data layer it can use. This chapter makes that explicit.
- CH.14Post-training: adaptation, RL, steerability, memoryFine-tuning economics, three routes to reinforcement learning, steerability and memory.
- CH.15Deployment, evaluation and industrial readinessLatency budgets, compute placement, benchmarks and what a readiness assessment looks like for a real cell.
- CH.16Applications by domainManipulation, humanoids, navigation, driving, industrial and service robots: which family dominates and what shipped.
Open problems and futures
- CH.17Open problemsThe ten problems the 2026 surveys converge on, each with a concrete failure example.
- CH.18Candidate solutions and research betsFor each open problem: the approaches being tried, who is trying them, and how mature they are.
- CH.19Three scenarios for 2027 to 2030Turn the open problems into futures the reader can watch for, without picking one.