Ch.12Part II · The seven families
Choosing between families
The practitioner chapter: a comparison matrix, a decision flow and three worked scenarios.
- Math
- L0
- Sources
- 6
- Figures
- 3
- Exercises
- 3
- Last checked
- 9 Oct 2026
No equations
3 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
| Criterion | Reactive VLAsch. 5 | World-action modelsch. 6 | Latent predictionch. 7 | World simulatorsch. 8 | Geometry-firstch. 9 | Embodied reasoningch. 10 | Large behavior modelsch. 11 |
|---|---|---|---|---|---|---|---|
| Cheap inference | |||||||
| Little robot data needed | |||||||
| Language grounding | |||||||
| Reasons about consequences | |||||||
| Cross-embodiment transfer | |||||||
| Maturity | |||||||
| Open weights | |||||||
| Evidence quality | |||||||
| Weighted score | 2.38 | 2.13 | 2.13 | 2.13 | 1.50 | 2.13 | 2.13 |
Reactive VLAs: reasons about consequences (low)
Predicts no futures at all.
Weights: how much each criterion matters for your job. With these weights the top family is Reactive VLAs (2.38), then World-action models (2.13).
Why this chapter exists
Part II described seven families one at a time. A practitioner does not meet them that way. They meet a job: a cell to automate, a robot to make useful, a research question to answer. They want to know which family to start with, what it will cost in data and compute, and how it will fail first.
This chapter adds no new models. It puts the seven families side by side, turns the comparison into a short sequence of questions, and walks three realistic jobs through it. Treat it as a map drawn on a particular date. The matrix encodes the authors' judgement as of October 2026, and some of it will be wrong within a year.
After this chapter you will be able to:
- read the comparison matrix of figure 12.1, including which cells rest on independent evidence and which on vendor claims;
- walk a new job through the decision flow of figure 12.2 and name a first family to try;
- predict the first failure mode for that choice, and what to measure to catch it.
We've been trying to figure out how do we not just make a humanoid robot, but also make a humanoid robot that does useful work.
Reading the matrix
Each column of figure 12.1 is a family and each row a criterion that a deployment team cares about. The scores have three levels, so that nobody mistakes them for measurements. Higher is always better, so "cheap inference" scores high for a small, fast model and low for one that generates video.
Three rows deserve a warning. Little robot data needed mixes two things: how much action-labelled data the family needs to pretrain, and how much a new task needs. Large behavior models need few demonstrations per task once pretrained, but every task must still be demonstrated, so the row scores them low. Open weights counts what can be downloaded today, not what is promised. Evidence quality is the row to read first. It says whether a family's other scores rest on shared benchmarks and independent evaluation, or mostly on what the makers report. Reactive VLAs and large behavior models score high there. World models used as simulators and embodied reasoning models score low, because most of what is known about them comes from vendor posts and surveys.
Two things are worth trying in the figure. First, with all weights equal, reactive VLAs come out on top, mostly on maturity, open weights and evidence. Second, the ranking is fragile. Raising the weight on reasoning about consequences from 1 to 2.25 hands first place to world-action models, level with world models used as simulators, and raising the weight on cheap inference by the same amount hands it to large behavior models. Exercise 12.3 measures how much each criterion must move to change the answer. One structural fact survives every weighting: no family is dominated, meaning none is matched or beaten on every criterion by another. That is the strongest argument that the seven families are real alternatives rather than a ranking in disguise. Exercise 12.2 checks it.
The decision walk-through
Figure 12.2 turns the matrix into four questions. They are ordered from the cheapest answer to the most expensive, so that a job stops at the simplest family that fits.
Is it a short-horizon task with objects like those in the training data? If yes, start with a reactive VLA. It is the most documented, best-supported and cheapest family to fine-tune, and chapter 5 showed that hundreds of demonstrations often suffice. Most industrial jobs that are automated today look like this.
Does it involve contact, occlusion, or effects that appear only later? If yes, a reactive policy will make locally sensible moves that ruin later steps, which is the gap chapter 6 opened with. Start with a world-action model, or a latent-prediction model if inference cost matters (chapter 7). If the contact state itself is what matters, as in insertion or cloth handling, add geometry-first models to the list (chapter 9).
Is there action-labelled robot data for this body? If no, the realistic route is to pretrain on video with latent actions and adapt with a small decoder (chapter 7), and to use a world model to multiply whatever data you collect (chapter 8).
Is the robot restricted to onboard compute? If yes, a quantised VLA on the robot, with an onboard reasoning model only if it fits. If a network is available, put an embodied reasoning model above a controller (chapter 10). The controller can be a VLA, or a large behavior model where the tasks are fixed and speed matters most (chapter 11).
The flow is deliberately simple, and real jobs blur its questions. A household task can be short-horizon in one room and long-horizon in the next. The questions are still the right ones to ask first, because each one corresponds to a property that separates the families: horizon, contact, data and compute.
Three worked scenarios
A warehouse pick-and-place cell
Task. Pick known products from totes and place them in boxes, for eight hours a day.
- Reactive VLAs. A reactive VLA fine-tuned on a few hundred demonstrations of this cell (chapter 5).
First failure to expect. New packaging: transparent, reflective or deformable items the demonstrations never covered.
What to measure. Success with confidence intervals over enough trials to see a few-point drop (chapter 11), re-measured whenever the product range changes.
A humanoid tidying a home
Task. Put away groceries and tidy a living room, from spoken instructions, in homes it has not seen.
- World-action models. A world-action or latent-prediction model as the controller (chapters 6 and 7).
- Embodied reasoning. An embodied reasoning model above it to plan and to judge each step (chapter 10).
First failure to expect. A failed step judged done, so the robot moves on and the task silently fails (figure 10.2).
What to measure. False successes per task, not only completion rate; and behaviour after a nudge or a dropped object.
A mobile inspection robot
Task. Walk an industrial facility, read gauges and sight glasses, and report anomalies.
- Embodied reasoning. A reasoning model orchestrating the robot's navigation and arm APIs, and reading the instruments (chapter 10).
First failure to expect. Misread instruments in poor light or at oblique angles, and stalls when the network drops.
What to measure. Reading error against a human inspector's readings, and behaviour with the network cut.
The three cards share a pattern worth making explicit. The first failure is rarely the model's headline weakness. For the warehouse VLA it is not reasoning about consequences, which the job does not need, but distribution shift in the products. For the household humanoid it is not the controller but the judgement of when a step is done. For the inspection robot it is perception and connectivity. In each case the useful question is not "which family is best?" but "how will this family fail on this job, and how will I know?"
In code
The decision flow fits in a dozen lines. Writing it as code makes its assumptions inspectable, and easy to argue with.
#[derive(Clone, Copy, Debug, PartialEq, Eq, PartialOrd, Ord, Hash)]
pub enum Family {
ReactiveVla, // chapter 5
WorldAction, // chapter 6
Latent, // chapter 7
Simulator, // chapter 8
Geometry, // chapter 9
Reasoning, // chapter 10
LargeBehavior, // chapter 11
}
/// What you know about the job before choosing.
#[derive(Clone, Copy, Debug, Default)]
pub struct Job {
pub short_horizon_in_distribution: bool,
pub contact_occlusion_or_delayed_effects: bool,
pub robot_action_data_for_this_body: bool,
pub edge_compute_only: bool,
pub precise_contact_state_matters: bool,
}
/// The decision flow of figure 12.2: a first family to try, and what to consider next.
pub fn recommend(job: &Job) -> Vec<Family> {
use Family::*;
if job.short_horizon_in_distribution {
return vec![ReactiveVla];
}
if job.contact_occlusion_or_delayed_effects {
let mut v = vec![WorldAction, Latent];
if job.precise_contact_state_matters {
v.push(Geometry);
}
return v;
}
if !job.robot_action_data_for_this_body {
return vec![Latent, Simulator];
}
if job.edge_compute_only {
return vec![ReactiveVla, Reasoning];
}
vec![Reasoning, ReactiveVla, LargeBehavior]
}Exercises
The stubs are in rust/ch12-choosing/src/exercises.rs. This chapter is equation-free; the exercises are about weighing, comparing and stress-testing a judgement.
Exercise 12.1
A weighted score
Implement weighted_score and best. Then choose weights for a job you know, and check whether the answer agrees with the decision flow. If it does not, which of the two do you trust, and why?
Check your answer from the rust/ folder:
cargo test -p ch12-choosing --test exercises ex12_1Hint 1
Multiply each score by its weight, add them up, and divide by the total weight.
Hint 2
Break ties in favour of the family that comes first in ALL.
Exercise 12.2
Every family wins on something
Implement dominated. The book's matrix has no dominated family. Find the smallest change to one cell that makes one family dominated, and say whether you would believe the matrix more or less after that change.
Check your answer from the rust/ folder:
cargo test -p ch12-choosing --test exercises ex12_2Hint 1
A family is dominated if another is at least as good everywhere and strictly better somewhere.
Hint 2
Compare every pair; a family cannot dominate itself.
Exercise 12.3
How fragile is the ranking?
Implement flip. For each criterion, find how much its weight must rise from equal weights to change the winner. Which criteria can never dethrone the leader, and what does that tell you about why the leader leads?
Check your answer from the rust/ folder:
cargo test -p ch12-choosing --test exercises ex12_3Hint 1
Increase one weight in steps of 0.25 and recompute the winner each time.
Hint 2
Stop at the first change of winner, or give up at an increase of 10.
What we are not sure about
Whether the families stay separate. Chapters 5 to 11 each ended by noting that the family's boundary is blurring: VLAs adding future prediction, reasoning moving into controllers, imitation primitives everywhere. A matrix with seven columns may become a matrix of design choices within one model.
Whether the criteria are the right ones. The industrial survey uses 149 criteria; this matrix uses eight. Safety, which this book covers in chapters 15 and 17, is absent as a row because no family has enough evidence to score it.
How quickly the scores age. The evidence-quality and maturity rows are the most likely to change, in both directions, as independent evaluations of world-action models and reasoning models appear. Check the date on this chapter before relying on it.
Further reading
- The industrial readiness survey, for a far more detailed set of criteria (Kube et al. 2026, arXiv 2603.06749).
- The WAM survey, for the taxonomy that most of Part II builds on (Lu et al. 2026, arXiv 2609.16074).
- The April 2026 comprehensive review of foundation models in robotics, for the broader landscape (Psiris et al. 2026, arXiv 2604.15395).
- TRI's careful examination of large behavior models, for how to evaluate a choice once made (TRI LBM Team et al. 2025, arXiv 2507.05331).
- Rodney Brooks' annual predictions scorecard, for a long view on hype and deployment (Rodney Brooks 2026).
References
6 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- D. Kube, S. Hadwiger, T. Meisen Robotic Foundation Models for Industrial Control: A Comprehensive Survey and Readiness Assessment Framework. arXiv 2603.06749 v1, 2026-03-06.
- RoboCloud Hub NVIDIA GR00T N2 Explained; VLA Tutorial 2026. robocloudhub.tech, 2026.
- Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.
- A. Psiris, V. Argyriou, E. K. Markakis and 5 others Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions. Transactions on Machine Learning Research (TMLR), 07/2026. arXiv 2604.15395 v3, 2026-04-16.
- TRI LBM Team, J. Barreiros, A. Beaulieu and 79 others A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation. arXiv 2507.05331 v1, 2025-07-07.
- Rodney Brooks Predictions Scorecard, 2026 January 01. rodneybrooks.com, 2026-01-01.