Skip to content

Ch.12Part II · The seven families

Choosing between families

The practitioner chapter: a comparison matrix, a decision flow and three worked scenarios.

Math
L0

No equations

Sources
6

3 papers from 2026

Figures
3

Most are interactive

Exercises
3

Rust, with tests

Last checked
9 Oct 2026

The field moves monthly

Papers cited, by half-year of first arXiv version

CriterionReactive VLAsch. 5World-action modelsch. 6Latent predictionch. 7World simulatorsch. 8Geometry-firstch. 9Embodied reasoningch. 10Large behavior modelsch. 11
Cheap inference
Little robot data needed
Language grounding
Reasons about consequences
Cross-embodiment transfer
Maturity
Open weights
Evidence quality
Weighted score2.382.132.132.131.502.132.13

Reactive VLAs: reasons about consequences (low)

Predicts no futures at all.

Weights: how much each criterion matters for your job. With these weights the top family is Reactive VLAs (2.38), then World-action models (2.13).

Figure 12.1No family wins on every criterion; which one comes out on top depends on what your job weighs most, and modest changes in those weights change the answer. Seven families against eight criteria, scored low, medium or high, where high is always better. Select any cell for the reason behind it. Change the weights, or try the preset jobs, and watch the ranking. Source: the authors' judgement as of October 2026, drawn from the evidence in chapters 5 to 11; every cell's reason names its source. Scores also in rust/ch12-choosing.

Why this chapter exists

Part II described seven families one at a time. A practitioner does not meet them that way. They meet a job: a cell to automate, a robot to make useful, a research question to answer. They want to know which family to start with, what it will cost in data and compute, and how it will fail first.

This chapter adds no new models. It puts the seven families side by side, turns the comparison into a short sequence of questions, and walks three realistic jobs through it. Treat it as a map drawn on a particular date. The matrix encodes the authors' judgement as of October 2026, and some of it will be wrong within a year.

After this chapter you will be able to:

  • read the comparison matrix of figure 12.1, including which cells rest on independent evidence and which on vendor claims;
  • walk a new job through the decision flow of figure 12.2 and name a first family to try;
  • predict the first failure mode for that choice, and what to measure to catch it.

We've been trying to figure out how do we not just make a humanoid robot, but also make a humanoid robot that does useful work.

Pras Velagapudi, chief technology officer at Agility Robotics. The Wall Street Journal, 25 December 2025, as quoted in Rodney Brooks, Predictions Scorecard, 2026 January 01, 25 December 2025.

Reading the matrix

Each column of figure 12.1 is a family and each row a criterion that a deployment team cares about. The scores have three levels, so that nobody mistakes them for measurements. Higher is always better, so "cheap inference" scores high for a small, fast model and low for one that generates video.

Three rows deserve a warning. Little robot data needed mixes two things: how much action-labelled data the family needs to pretrain, and how much a new task needs. Large behavior models need few demonstrations per task once pretrained, but every task must still be demonstrated, so the row scores them low. Open weights counts what can be downloaded today, not what is promised. Evidence quality is the row to read first. It says whether a family's other scores rest on shared benchmarks and independent evaluation, or mostly on what the makers report. Reactive VLAs and large behavior models score high there. World models used as simulators and embodied reasoning models score low, because most of what is known about them comes from vendor posts and surveys.

Two things are worth trying in the figure. First, with all weights equal, reactive VLAs come out on top, mostly on maturity, open weights and evidence. Second, the ranking is fragile. Raising the weight on reasoning about consequences from 1 to 2.25 hands first place to world-action models, level with world models used as simulators, and raising the weight on cheap inference by the same amount hands it to large behavior models. Exercise 12.3 measures how much each criterion must move to change the answer. One structural fact survives every weighting: no family is dominated, meaning none is matched or beaten on every criterion by another. That is the strongest argument that the seven families are real alternatives rather than a ranking in disguise. Exercise 12.2 checks it.

The decision walk-through

Figure 12.2 turns the matrix into four questions. They are ordered from the cheapest answer to the most expensive, so that a job stops at the simplest family that fits.

yesnoyesnonoyesyesnoShort-horizon task with in-distribution objects?Contact-rich, occluded, or delayed effects?Any action-labelled robot data for this body?Edge compute only, no cloud?Reactive VLA or flow-matching action expertchapter 5World-action model or latent-prediction modelchapter 6chapter 7Latent-action pretraining from video, then adaptchapter 7chapter 8Quantised VLA on device; reasoning model optionalchapter 5chapter 10Reasoning model over a controllerchapter 10chapter 5chapter 11If precise contact state matters more than appearance,also consider geometry-first models (chapter 9).
Short-horizon task with in-distribution objects?
Figure 12.2Four questions about the job, asked in order, name a first family to try; most jobs stop at the first or second question. The decision flow from Book.md chapter 12. Answer each question, or pick a worked scenario, and follow the path to a starting point and the chapters to read.schematic, not measured Source: the authors' synthesis of Part II; the same flow is implemented in rust/ch12-choosing.

Is it a short-horizon task with objects like those in the training data? If yes, start with a reactive VLA. It is the most documented, best-supported and cheapest family to fine-tune, and chapter 5 showed that hundreds of demonstrations often suffice. Most industrial jobs that are automated today look like this.

Does it involve contact, occlusion, or effects that appear only later? If yes, a reactive policy will make locally sensible moves that ruin later steps, which is the gap chapter 6 opened with. Start with a world-action model, or a latent-prediction model if inference cost matters (chapter 7). If the contact state itself is what matters, as in insertion or cloth handling, add geometry-first models to the list (chapter 9).

Is there action-labelled robot data for this body? If no, the realistic route is to pretrain on video with latent actions and adapt with a small decoder (chapter 7), and to use a world model to multiply whatever data you collect (chapter 8).

Is the robot restricted to onboard compute? If yes, a quantised VLA on the robot, with an onboard reasoning model only if it fits. If a network is available, put an embodied reasoning model above a controller (chapter 10). The controller can be a VLA, or a large behavior model where the tasks are fixed and speed matters most (chapter 11).

The flow is deliberately simple, and real jobs blur its questions. A household task can be short-horizon in one room and long-horizon in the next. The questions are still the right ones to ask first, because each one corresponds to a property that separates the families: horizon, contact, data and compute.

Three worked scenarios

A warehouse pick-and-place cell

A fixed arm with a suction or parallel gripper over a conveyor.

Task. Pick known products from totes and place them in boxes, for eight hours a day.

Short horizon, in-distribution objects: yes.

  • Reactive VLAs. A reactive VLA fine-tuned on a few hundred demonstrations of this cell (chapter 5).

First failure to expect. New packaging: transparent, reflective or deformable items the demonstrations never covered.

What to measure. Success with confidence intervals over enough trials to see a few-point drop (chapter 11), re-measured whenever the product range changes.

A humanoid tidying a home

A two-armed humanoid with head and wrist cameras.

Task. Put away groceries and tidy a living room, from spoken instructions, in homes it has not seen.

Short horizon: no. Occlusion and delayed effects (cupboards, stacked items): yes.

  • World-action models. A world-action or latent-prediction model as the controller (chapters 6 and 7).
  • Embodied reasoning. An embodied reasoning model above it to plan and to judge each step (chapter 10).

First failure to expect. A failed step judged done, so the robot moves on and the task silently fails (figure 10.2).

What to measure. False successes per task, not only completion rate; and behaviour after a nudge or a dropped object.

A mobile inspection robot

A quadruped with a manipulator arm and a navigation stack.

Task. Walk an industrial facility, read gauges and sight glasses, and report anomalies.

Short horizon: no. Contact-rich: no. Robot action data: yes, through its own APIs. Edge only: no, the site has a network.

  • Embodied reasoning. A reasoning model orchestrating the robot's navigation and arm APIs, and reading the instruments (chapter 10).

First failure to expect. Misread instruments in poor light or at oblique angles, and stalls when the network drops.

What to measure. Reading error against a human inspector's readings, and behaviour with the network cut.

Figure 12.3The family follows from the job, and so does the first thing that will go wrong; plan the measurement for that failure before deployment. Three jobs walked through the decision flow: the robot, the task, the answers to the four questions, the families chosen, the first failure to expect and what to measure. The same scenarios are presets in figure 12.2. Source: the authors' synthesis; failure modes from chapters 5, 10 and 11, and from the Gemini Robotics-ER announcements for instrument reading.

The three cards share a pattern worth making explicit. The first failure is rarely the model's headline weakness. For the warehouse VLA it is not reasoning about consequences, which the job does not need, but distribution shift in the products. For the household humanoid it is not the controller but the judgement of when a step is done. For the inspection robot it is perception and connectivity. In each case the useful question is not "which family is best?" but "how will this family fail on this job, and how will I know?"

In code

The decision flow fits in a dozen lines. Writing it as code makes its assumptions inspectable, and easy to argue with.

rust/ch12-choosing/src/lib.rsThe decision flow of figure 12.2
#[derive(Clone, Copy, Debug, PartialEq, Eq, PartialOrd, Ord, Hash)]
pub enum Family {
    ReactiveVla,      // chapter 5
    WorldAction,      // chapter 6
    Latent,           // chapter 7
    Simulator,        // chapter 8
    Geometry,         // chapter 9
    Reasoning,        // chapter 10
    LargeBehavior,    // chapter 11
}

/// What you know about the job before choosing.
#[derive(Clone, Copy, Debug, Default)]
pub struct Job {
    pub short_horizon_in_distribution: bool,
    pub contact_occlusion_or_delayed_effects: bool,
    pub robot_action_data_for_this_body: bool,
    pub edge_compute_only: bool,
    pub precise_contact_state_matters: bool,
}

/// The decision flow of figure 12.2: a first family to try, and what to consider next.
pub fn recommend(job: &Job) -> Vec<Family> {
    use Family::*;
    if job.short_horizon_in_distribution {
        return vec![ReactiveVla];
    }
    if job.contact_occlusion_or_delayed_effects {
        let mut v = vec![WorldAction, Latent];
        if job.precise_contact_state_matters {
            v.push(Geometry);
        }
        return v;
    }
    if !job.robot_action_data_for_this_body {
        return vec![Latent, Simulator];
    }
    if job.edge_compute_only {
        return vec![ReactiveVla, Reasoning];
    }
    vec![Reasoning, ReactiveVla, LargeBehavior]
}

Exercises

The stubs are in rust/ch12-choosing/src/exercises.rs. This chapter is equation-free; the exercises are about weighing, comparing and stress-testing a judgement.

Exercise 12.1

A weighted score

Implement weighted_score and best. Then choose weights for a job you know, and check whether the answer agrees with the decision flow. If it does not, which of the two do you trust, and why?

Check your answer from the rust/ folder:

cargo test -p ch12-choosing --test exercises ex12_1
Hint 1

Multiply each score by its weight, add them up, and divide by the total weight.

Hint 2

Break ties in favour of the family that comes first in ALL.

Exercise 12.2

Every family wins on something

Implement dominated. The book's matrix has no dominated family. Find the smallest change to one cell that makes one family dominated, and say whether you would believe the matrix more or less after that change.

Check your answer from the rust/ folder:

cargo test -p ch12-choosing --test exercises ex12_2
Hint 1

A family is dominated if another is at least as good everywhere and strictly better somewhere.

Hint 2

Compare every pair; a family cannot dominate itself.

Exercise 12.3

How fragile is the ranking?

Implement flip. For each criterion, find how much its weight must rise from equal weights to change the winner. Which criteria can never dethrone the leader, and what does that tell you about why the leader leads?

Check your answer from the rust/ folder:

cargo test -p ch12-choosing --test exercises ex12_3
Hint 1

Increase one weight in steps of 0.25 and recompute the winner each time.

Hint 2

Stop at the first change of winner, or give up at an increase of 10.

What we are not sure about

Whether the families stay separate. Chapters 5 to 11 each ended by noting that the family's boundary is blurring: VLAs adding future prediction, reasoning moving into controllers, imitation primitives everywhere. A matrix with seven columns may become a matrix of design choices within one model.

Whether the criteria are the right ones. The industrial survey uses 149 criteria; this matrix uses eight. Safety, which this book covers in chapters 15 and 17, is absent as a row because no family has enough evidence to score it.

How quickly the scores age. The evidence-quality and maturity rows are the most likely to change, in both directions, as independent evaluations of world-action models and reasoning models appear. Check the date on this chapter before relying on it.

Further reading

References

6 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.

  1. D. Kube, S. Hadwiger, T. Meisen Robotic Foundation Models for Industrial Control: A Comprehensive Survey and Readiness Assessment Framework. arXiv 2603.06749 v1, 2026-03-06.
  2. RoboCloud Hub NVIDIA GR00T N2 Explained; VLA Tutorial 2026. robocloudhub.tech, 2026.
  3. Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.
  4. A. Psiris, V. Argyriou, E. K. Markakis and 5 others Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions. Transactions on Machine Learning Research (TMLR), 07/2026. arXiv 2604.15395 v3, 2026-04-16.
  5. TRI LBM Team, J. Barreiros, A. Beaulieu and 79 others A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation. arXiv 2507.05331 v1, 2025-07-07.
  6. Rodney Brooks Predictions Scorecard, 2026 January 01. rodneybrooks.com, 2026-01-01.