Skip to content

Ch.10Part II · The seven familiesEmbodied reasoning

Embodied reasoning models

VLM-level models for spatial understanding, task decomposition and success detection that sit above a controller from another family.

Math
L0

No equations

Sources
10

3 papers from 2026

Figures
3

Most are interactive

Exercises
3

Rust, with tests

Last checked
9 Oct 2026

The field moves monthly

Papers cited, by half-year of first arXiv version

Task: put the fruit in the bowl✓ 1. pick up the apple✓ 2. place it in the bowl› 3. pick up the banana 4. place it in the bowlrobot: motors and camerasembodied reasoning modelplans, calls tools, judges progressVLAnavigation APIsearch, user functionsaction chunksvideoprogress, judged from video20-40%

A separate reasoning model plans, sends one subgoal at a time to a controller it calls as a tool, watches the video stream, and decides when to move on. Gemini Robotics ER 2 (July 2026) works this way, calling VLAs, navigation APIs or search through the Live API.

What crosses: down: a subgoal in words; up: video, from which the reasoner judges progress.

Figure 10.1An embodied reasoning model does not move the robot; it decides what the controller should do next and whether it worked, and the three ways of wiring it in differ in what crosses between thinking and acting. Three places to put the reasoning: a separate model over controllers, a reasoning backbone inside a VLA, or one model that thinks and then acts. Choose an arrangement and watch a four-step task.schematic, not measured Source: Gemini Robotics ER 2 announcement (Google, 30 July 2026); NVIDIA Isaac-GR00T repository (GR00T N1.7); π0.5 (arXiv 2504.16054); Gemini Robotics 1.5 (2510.03342); MolmoAct2 (2605.02881).

Why this chapter exists

Every family so far ends in motor commands. This one does not. An embodied reasoning model is a vision-language model trained to understand space, plan multi-step tasks and judge progress, which then hands the motion to a controller from one of the other families. It is the part of the robot that knows the cup is behind the kettle, that "tidy the desk" means four separate pickups, and, above all, that the third pickup failed and needs another try.

That last ability matters more than it sounds. A controller that has tried a step has no reliable way of knowing whether it succeeded, so it either moves on regardless or never moves on at all. A reasoning model that can look at the scene and say "done" or "not yet" is what turns a sequence of attempts into a task. Google DeepMind's announcement of Gemini Robotics-ER 1.6 calls success detection "a cornerstone of autonomy" (Google DeepMind 2026).

After this chapter you will be able to:

  • name the four jobs an embodied reasoning model does, and say which of them a controller cannot do for itself;
  • explain with figure 10.2 why a judge that is right most of the time still matters enormously over a ten-step task, and why an optimistic judge is worse than a cautious one;
  • read an embodied reasoning announcement critically, knowing that almost all evidence so far comes from the model makers.

In robotics, knowing when a task is finished is just as important as knowing how to start it.

Laura Graesser and Peng Xu, Google DeepMind, authors of the Gemini Robotics-ER 1.6 announcement. Gemini Robotics-ER 1.6: Powering real-world robotics tasks through enhanced embodied reasoning, Google DeepMind blog, 14 April 2026.

Four jobs

Spatial understanding. Where things are, how many there are, which is the smallest, where to grasp. The Gemini Robotics-ER models express much of this by pointing: returning image coordinates for "every object small enough to fit inside the blue cup". Pointing also serves as an intermediate step for counting and for metric estimates (Google DeepMind 2026). The first Gemini Robotics-ER, introduced alongside Gemini Robotics in March 2025, already covered object detection, pointing, trajectories and grasps (Gemini Robotics Team et al. 2025, arXiv 2503.20020).

Task decomposition. Turning "put the fruit in the bowl" into a sequence of steps a controller can execute, and changing the plan when something unexpected happens.

Progress and success detection. Watching the scene and deciding whether the current step is done. Gemini Robotics ER 2 classifies each video frame into five levels of progress, from 0-20 percent to 80-100 percent, so the robot can "retry failed steps without restarting an entire workflow" (Google DeepMind 2026). Multiple cameras make this harder and more useful: ER 1.6's main advance was reasoning across an overhead and a wrist-mounted view at once (Google DeepMind 2026).

Orchestration. Calling the tools that do the work. ER 2 treats VLAs, navigation APIs, Google Search and any user-defined function as tools, and streams video, audio or text through a low-latency endpoint so it can "think" about the next step while the robot is still acting (Google DeepMind 2026).

None of the four is motor control. That is the defining line of the family, and the reason it sits on its own in the book's family grid: what it predicts is a plan in language, not an action.

Where the reasoning sits

Figure 10.1 shows three ways to connect reasoning to motion, and they are converging.

A separate reasoner over controllers. The cleanest split. The reasoning model runs in the cloud or on a large onboard computer, sends subgoals in words to whatever controller fits the step, and watches the result. Gemini Robotics ER 2, which became publicly available to developers through the Gemini API on 30 July 2026, is designed for exactly this (Google DeepMind 2026). The advantage is that the reasoner is embodiment-agnostic: the same model can orchestrate a robot arm's VLA, a quadruped's navigation API and a search engine.

Reasoning inside the VLA. The reasoning model becomes the VLA's backbone. NVIDIA's Cosmos-Reason models were trained for physical common sense and embodied decisions (NVIDIA et al. 2025, arXiv 2503.15558), and GR00T N1.7 builds its 3-billion-parameter VLA on Cosmos-Reason2-2B (NVIDIA 2026). π0.5 goes a step further and predicts a semantic subtask before the actions for it, within one model (Physical Intelligence et al. 2025, arXiv 2504.16054). Nothing readable crosses between thinking and acting, which is faster and harder to inspect.

Think, then act. One model writes its reasoning out and then acts. Gemini Robotics 1.5 "interleaves actions with a multi-level internal reasoning process in natural language" (Gemini Robotics Team et al. 2025, arXiv 2510.03342), and MolmoAct2 is built as an action reasoning model on a backbone specialised for embodied reasoning (Fang et al. 2026, arXiv 2605.02881). The reasoning is readable, which helps debugging, and every thought costs time.

Why success detection matters so much

The value of a judge is easiest to see in a long task. Figure 10.2 runs one: a controller that completes each step on 70 percent of attempts, a judge that looks after every attempt, and up to three attempts before the robot stops and asks for help.

One sampled run of 10 steps; the bars below are exact probabilities.
judged not done, attempts left: retryActcontroller tries the stepObservecameras over timeJudgereasoning model: done?Next stepjudged doneAsk for helpout of attemptsstepsno judge, one try97%this judge59%14%27%exact probabilities for the whole 10-step task
One step: completed
94.8%
One step: failed without knowing
2.0%
Attempts per step, on average
1.46what checking costs in time
Figure 10.2Over a long task, a judge that catches failures turns a near-certain failure into a likely success; a judge that is too ready to say 'done' gives most of that back as failures nobody notices. The success-detection loop as a state machine: act, observe, judge, then move on, retry or ask for help. Run a sampled ten-step task, and change how good the controller and the judge are; the bars compare the exact outcomes with a robot that never checks.toy simulation Source: this book's toy, in rust/ch10-reasoning; the loop follows the Gemini Robotics-ER 1.6 and ER 2 announcements (Google DeepMind, April and July 2026).

The numbers are stark. A robot that never checks, and simply moves on after one try, completes a ten-step task about 3 percent of the time, and in the other 97 percent it fails without knowing. With a judge that wrongly says "done" 5 percent of the time and wrongly says "not done" 10 percent of the time, the same controller completes the task about 59 percent of the time, asks for help 27 percent of the time, and fails silently 14 percent of the time.

The two kinds of judging error are not equally bad. Raise the chance of a hallucinated success to 30 percent and silent failures jump to about 61 percent: the robot reports a tidy desk that is not tidy. Raise the chance of a wrong retry to 40 percent instead and silent failures stay low, but the robot asks for help on most tasks and spends more attempts per step. A robot that asks too often is annoying. A robot that is confidently wrong is dangerous, and that asymmetry is the strongest argument for evaluating embodied reasoning models on false successes specifically. Exercises 10.1 to 10.3 compute these numbers exactly.

Representative models

20252026Reasoning modelsInside controllersIndependent evaluationCosmos-Reason1Gemini Robotics-ERER 1.5Cosmos-Reason2ER 1.6ER 2π0.5Gemini Robotics 1.5GR00T N1.7MolmoAct2RoboRewardBeTTER

BeTTER 2026-04

A diagnostic benchmark that finds state-of-the-art VLAs fail catastrophically when layouts shift or tasks require temporal extrapolation.

Figure 10.3The family moved from research reports to public developer APIs in about eighteen months; the independent checks started arriving only in 2026. Embodied reasoning models, controllers with reasoning built in, and the first independent evaluations, 2025 to 2026. Select a milestone for what it introduced. Source: each model's report or announcement as cited in this chapter; RoboReward (arXiv 2601.00675); BeTTER (2604.18000). Capabilities as stated by the makers.

Why it exists, good at, bad at

Why it exists. Knowing when a task is finished is what lets a robot decide whether to retry or move on, and controllers are poor at it by themselves. Long-horizon tasks also need planning that a reactive controller does not do.

What it is good at. It reuses frontier VLM capability directly, so it improves whenever the underlying VLM improves. It is fast to iterate on, because changing a prompt or a tool list needs no robot data. And it is largely embodiment-agnostic: the same reasoning model can sit above very different robots.

What it is bad at. Latency and, for the largest models, dependence on the cloud: ER 2's low-latency streaming endpoint exists because "high-level reasoning depends on execution speed" (Google DeepMind 2026). It has no direct control, so it is only as good as the controllers it calls. And its failure modes are those of the underlying VLM: hallucinated success, brittle counting, misreading a cluttered scene.

In code

The crate's loop is the state machine of figure 10.2, with a reproducible random generator so traces can be compared.

rust/ch10-reasoning/src/lib.rsA judge, an attempt, and one step of the success-detection loop
/// One step of a task, as the reasoning model sees it.
#[derive(Clone, Copy, Debug, PartialEq)]
pub struct Judge {
    /// Chance that one attempt by the controller completes the step.
    pub p_success: f64,
    /// Chance the judge says "done" when the step is not done: a hallucinated success.
    pub false_done: f64,
    /// Chance the judge says "not done" when the step is done: a wasted retry.
    pub false_retry: f64,
    /// Attempts allowed before the robot stops and asks a person for help.
    pub max_attempts: u32,
}

impl Judge {
    /// No success detection at all: try once and move on.
    pub fn blind(p_success: f64) -> Judge {
        Judge { p_success, false_done: 1.0, false_retry: 0.0, max_attempts: 1 }
    }
}

/// How a step ends.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum StepEnd {
    /// Judged done, and really done.
    Done,
    /// Judged done, but not done: the robot moves on without knowing.
    SilentFailure,
    /// Out of attempts: the robot knows it failed and asks for help.
    AskedForHelp,
}

/// What happened in one attempt, for drawing a trace.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub struct Attempt {
    pub done_after: bool,
    pub judged_done: bool,
}

/// A small deterministic generator (SplitMix64) so traces are reproducible.
pub struct Rng(u64);

impl Rng {
    pub fn new(seed: u64) -> Rng {
        Rng(seed)
    }
    pub fn next_f64(&mut self) -> f64 {
        self.0 = self.0.wrapping_add(0x9E37_79B9_7F4A_7C15);
        let mut z = self.0;
        z = (z ^ (z >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9);
        z = (z ^ (z >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB);
        ((z ^ (z >> 31)) >> 11) as f64 / (1u64 << 53) as f64
    }
}

/// Run one step: attempt, judge, and retry until judged done or out of attempts. A step that is
/// already done stays done if the controller is asked to do it again.
pub fn run_step(j: &Judge, rng: &mut Rng) -> (StepEnd, Vec<Attempt>) {
    let mut done = false;
    let mut log = Vec::new();
    for _ in 0..j.max_attempts {
        if !done && rng.next_f64() < j.p_success {
            done = true;
        }
        let says_done = if done { rng.next_f64() >= j.false_retry } else { rng.next_f64() < j.false_done };
        log.push(Attempt { done_after: done, judged_done: says_done });
        if says_done {
            return (if done { StepEnd::Done } else { StepEnd::SilentFailure }, log);
        }
    }
    (StepEnd::AskedForHelp, log)
}

Exercises

The stubs are in rust/ch10-reasoning/src/exercises.rs. This chapter is equation-free; the exercises ask for probabilities, which you can compute by following the state machine one attempt at a time.

Exercise 10.1

One step, exactly

Implement step_outcomes, the exact chances that a step ends done, silently failed, or with a request for help. The test compares your answer with 200,000 sampled steps. Then find the attempts limit beyond which more attempts stop helping, for the default judge.

Check your answer from the rust/ folder:

cargo test -p ch10-reasoning --test exercises ex10_1
Hint 1

Keep two numbers: the chance the step is done and the judge has not yet stopped, and the chance it is not done and the judge has not yet stopped.

Hint 2

After each attempt, move some of the not-done chance to done, then let the judge stop part of each.

Exercise 10.2

A whole task

Implement task_outcomes for a task of many steps. Use it to find how good a controller must be, per attempt, for a robot without any judge to complete a ten-step task half the time.

Check your answer from the rust/ folder:

cargo test -p ch10-reasoning --test exercises ex10_2
Hint 1

A task completes only if every step completes.

Hint 2

The robot stops at the first request for help, so later steps never run.

Exercise 10.3

What checking costs

Implement expected_attempts. A judge that retries finished steps wastes time; a judge that accepts unfinished ones wastes trust. If each attempt takes 20 seconds, how much longer does the default judge make a ten-step task than no judge, and what does it buy?

Check your answer from the rust/ folder:

cargo test -p ch10-reasoning --test exercises ex10_3
Hint

Add up, for each attempt, the chance that the step is still going when the attempt starts.

What we are not sure about

How good the judges really are. The frontier models' success-detection and progress numbers are reported by their makers on their own evaluations. The one large independent benchmark of VLMs as judges, RoboReward, found no model reliable across tasks. Whether today's ER models would pass a third-party test on false successes is unknown.

Whether the separate reasoner survives. Reasoning is moving into controllers: GR00T N1.7's backbone, π0.5's subtasks, Gemini Robotics 1.5's thoughts. It is possible that within a few years the separate reasoning model becomes a component of every VLA rather than a family of its own.

What the latency costs in practice. Reasoning over video in the cloud adds delay at exactly the moment a robot needs to decide whether to retry. Vendors describe streaming designs that hide it; there are no independent measurements of end-to-end delay in real deployments.

Further reading

References

10 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.

  1. Google DeepMind Gemini Robotics-ER 1.6. blog.google, 2026-04-14.
  2. Gemini Robotics Team, S. Abeyruwan, J. Ainslie and 115 others Gemini Robotics: Bringing AI into the Physical World. arXiv 2503.20020 v1, 2025-03-25.
  3. Google DeepMind Introducing Gemini Robotics ER 2. blog.google, 2026-07-30.
  4. NVIDIA, A. Azzolini, J. Bai and 50 others Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning. arXiv 2503.15558 v3, 2025-03-18.
  5. NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
  6. Physical Intelligence, K. Black, N. Brown and 33 others π0.5: a Vision-Language-Action Model with Open-World Generalization. arXiv 2504.16054 v1, 2025-04-22.
  7. Gemini Robotics Team, A. Abdolmaleki, S. Abeyruwan and 169 others Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer. arXiv 2510.03342 v3, 2025-10-02.
  8. H. Fang, J. Duan, D. Clay and 26 others MolmoAct2: Action Reasoning Models for Real-world Deployment. arXiv 2605.02881 v2, 2026-05-04.
  9. T. Lee, A. Wagenmaker, K. Pertsch and 3 others RoboReward: General-Purpose Vision-Language Reward Models for Robotics. arXiv 2601.00675 v2, 2026-01-02.
  10. H. Xu, S. Zheng, H. Luo and 3 others Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models. arXiv 2604.18000 v1, 2026-04-20.