Ch.10Part II · The seven familiesEmbodied reasoning
Embodied reasoning models
VLM-level models for spatial understanding, task decomposition and success detection that sit above a controller from another family.
- Math
- L0
- Sources
- 10
- Figures
- 3
- Exercises
- 3
- Last checked
- 9 Oct 2026
No equations
3 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
A separate reasoning model plans, sends one subgoal at a time to a controller it calls as a tool, watches the video stream, and decides when to move on. Gemini Robotics ER 2 (July 2026) works this way, calling VLAs, navigation APIs or search through the Live API.
What crosses: down: a subgoal in words; up: video, from which the reasoner judges progress.
Why this chapter exists
Every family so far ends in motor commands. This one does not. An embodied reasoning model is a vision-language model trained to understand space, plan multi-step tasks and judge progress, which then hands the motion to a controller from one of the other families. It is the part of the robot that knows the cup is behind the kettle, that "tidy the desk" means four separate pickups, and, above all, that the third pickup failed and needs another try.
That last ability matters more than it sounds. A controller that has tried a step has no reliable way of knowing whether it succeeded, so it either moves on regardless or never moves on at all. A reasoning model that can look at the scene and say "done" or "not yet" is what turns a sequence of attempts into a task. Google DeepMind's announcement of Gemini Robotics-ER 1.6 calls success detection "a cornerstone of autonomy" (Google DeepMind 2026).
After this chapter you will be able to:
- name the four jobs an embodied reasoning model does, and say which of them a controller cannot do for itself;
- explain with figure 10.2 why a judge that is right most of the time still matters enormously over a ten-step task, and why an optimistic judge is worse than a cautious one;
- read an embodied reasoning announcement critically, knowing that almost all evidence so far comes from the model makers.
In robotics, knowing when a task is finished is just as important as knowing how to start it.
Four jobs
Spatial understanding. Where things are, how many there are, which is the smallest, where to grasp. The Gemini Robotics-ER models express much of this by pointing: returning image coordinates for "every object small enough to fit inside the blue cup". Pointing also serves as an intermediate step for counting and for metric estimates (Google DeepMind 2026). The first Gemini Robotics-ER, introduced alongside Gemini Robotics in March 2025, already covered object detection, pointing, trajectories and grasps (Gemini Robotics Team et al. 2025, arXiv 2503.20020).
Task decomposition. Turning "put the fruit in the bowl" into a sequence of steps a controller can execute, and changing the plan when something unexpected happens.
Progress and success detection. Watching the scene and deciding whether the current step is done. Gemini Robotics ER 2 classifies each video frame into five levels of progress, from 0-20 percent to 80-100 percent, so the robot can "retry failed steps without restarting an entire workflow" (Google DeepMind 2026). Multiple cameras make this harder and more useful: ER 1.6's main advance was reasoning across an overhead and a wrist-mounted view at once (Google DeepMind 2026).
Orchestration. Calling the tools that do the work. ER 2 treats VLAs, navigation APIs, Google Search and any user-defined function as tools, and streams video, audio or text through a low-latency endpoint so it can "think" about the next step while the robot is still acting (Google DeepMind 2026).
None of the four is motor control. That is the defining line of the family, and the reason it sits on its own in the book's family grid: what it predicts is a plan in language, not an action.
Where the reasoning sits
Figure 10.1 shows three ways to connect reasoning to motion, and they are converging.
A separate reasoner over controllers. The cleanest split. The reasoning model runs in the cloud or on a large onboard computer, sends subgoals in words to whatever controller fits the step, and watches the result. Gemini Robotics ER 2, which became publicly available to developers through the Gemini API on 30 July 2026, is designed for exactly this (Google DeepMind 2026). The advantage is that the reasoner is embodiment-agnostic: the same model can orchestrate a robot arm's VLA, a quadruped's navigation API and a search engine.
Reasoning inside the VLA. The reasoning model becomes the VLA's backbone. NVIDIA's Cosmos-Reason models were trained for physical common sense and embodied decisions (NVIDIA et al. 2025, arXiv 2503.15558), and GR00T N1.7 builds its 3-billion-parameter VLA on Cosmos-Reason2-2B (NVIDIA 2026). π0.5 goes a step further and predicts a semantic subtask before the actions for it, within one model (Physical Intelligence et al. 2025, arXiv 2504.16054). Nothing readable crosses between thinking and acting, which is faster and harder to inspect.
Think, then act. One model writes its reasoning out and then acts. Gemini Robotics 1.5 "interleaves actions with a multi-level internal reasoning process in natural language" (Gemini Robotics Team et al. 2025, arXiv 2510.03342), and MolmoAct2 is built as an action reasoning model on a backbone specialised for embodied reasoning (Fang et al. 2026, arXiv 2605.02881). The reasoning is readable, which helps debugging, and every thought costs time.
Why success detection matters so much
The value of a judge is easiest to see in a long task. Figure 10.2 runs one: a controller that completes each step on 70 percent of attempts, a judge that looks after every attempt, and up to three attempts before the robot stops and asks for help.
- One step: completed
- 94.8%
- One step: failed without knowing
- 2.0%
- Attempts per step, on average
- 1.46what checking costs in time
The numbers are stark. A robot that never checks, and simply moves on after one try, completes a ten-step task about 3 percent of the time, and in the other 97 percent it fails without knowing. With a judge that wrongly says "done" 5 percent of the time and wrongly says "not done" 10 percent of the time, the same controller completes the task about 59 percent of the time, asks for help 27 percent of the time, and fails silently 14 percent of the time.
The two kinds of judging error are not equally bad. Raise the chance of a hallucinated success to 30 percent and silent failures jump to about 61 percent: the robot reports a tidy desk that is not tidy. Raise the chance of a wrong retry to 40 percent instead and silent failures stay low, but the robot asks for help on most tasks and spends more attempts per step. A robot that asks too often is annoying. A robot that is confidently wrong is dangerous, and that asymmetry is the strongest argument for evaluating embodied reasoning models on false successes specifically. Exercises 10.1 to 10.3 compute these numbers exactly.
Representative models
BeTTER 2026-04
A diagnostic benchmark that finds state-of-the-art VLAs fail catastrophically when layouts shift or tasks require temporal extrapolation.
Why it exists, good at, bad at
Why it exists. Knowing when a task is finished is what lets a robot decide whether to retry or move on, and controllers are poor at it by themselves. Long-horizon tasks also need planning that a reactive controller does not do.
What it is good at. It reuses frontier VLM capability directly, so it improves whenever the underlying VLM improves. It is fast to iterate on, because changing a prompt or a tool list needs no robot data. And it is largely embodiment-agnostic: the same reasoning model can sit above very different robots.
What it is bad at. Latency and, for the largest models, dependence on the cloud: ER 2's low-latency streaming endpoint exists because "high-level reasoning depends on execution speed" (Google DeepMind 2026). It has no direct control, so it is only as good as the controllers it calls. And its failure modes are those of the underlying VLM: hallucinated success, brittle counting, misreading a cluttered scene.
In code
The crate's loop is the state machine of figure 10.2, with a reproducible random generator so traces can be compared.
/// One step of a task, as the reasoning model sees it.
#[derive(Clone, Copy, Debug, PartialEq)]
pub struct Judge {
/// Chance that one attempt by the controller completes the step.
pub p_success: f64,
/// Chance the judge says "done" when the step is not done: a hallucinated success.
pub false_done: f64,
/// Chance the judge says "not done" when the step is done: a wasted retry.
pub false_retry: f64,
/// Attempts allowed before the robot stops and asks a person for help.
pub max_attempts: u32,
}
impl Judge {
/// No success detection at all: try once and move on.
pub fn blind(p_success: f64) -> Judge {
Judge { p_success, false_done: 1.0, false_retry: 0.0, max_attempts: 1 }
}
}
/// How a step ends.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum StepEnd {
/// Judged done, and really done.
Done,
/// Judged done, but not done: the robot moves on without knowing.
SilentFailure,
/// Out of attempts: the robot knows it failed and asks for help.
AskedForHelp,
}
/// What happened in one attempt, for drawing a trace.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub struct Attempt {
pub done_after: bool,
pub judged_done: bool,
}
/// A small deterministic generator (SplitMix64) so traces are reproducible.
pub struct Rng(u64);
impl Rng {
pub fn new(seed: u64) -> Rng {
Rng(seed)
}
pub fn next_f64(&mut self) -> f64 {
self.0 = self.0.wrapping_add(0x9E37_79B9_7F4A_7C15);
let mut z = self.0;
z = (z ^ (z >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9);
z = (z ^ (z >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB);
((z ^ (z >> 31)) >> 11) as f64 / (1u64 << 53) as f64
}
}
/// Run one step: attempt, judge, and retry until judged done or out of attempts. A step that is
/// already done stays done if the controller is asked to do it again.
pub fn run_step(j: &Judge, rng: &mut Rng) -> (StepEnd, Vec<Attempt>) {
let mut done = false;
let mut log = Vec::new();
for _ in 0..j.max_attempts {
if !done && rng.next_f64() < j.p_success {
done = true;
}
let says_done = if done { rng.next_f64() >= j.false_retry } else { rng.next_f64() < j.false_done };
log.push(Attempt { done_after: done, judged_done: says_done });
if says_done {
return (if done { StepEnd::Done } else { StepEnd::SilentFailure }, log);
}
}
(StepEnd::AskedForHelp, log)
}Exercises
The stubs are in rust/ch10-reasoning/src/exercises.rs. This chapter is equation-free; the exercises ask for probabilities, which you can compute by following the state machine one attempt at a time.
Exercise 10.1
One step, exactly
Implement step_outcomes, the exact chances that a step ends done, silently failed, or with a request for help. The test compares your answer with 200,000 sampled steps. Then find the attempts limit beyond which more attempts stop helping, for the default judge.
Check your answer from the rust/ folder:
cargo test -p ch10-reasoning --test exercises ex10_1Hint 1
Keep two numbers: the chance the step is done and the judge has not yet stopped, and the chance it is not done and the judge has not yet stopped.
Hint 2
After each attempt, move some of the not-done chance to done, then let the judge stop part of each.
Exercise 10.2
A whole task
Implement task_outcomes for a task of many steps. Use it to find how good a controller must be, per attempt, for a robot without any judge to complete a ten-step task half the time.
Check your answer from the rust/ folder:
cargo test -p ch10-reasoning --test exercises ex10_2Hint 1
A task completes only if every step completes.
Hint 2
The robot stops at the first request for help, so later steps never run.
Exercise 10.3
What checking costs
Implement expected_attempts. A judge that retries finished steps wastes time; a judge that accepts unfinished ones wastes trust. If each attempt takes 20 seconds, how much longer does the default judge make a ten-step task than no judge, and what does it buy?
Check your answer from the rust/ folder:
cargo test -p ch10-reasoning --test exercises ex10_3Hint
Add up, for each attempt, the chance that the step is still going when the attempt starts.
What we are not sure about
How good the judges really are. The frontier models' success-detection and progress numbers are reported by their makers on their own evaluations. The one large independent benchmark of VLMs as judges, RoboReward, found no model reliable across tasks. Whether today's ER models would pass a third-party test on false successes is unknown.
Whether the separate reasoner survives. Reasoning is moving into controllers: GR00T N1.7's backbone, π0.5's subtasks, Gemini Robotics 1.5's thoughts. It is possible that within a few years the separate reasoning model becomes a component of every VLA rather than a family of its own.
What the latency costs in practice. Reasoning over video in the cloud adds delay at exactly the moment a robot needs to decide whether to retry. Vendors describe streaming designs that hide it; there are no independent measurements of end-to-end delay in real deployments.
Further reading
- The Gemini Robotics 1.5 report, for the division of labour between a reasoning model and a VLA (Gemini Robotics Team et al. 2025, arXiv 2510.03342).
- The Gemini Robotics-ER 1.6 post, for success detection and multi-view reasoning (Google DeepMind 2026).
- The Gemini Robotics ER 2 announcement and model card, for progress tracking and orchestration (Google DeepMind 2026).
- Cosmos-Reason1, for an open reasoning model and its ontology of physical common sense (NVIDIA et al. 2025, arXiv 2503.15558).
- RoboReward, for how well VLMs judge robot episodes when someone else checks (Lee et al. 2026, arXiv 2601.00675).
- BeTTER, for what static benchmarks hide about embodied reasoning (Xu et al. 2026, arXiv 2604.18000).
References
10 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- Google DeepMind Gemini Robotics-ER 1.6. blog.google, 2026-04-14.
- Gemini Robotics Team, S. Abeyruwan, J. Ainslie and 115 others Gemini Robotics: Bringing AI into the Physical World. arXiv 2503.20020 v1, 2025-03-25.
- Google DeepMind Introducing Gemini Robotics ER 2. blog.google, 2026-07-30.
- NVIDIA, A. Azzolini, J. Bai and 50 others Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning. arXiv 2503.15558 v3, 2025-03-18.
- NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
- Physical Intelligence, K. Black, N. Brown and 33 others π0.5: a Vision-Language-Action Model with Open-World Generalization. arXiv 2504.16054 v1, 2025-04-22.
- Gemini Robotics Team, A. Abdolmaleki, S. Abeyruwan and 169 others Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer. arXiv 2510.03342 v3, 2025-10-02.
- H. Fang, J. Duan, D. Clay and 26 others MolmoAct2: Action Reasoning Models for Real-world Deployment. arXiv 2605.02881 v2, 2026-05-04.
- T. Lee, A. Wagenmaker, K. Pertsch and 3 others RoboReward: General-Purpose Vision-Language Reward Models for Robotics. arXiv 2601.00675 v2, 2026-01-02.
- H. Xu, S. Zheng, H. Luo and 3 others Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models. arXiv 2604.18000 v1, 2026-04-20.