Skip to content

Ch.16Part III · Cross-cutting engineering

Applications by domain

Manipulation, humanoids, navigation, driving, industrial and service robots: which family dominates and what shipped.

Math
L0

No equations

Sources
30

4 papers from 2026

Figures
5

Most are interactive

Exercises
3

Rust, with tests

Last checked
9 Oct 2026

The field moves monthly

Papers cited, by half-year of first arXiv version

Tabletop and bimanualHumanoidsMobile manipulation and navigationAutonomous drivingIndustrial and logisticsMobile service robots
Reactive VLAs
World-action models
Latent prediction
World simulators
Geometry-first
Embodied reasoning
Large behavior models

World simulators in autonomous driving

GAIA-1, GAIA-2 and DriveDreamer generate driving scenarios, common and rare, for training and testing.

Figure 16.1Each domain leans on different families: VLAs dominate manipulation and industry, world models dominate driving, and reasoning models lead in service robots. Six application domains against the seven families. Filled cells are a main approach in this book's sources, outlined cells are used or researched, dashes are absent from the sources. Select a cell for the systems behind it. Source: the authors' reading of the sources cited in this chapter and in chapters 5 to 11; deployment counts are industry-survey or vendor claims.

Why this chapter exists

Part II described the families as ideas. This chapter describes where they are used. Each domain imposes its own constraints: a humanoid must not fall, a car must handle the rare event, a warehouse cell must pick for eight hours without a supervisor, a service robot must understand a person. Those constraints decide which family dominates, and each domain has a failure mode of its own that a general benchmark does not show.

One section per domain follows, each with the family that dominates, a named system, the failure mode to expect and, where it adds something, a figure. Deployment counts and adopter lists in this chapter come from vendors and industry surveys, and are named as claims.

After this chapter you will be able to:

  • say which families dominate in each of six domains, and why the domain's constraints favour them;
  • name a representative system per domain and the first failure mode to plan for;
  • explain with figure 16.3 why look-ahead planning is only as good as the world model doing the looking.

The future of humanoids is about adaptability and learning.

Bernt Børnich, chief executive officer of 1X Technologies. NVIDIA Announces Isaac GR00T N1, NVIDIA Newsroom, 18 March 2025.

Tabletop and bimanual manipulation

This is the home turf of reactive VLAs (chapter 5) and the large behavior models they grew from (chapter 11). The π0 line, GR00T, OpenVLA and their open relatives are developed and benchmarked here, and Diffusion Policy and ACT remain the workhorses of bimanual teleoperation research (Physical Intelligence et al. 2026, arXiv 2604.15483; NVIDIA 2026; Chi et al. 2023, arXiv 2303.04137; Zhao et al. 2023, arXiv 2304.13705). The frontier claims are about learning quickly: Generalist AI reports GEN-1.5 learning a task from a single 3 to 12 second demonstration, at 59 percent average success across 10 tasksvendor claim (Generalist AI 2026).

The failure mode to expect is distribution shift that looks small to a person. LIBERO-Plus showed success falling from 95 percent to below 30 percent under modest changes of camera viewpoint and initial robot state (Fei et al. 2025, arXiv 2510.13626). Figures 5.1 and 11.3 cover the mechanics of this domain; this chapter does not repeat them.

Humanoids and whole-body control

Humanoids stack three layers: something that reasons about the task, something that turns it into motion, and something that keeps a two-legged body balanced while it moves. Figure 16.2 shows three ways the layers are arranged today.

reasoning backboneCosmos-Reason2-2B inside the VLAwith the VLAVLAemits compact latent action tokenschunk ratewhole-body controllerSONIC turns tokens into joint commandscontroller ratehighlighted: what the stack drives

One policy produces coordinated manipulation and locomotion: the VLA decides what the body should do, the learned controller how.

Figure 16.2Every humanoid stack separates deciding what to do from keeping the body upright; the systems differ in where the line falls and what crosses it. Three humanoid control stacks: GR00T N1.7 with its SONIC whole-body controller, Figure's Helix, and a reasoning model orchestrating a humanoid's own controllers. Select one.schematic, not measured Source: NVIDIA Isaac-GR00T repository (GR00T N1.7); SONIC (arXiv 2511.07820); Figure AI's Helix post (2025); Gemini Robotics ER 2 announcement (Google, 30 July 2026).

GR00T N1.7 drives a Unitree G1 through SONIC: the VLA emits compact latent action tokens and a learned whole-body controller turns them into joint commands for legs, arms and hands (NVIDIA 2026; Luo et al. 2025, arXiv 2511.07820). Helix splits a 7-billion-parameter reasoner at 7 to 9 Hz from an 80-million-parameter controller at 200 Hz (Figure AI 2025). Gemini Robotics ER 2 orchestrates a humanoid's own controllers as tools (Google DeepMind 2026). NVIDIA previewed GR00T N2 at GTC in March 2026 and named AGIBOT, Humanoid, LG Electronics, NEURA Robotics and Noble Machines as adopting GR00T models for industrial humanoidsvendor claim (NVIDIA Newsroom 2026). For research, NVIDIA's Isaac GR00T reference humanoid combines a Unitree H2 Plus with Sharpa five-fingered hands and Jetson Thor compute, with Ai2, ETH Zurich, Stanford and UC San Diego among the first users (NVIDIA Newsroom 2026).

The failure mode to expect is physical, and it collides with how industrial safety works. Every industrial robot has a safety stop that cuts power to the motors, but cutting power to a balancing robot can make it fall. Rodney Brooks recalls that at Rethink Robotics his team kept the arms of Baxter and Sawyer from dropping on power cut-off with a circuit that braked the motors without active power, and adds: "Perhaps there are similar possible solutions for humanoid robots and falling, but they need to be invented yet" (Rodney Brooks 2026).

Mobile manipulation and navigation

Navigation was one of the first places world models earned their keep, because a robot that can imagine where a path leads can compare paths before taking one. Navigation World Models, a 1-billion-parameter video diffusion model trained on egocentric video from humans and robots, plans by simulating candidate trajectories and evaluating whether they reach the goal, either from scratch or by ranking trajectories proposed by another policy (Bar et al. 2024, arXiv 2412.03572). X-Mobility learns a latent world model separately from its action policy, so it can train on data with or without expert actions (Liu et al. 2024, arXiv 2410.17491). InternVLA-N1, published as DualVLN, puts a VLM that predicts mid-term waypoints over a fast diffusion policy, the fast-slow split of chapter 5 applied to navigation (Wei et al. 2025, arXiv 2512.08186). For mobile manipulation, DreamTrajectory predicts an end-effector trajectory together with the whole-body action chunk, and checks the motion it intended against the motion it achieved (Yang et al. 2026, arXiv 2608.01381).

pillarpillarrobotgoal
Chosen curvature
0.20
What really happens
0.18 m from the goal
Figure 16.3Look-ahead planning picks the path whose imagined outcome is best; when the world model is wrong about how the robot moves, the best-looking path can be the one that crashes. A robot, two pillars and a goal. A navigation world model imagines thirteen candidate arcs and the robot drives the one it imagines ending nearest the goal. Make the model's sense of turning wrong, and compare the imagined path with the driven one.toy simulation Source: this book's toy, in rust/ch16-applications; look-ahead by ranking imagined trajectories follows Navigation World Models (arXiv 2412.03572).

The failure mode to expect is the one in figure 16.3. With an accurate model the robot threads past the first pillar to the goal. If the model believes the robot turns 25 percent more sharply than it does, the chosen path is really straighter than imagined and clips the first pillar. If it believes the robot turns 40 percent less sharply, the path turns too far and hits the second. This is chapter 8's model bias in a form a reader can drive.

Autonomous driving

Driving is where world models are most established, and they are used mainly to generate data and scenarios rather than to drive. A car must handle events too rare to collect, and a world model that can produce them on demand is worth more than one more policy. Figure 16.4 shows three roles.

ego cardashed: agents the model generates,differently each time

Generate scenarios

Generate realistic multi-camera driving video under control: ego motion, other agents, weather and road layout. GAIA-2 conditions on all of these and covers the UK, US and Germany; its makers frame it as a way to simulate both common and rare scenarios.

GAIA-1 (September 2023), GAIA-2 (March 2025).

Figure 16.4Driving uses world models mostly to produce the scenarios a car must be tested against: generated video, forecast occupancy, and futures learned from real driving. Three roles of world models around one intersection: generating scenarios, forecasting occupancy, and predicting future camera frames together with a path. Select a role.schematic, not measured Source: GAIA-1 (arXiv 2309.17080); GAIA-2 (2503.20523); OccWorld (2311.16038); DriveDreamer (2309.09777).

GAIA-1 cast driving world modelling as next-token prediction over video, text and action (Hu et al. 2023, arXiv 2309.17080). GAIA-2, a latent diffusion model, generates consistent multi-camera video conditioned on ego dynamics, other agents, environment and road semantics, across the UK, the US and Germany, so that "both common and rare driving scenarios" can be simulated (Russell et al. 2025, arXiv 2503.20523). OccWorld forecasts 3D occupancy and the ego path with a GPT-like transformer, a geometry-first model in chapter 9's sense (Zheng et al. 2023, arXiv 2311.16038). DriveDreamer builds its world model entirely from real-world driving (Wang et al. 2023, arXiv 2309.09777). Public evaluation runs on datasets and benchmarks such as nuScenes and NAVSIM (Caesar et al. 2019, arXiv 1903.11027; Dauner et al. 2024, arXiv 2406.15349).

The failure mode to expect is the one rare scenarios are meant to fix: a generated scenario looks plausible but does not behave like the real world, so a system validated on it is less safe than it appears. Chapter 8's gap between plausible and feasible applies with higher stakes.

Industrial and logistics

Warehouse picking was among the first commercial uses of robot foundation models. Covariant's RFM-1, an 8-billion-parameter any-to-any model trained on data from a fleet of warehouse robots, was announced in March 2024. Its makers described a data collection system that had gathered "tens of millions of trajectories" from robots deployed to dozens of customersvendor claim (Covariant 2024). By the first quarter of 2026, one industry survey counts at least eleven commercial deployments with VLAs as the primary policyindustry survey (SVRC (Robotics Center of Silicon Valley) 2026). Inspection is a second industrial use: Gemini Robotics-ER 1.6's instrument reading was developed with Boston Dynamics for facility inspection with Spotvendor claim (Google DeepMind 2026).

shared zone behind a light curtainsource toteoutgoing boxPLCI9I4I5I6I2I3I7

I2. Safety and compliance

Where it bites: the fence and the shared zone. What to check: An independent, certified safety layer between the policy and the motors (figure 15.5).

Figure 16.5In a real cell the industrial readiness implications are not abstract: each one bites at a specific place, and each suggests a specific test. A picking cell with seven of the survey's eleven implications marked where they apply. Select a marker for what to check.schematic, not measured Source: implications from the industrial readiness survey (Kube et al. 2026, arXiv 2603.06749); the cell and the checks are the authors' illustration.

The failure mode to expect is integration, not intelligence. Chapter 15's readiness survey found even the highest-rated models fulfilling none of its criteria on precision, real-time performance or cost-effective integration (Kube et al. 2026, arXiv 2603.06749). A cell fails in production because it cannot meet the conveyor's cycle time or talk to the warehouse system, more often than because the model cannot recognise a product.

Mobile service robots

Service robots in homes, hospitals and shops are where language matters most, and where embodied reasoning models (chapter 10) lead. A systematic review published in Robotics in 2026 names the domain's core challenges: translating natural-language instructions into executable actions, perceiving reliably in human-centred environments, estimating uncertainty for safe decisions, and fitting computation on board (Lisondra et al. 2025, arXiv 2505.20503). Benchmarks for instruction-following navigation, such as R2R and VLN-CE, come from this line of work (Anderson et al. 2017, arXiv 1711.07280; Krantz et al. 2020, arXiv 2004.02857).

The failure mode to expect is misunderstanding. A service robot that executes the wrong interpretation of a request confidently is worse than one that asks. Chapter 10's success detection and chapter 14's memory are the tools; uncertainty estimation, which the review lists as a core challenge, is the missing one.

In code

The crate's navigation world is two pillars, a goal and arcs of constant curvature.

rust/ch16-applications/src/lib.rsThe room, the candidate arcs and the step size behind figure 16.3
pub type P = [f64; 2];

/// A robot's pose: position and heading in radians.
#[derive(Clone, Copy, Debug, PartialEq)]
pub struct Pose {
    pub x: f64,
    pub y: f64,
    pub heading: f64,
}

/// A round obstacle on the floor.
#[derive(Clone, Copy, Debug)]
pub struct Obstacle {
    pub c: P,
    pub r: f64,
}

/// A room with two pillars: one beside the straight line to the goal, one beyond the goal.
pub fn scene() -> (Pose, P, Vec<Obstacle>) {
    (
        Pose { x: 0.0, y: 0.0, heading: 0.0 },
        [4.6, 3.1],
        vec![Obstacle { c: [2.6, -0.2], r: 0.5 }, Obstacle { c: [3.6, 4.6], r: 0.5 }],
    )
}

/// Candidate paths are arcs of constant curvature, as a sampling planner or a policy would propose.
pub const CURVATURES: [f64; 13] = [-0.2, -0.15, -0.1, -0.05, 0.0, 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4];
pub const STEPS: usize = 30;
pub const STEP: f64 = 0.2;

pub fn dist(a: P, b: P) -> f64 {
    ((a[0] - b[0]).powi(2) + (a[1] - b[1]).powi(2)).sqrt()
}

Exercises

The stubs are in rust/ch16-applications/src/exercises.rs. This chapter is equation-free; the exercises build the look-ahead planner of figure 16.3 one piece at a time.

Exercise 16.1

Roll a path out

Implement arc. Then explain why a real mobile base would also need a limit on how fast its heading can change.

Check your answer from the rust/ folder:

cargo test -p ch16-applications --test exercises ex16_1
Hint 1

Turn first, then move: heading += curvature × step, then x += step × cos(heading).

Hint 2

Include the starting position in the result.

Exercise 16.2

Does the path hit anything?

Implement collides. What does checking only the sampled points miss, and how would you fix it?

Check your answer from the rust/ folder:

cargo test -p ch16-applications --test exercises ex16_2
Hint

A point collides if its distance to an obstacle's centre is less than the obstacle's radius plus the robot's.

Exercise 16.3

Choose by imagining

Implement choose. Then find the range of turning errors for which the robot still reaches the goal safely, and suggest one check a planner could run before trusting its world model.

Check your answer from the rust/ folder:

cargo test -p ch16-applications --test exercises ex16_3
Hint 1

Imagine each candidate with its curvature multiplied by the model's turning error.

Hint 2

Skip candidates whose imagined path collides; keep the one whose imagined end is nearest the goal.

What we are not sure about

How much is deployed. Deployment counts and adopter lists in this chapter are vendor statements or come from industry surveys. Independent counts of robot foundation models in production do not exist in this book's sources.

Whether driving's lessons transfer. Driving built world models for scenario generation years before manipulation did. Whether manipulation will follow the same path, generating rare events for validation, or a different one is not yet clear.

Which domains come next. The matrix in figure 16.1 has many dashes. Some mark real gaps, such as geometry-first models in service robots; others only mark what this book's sources did not cover.

Further reading

References

30 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.

  1. Physical Intelligence, B. Ai, A. Amin and 85 others π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities. arXiv 2604.15483 v2, 2026-04-16.
  2. NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
  3. C. Chi, Z. Xu, S. Feng and 5 others Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. arXiv 2303.04137 v5, 2023-03-07.
  4. T. Z. Zhao, V. Kumar, S. Levine, C. Finn Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv 2304.13705 v1, 2023-04-23.
  5. Generalist AI GEN-1.5: Embodied Foundation Models are One-Shot Learners. generalistai.com, 2026-08-19.
  6. S. Fei, S. Wang, J. Shi and 10 others LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models. arXiv 2510.13626 v3, 2025-10-15.
  7. Z. Luo, Y. Yuan, T. Wang and 25 others SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control. Science Robotics 11 (117), eaed4592 (2026). arXiv 2511.07820 v4, 2025-11-11.
  8. Figure AI Helix: A Vision-Language-Action Model for Generalist Humanoid Control. figure.ai, 2025-02-20.
  9. Google DeepMind Introducing Gemini Robotics ER 2. blog.google, 2026-07-30.
  10. NVIDIA Newsroom NVIDIA and Global Robotics Leaders Take Physical AI to the Real World (GTC 2026, GR00T N2 preview). nvidianews.nvidia.com, 2026-03-16.
  11. NVIDIA Newsroom NVIDIA Announces NVIDIA Isaac GR00T Reference Humanoid Robot for Academic Research. nvidianews.nvidia.com, 2026-05-31.
  12. Rodney Brooks Predictions Scorecard, 2026 January 01. rodneybrooks.com, 2026-01-01.
  13. A. Bar, G. Zhou, D. Tran and 2 others Navigation World Models. arXiv 2412.03572 v2, 2024-12-04.
  14. W. Liu, H. Zhao, C. Li and 5 others X-MOBILITY: End-To-End Generalizable Navigation via World Modeling. arXiv 2410.17491 v3, 2024-10-23.
  15. M. Wei, C. Wan, J. Peng and 8 others Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation. arXiv 2512.08186 v1, 2025-12-09.
  16. Z. Yang, W. Zhang, X. Chen and 9 others DreamTrajectory: Trajectory-Guided Action Generation with World Model Alignment for Mobile Manipulation. arXiv 2608.01381 v2, 2026-08-02.
  17. A. Hu, L. Russell, H. Yeo and 5 others GAIA-1: A Generative World Model for Autonomous Driving. arXiv 2309.17080 v1, 2023-09-29.
  18. L. Russell, A. Hu, L. Bertoni and 4 others GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving. arXiv 2503.20523 v1, 2025-03-26.
  19. W. Zheng, W. Chen, Y. Huang and 3 others OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving. arXiv 2311.16038 v1, 2023-11-27.
  20. X. Wang, Z. Zhu, G. Huang and 3 others DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving. arXiv 2309.09777 v2, 2023-09-18.
  21. H. Caesar, V. Bankiti, A. H. Lang and 7 others nuScenes: A multimodal dataset for autonomous driving. arXiv 1903.11027 v5, 2019-03-26.
  22. D. Dauner, M. Hallgarten, T. Li and 9 others NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking. arXiv 2406.15349 v2, 2024-06-21.
  23. Covariant Covariant introduces RFM-1 to give robots the human-like ability to reason. covariant.ai, 2024-03-11.
  24. SVRC (Robotics Center of Silicon Valley) State of Robotics 2026. roboticscenter.ai, 2026-03.
  25. Google DeepMind Gemini Robotics-ER 1.6. blog.google, 2026-04-14.
  26. D. Kube, S. Hadwiger, T. Meisen Robotic Foundation Models for Industrial Control: A Comprehensive Survey and Readiness Assessment Framework. arXiv 2603.06749 v1, 2026-03-06.
  27. M. Lisondra, B. Benhabib, G. Nejat Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review. Robotics 2026, 15(3), 55. arXiv 2505.20503 v2, 2025-05-26.
  28. P. Anderson, Q. Wu, D. Teney and 6 others Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments. arXiv 1711.07280 v3, 2017-11-20.
  29. J. Krantz, E. Wijmans, A. Majumdar and 2 others Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments. arXiv 2004.02857 v2, 2020-04-06.
  30. Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.