Skip to content

Ch.15Part III · Cross-cutting engineering

Deployment, evaluation and industrial readiness

Latency budgets, compute placement, benchmarks and what a readiness assessment looks like for a real cell.

Math
L0

No equations

Sources
14

4 papers from 2026

Figures
5

Most are interactive

Exercises
3

Rust, with tests

Last checked
9 Oct 2026

The field moves monthly

Papers cited, by half-year of first arXiv version

One pass, camera to motorcamera frameVLM backboneaction expert, 10 steps133 ms100 ms200 ms300 msWorst case, from a change to a reactionfinishing the current chunk1133 ms0 s0.5 s1 s1.5 s2 s2.5 sIn that time a person walking briskly (an illustrative 1.5 m/s) moves 1.70 m towards the robot.
Pipeline latency
133 ms
Worst-case reaction
1.13 s
Figure 15.1A robot that stops to think while executing long chunks can be blind to a change for more than a second; executing while thinking cuts that to a few hundred milliseconds. Top: one pass from camera to motor, stage by stage, on the robot or in the cloud. Bottom: the worst-case time from a change in the scene to the robot acting on it. Change the chunk length, the network and the execution mode.toy simulation Source: illustrative stage timings from this book, in rust/ch15-deployment; asynchronous execution after real-time chunking (arXiv 2506.07339).

Why this chapter exists

A model that works in a paper is not yet a model that works in a factory. Between the two lie three kinds of engineering. The model must run fast enough, on hardware the robot can carry or reach. Someone must measure whether it works, in a way that predicts how it will behave on the job. And it must fit into a safety case that a regulator, an insurer and a works council will accept. This chapter covers each in turn, at the level of what a deployment team needs to ask, not how to implement the answers.

After this chapter you will be able to:

  • break a robot's reaction time into its parts with figure 15.1, and explain why long action chunks executed synchronously are dangerous;
  • choose where to run each part of a model, on the robot or in the cloud, and say what each choice depends on;
  • say what each common benchmark measures and does not, and read an industrial readiness assessment.

Advancing robotics research for real-world problems requires humanoids that can move, interact and manipulate with precision in dynamic environments.

Michael Yip, professor at UC San Diego and director of the Advanced Robotics and Controls Laboratory. NVIDIA Announces NVIDIA Isaac GR00T Reference Humanoid Robot for Academic Research, NVIDIA Newsroom, 31 May 2026.

The latency budget

Every pass of a VLA has a cost in time: capture an image, prepare it, run the vision-language backbone, run the action head, send the result to the motor controller. Figure 15.1 uses illustrative timings, about 130 milliseconds on the robot, but the structure is general. A cross-accelerator study from April 2026 found the same two phases in real VLAs: a compute-bound backbone followed by a memory-bound action expert, which leaves hardware underused. It also found that right-sized edge devices can meet control rates at lower cost and energy than flagship GPUs (Zhou et al. 2026, arXiv 2604.24447).

The number that matters for safety is not one pass, but how long the robot is blind to a change. With action chunking (chapter 4), a robot that stops to think executes a whole chunk before it looks again. If the chunk lasts two seconds, a person who steps into the workspace just after the robot looked is not seen for two seconds, plus the time to think about it. Asynchronous execution looks again as soon as the last pass finishes, so the worst case is about two passes, whatever the chunk length. Real-time chunking does this for flow-based VLAs without retraining (Black et al. 2025, arXiv 2506.07339), and LingBot-VA uses asynchronous inference for the same reason (Li et al. 2026, arXiv 2601.21998). The other tools that shrink the budget are fewer denoising steps, quantisation, and export to optimised runtimes. GR00T N1.7 adds full-pipeline ONNX and TensorRT export for desktop GPUs and edge platforms (NVIDIA 2026). One industry survey reports quantised VLAs running at 10 to 25 Hz on consumer-grade GPUsindustry survey (SVRC (Robotics Center of Silicon Valley) 2026).

Where to run the model

Control cycles: on time or late, 200 in a rowPlan updates from the cloud reasoner, about once a second

The controller never misses its deadline; only plan updates arrive late when the network stalls. Gemini Robotics' cloud backbone with an on-robot decoder is this split.

Control cycles on time
100%
Depends on the network for
plans, not motion
Figure 15.2Running everything on the robot makes every cycle predictable; running in the cloud buys faster compute at the price of the network's worst moments; the split keeps motion local and puts only planning on the network. Two hundred control cycles over a shared network with occasional stalls. Compare three placements and change the control deadline.toy simulation Source: illustrative network model in rust/ch15-deployment; placements after Gemini Robotics (arXiv 2503.20020), Gemini Robotics On-Device (Google DeepMind, June 2025) and the GR00T platform on Jetson Thor (NVIDIA, July 2026).

There are three options. All on the robot makes timing predictable and removes the network as a failure mode. Google DeepMind's case for Gemini Robotics On-Device is exactly that it is "helpful for latency sensitive applications" and robust "in environments with intermittent or zero connectivity" (Google DeepMind 2025). NVIDIA's GR00T platform deploys post-trained policies to Jetson Thor for on-device inference (NVIDIA Technical Blog 2026). All in the cloud buys the biggest models and fastest GPUs, and makes every motion depend on the network. The split keeps the controller on the robot and puts the reasoning in the cloud. Gemini Robotics hosts its backbone in the cloud with a local action decoder on the robot (Gemini Robotics Team et al. 2025, arXiv 2503.20020), and the embodied reasoning models of chapter 10 sit above an on-robot controller. In the split, a network stall delays a plan, not a motion, which is usually the right failure to have.

Evaluation

Chapter 11 showed how many trials it takes to tell two policies apart. This section is about what the trials should be. Figure 15.3 places the common benchmarks by two properties: simulated or real, and short or long tasks.

real robots, long tasks:no shared benchmarksimulatedreal robotsshortlongplacement is the authors’ reading of each benchmark’s paperManiSkill3LIBEROLIBERO-PlusRoboTwin 2.0CALVINBEHAVIOR-1K challengeTRI blind A/B trialsRoboArena

LIBERO-Plus (2025)

LIBERO under perturbations of layout, camera, initial state, language, light, background and sensor noise.

What it does not measure: still simulated and short.

Figure 15.3Most benchmarks are simulated and short; real-robot evaluation exists but is either one lab's or short; nobody shares a benchmark for long tasks on real robots. Common evaluation setups placed by simulated versus real and by task horizon. Select one for what it measures and what it does not.schematic, not measured Source: each benchmark's paper: ManiSkill3 (arXiv 2410.00425), LIBERO (2306.03310), LIBERO-Plus (2510.13626), RoboTwin 2.0 (2506.18088), CALVIN (2112.03227), RoboArena (2506.18123), TRI LBM (2507.05331); BEHAVIOR-1K as studied in 2604.21192. Placement is the authors' reading.

Two findings sharpen the picture. LIBERO-Plus perturbed the setups VLAs are usually tested on, changing camera viewpoints, initial robot states, lighting and more. Performance dropped "from 95% to below 30% under modest perturbations", and the models largely ignored their language instructions (Fei et al. 2025, arXiv 2510.13626). A study of the BEHAVIOR-1K household challenge argued that scoring only the final state of objects "can potentially exaggerate reported performance", and proposed counting safety violations along the way (Rasouli et al. 2026, arXiv 2604.21192). On the real-robot side, RoboArena crowd-sources double-blind pairwise comparisons across evaluators who choose their own tasks, which scales diversity (Atreya et al. 2025, arXiv 2506.18123). TRI's blind A/B trials show what rigour looks like inside one lab (TRI LBM Team et al. 2025, arXiv 2507.05331). None of these measures contact quality, or recovery over a long task on a real robot.

Industrial readiness

A March 2026 survey asked a different question: not how well models do on benchmarks, but how well their papers address what industry needs. It distilled the industrial literature into eleven implications, from adaptability and safety to real-time performance and data, and operationalised them as 149 criteria. It then assessed 324 manipulation-capable models through 48,276 criterion-level decisions, made by an LLM-assisted pipeline validated against expert judgements (Kube et al. 2026, arXiv 2603.06749). Figure 15.4 shows its five highest-rated models.

10%30%50%Adaptability and flexibilitySafety and complianceHuman-robot interaction and collaborationRobustness and reliabilityPrecision and accuracyReal-time performanceCost-effectiveness and integrationExplainability and trustSensor fusion and perceptionStandardised benchmarking and evaluationData requirements and usage
Share of all 149 criteria fulfilled
12%
Strongest implication
Adaptability and flexibility
Implications with nothing fulfilled
3 of 11
Figure 15.4Even the highest-rated models satisfy about an eighth of the industrial criteria, with a peak on adaptability and near-zero on safety, precision and real-time performance. The share of each implication's criteria the survey rates as fulfilled, for its five highest-rated models. Select a model. Source: Kube et al. 2026 (arXiv 2603.06749), table 11; ratings from an LLM-assisted pipeline validated against expert judgements.

The highest-rated model, Gemini Robotics 1.5, fulfils 12 percent of the criteria overall, with 44 percent on adaptability and flexibility and none on precision, real-time performance or cost-effective integration (Kube et al. 2026, arXiv 2603.06749). The survey's reading is that recent work advances "one or two enabling dimensions at a time". Gemini's lead comes from combining a reasoning model with a VLA, which suggests that near-term deployments will be layered systems (Kube et al. 2026, arXiv 2603.06749). Its scores rate papers, not deployed systems, and a vendor may have solved in practice what it did not publish, so the numbers are a lower bound on readiness and an upper bound on public evidence.

Safety and certification

Industrial robots operate under functional-safety standards. The survey notes the role of ISO 10218 for the safe installation and operation of robots, and ISO/TS 15066 for collaborative operation and contact (Kube et al. 2026, arXiv 2603.06749). These standards were written for systems whose behaviour can be specified, analysed and tested part by part. A modular pipeline fits that pattern: perception, planning and control each have a specification, and much of the chain can sit inside the safety case. A learned end-to-end policy does not. It has no specification beyond its training data, and its failure modes cannot be enumerated in advance.

certifiable boundarycamerasraw pixelslearned policyone network, end to endlow-level controllertracks the commanded actionssafety monitorspeed, separation, force limitsrobotThe policy has no specification beyond its training data, so the safety case rests on an independent monitor that can stop it.
Figure 15.5With a learned policy, the certifiable boundary shrinks to an independent safety monitor that can stop the robot; the policy itself sits outside it. Where a certifiable boundary can be drawn around a modular pipeline and around a learned policy. Switch between the two.schematic, not measured Source: the authors' schematic; the role of independently certified supervision follows the industrial readiness survey (arXiv 2603.06749).

The practical answer today is the one in figure 15.5. Keep a conventional, certifiable safety layer, such as speed and separation monitoring and force limits, between the learned policy and the motors, independent of it and able to stop it. One of the survey's criteria describes exactly this: a system supervised "by independently certified controllers/modules" that can reject unsafe commands (Kube et al. 2026, arXiv 2603.06749). Model makers say the same in their own terms. Google's model card for Gemini Robotics ER 2 asks users not to use the robotics models for safety-critical applications such as healthcare and transportation (Google DeepMind 2026). This is the modular pipeline's advantage that learned policies gave up, and chapter 17 returns to it as an open problem.

In code

The crate's budget is a list of stages and a sum. The value is in making each stage's time explicit, so it can be measured and argued about.

rust/ch15-deployment/src/lib.rsTwo pipelines, a latency sum and a model of a shared network
/// One stage of the pipeline from camera to motor, with its time in milliseconds.
#[derive(Clone, Copy, Debug)]
pub struct Stage {
    pub name: &'static str,
    pub ms: f64,
}

/// An illustrative on-robot pipeline for a VLA with a flow-matching action expert.
pub fn on_device() -> Vec<Stage> {
    vec![
        Stage { name: "camera frame", ms: 33.0 },
        Stage { name: "preprocess and transfer", ms: 8.0 },
        Stage { name: "VLM backbone", ms: 60.0 },
        Stage { name: "action expert, 10 steps", ms: 30.0 },
        Stage { name: "send to controller", ms: 2.0 },
    ]
}

/// The same pipeline with the model in a data centre: a faster GPU, plus the network both ways.
pub fn in_the_cloud(round_trip_ms: f64) -> Vec<Stage> {
    vec![
        Stage { name: "camera frame", ms: 33.0 },
        Stage { name: "preprocess and transfer", ms: 8.0 },
        Stage { name: "network round trip", ms: round_trip_ms },
        Stage { name: "VLM backbone", ms: 25.0 },
        Stage { name: "action expert, 10 steps", ms: 12.0 },
        Stage { name: "send to controller", ms: 2.0 },
    ]
}

/// Total time from a frame being captured to its actions reaching the controller.
pub fn latency_ms(stages: &[Stage]) -> f64 {
    stages.iter().map(|s| s.ms).sum()
}

/// Round-trip times in milliseconds for a run of control cycles over a shared network:
/// mostly near `base`, with occasional long stalls. Deterministic, so runs can be compared.
pub fn network_samples(base: f64, n: usize) -> Vec<f64> {
    (0..n)
        .map(|i| {
            let x = ((i as u64).wrapping_mul(0x9E37_79B9_7F4A_7C15) >> 40) as f64 / (1u64 << 24) as f64;
            if x > 0.97 { base * 8.0 } else if x > 0.85 { base * 2.5 } else { base * (0.8 + 0.4 * x) }
        })
        .collect()
}

Exercises

The stubs are in rust/ch15-deployment/src/exercises.rs. This chapter is equation-free; the exercises are arithmetic about time and coverage.

Exercise 15.1

How long until the robot reacts?

Implement worst_reaction_ms. Then find the longest chunk for which a synchronous robot reacts within half a second, and compare it with the chunk lengths reported in chapters 4 and 5.

Check your answer from the rust/ folder:

cargo test -p ch15-deployment --test exercises ex15_1
Hint 1

Synchronous: the rest of the chunk, then one pass.

Hint 2

Asynchronous: the pass already running, then one more.

Exercise 15.2

What the network does to a deadline

Implement deadline_hit_rate. Then find the deadline the cloud placement needs to be on time 99 percent of the time with the crate's network, and say whether a motion controller could live with it.

Check your answer from the rust/ folder:

cargo test -p ch15-deployment --test exercises ex15_2
Hint

Count the cycles whose compute plus round trip fits the deadline, and divide by the number of cycles.

Exercise 15.3

The weakest link

Implement weakest. Then argue, in three sentences, whether a readiness review should rank models by their total score or by their weakest implication.

Check your answer from the rust/ folder:

cargo test -p ch15-deployment --test exercises ex15_3
Hint 1

Track the lowest value and its index in one pass.

Hint 2

Count the implications at or above the threshold separately.

What we are not sure about

What real deployments achieve. Latency, uptime and failure rates of deployed robot foundation models are rarely published. The figures in this chapter are illustrative or come from papers and surveys, not from production logs.

Whether a learned policy can ever sit inside the safety case. Today's answer is to wrap it in an independent safety layer. Whether formal methods, runtime verification or new standards will let some learned components inside the boundary is open, and chapter 17 discusses the candidates.

How to evaluate long tasks on real robots. The empty quadrant in figure 15.3 is the field's largest evaluation gap. Distributed efforts like RoboArena show one way to scale real-world evaluation; none yet covers long horizons.

Further reading

References

14 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.

  1. K. Zhou, Q. Chen, D. Peng and 3 others Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment. arXiv 2604.24447 v1, 2026-04-27.
  2. K. Black, M. Y. Galliker, S. Levine Real-Time Execution of Action Chunking Flow Policies. arXiv 2506.07339 v2, 2025-06-09.
  3. L. Li, Q. Zhang, Y. Luo and 9 others Causal World Modeling for Robot Control. arXiv 2601.21998 v2, 2026-01-29.
  4. NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
  5. SVRC (Robotics Center of Silicon Valley) State of Robotics 2026. roboticscenter.ai, 2026-03.
  6. Google DeepMind Gemini Robotics On-Device brings AI to local robotic devices. deepmind.google, 2025-06-24.
  7. NVIDIA Technical Blog Develop Humanoid Robot Policies End-to-End with NVIDIA Isaac GR00T. developer.nvidia.com, 2026-07-07.
  8. Gemini Robotics Team, S. Abeyruwan, J. Ainslie and 115 others Gemini Robotics: Bringing AI into the Physical World. arXiv 2503.20020 v1, 2025-03-25.
  9. S. Fei, S. Wang, J. Shi and 10 others LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models. arXiv 2510.13626 v3, 2025-10-15.
  10. A. Rasouli, Y. Wu, Z. Li and 4 others How VLAs (Really) Work In Open-World Environments. arXiv 2604.21192 v1, 2026-04-23.
  11. P. Atreya, K. Pertsch, T. Lee and 29 others RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies. arXiv 2506.18123 v2, 2025-06-22.
  12. TRI LBM Team, J. Barreiros, A. Beaulieu and 79 others A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation. arXiv 2507.05331 v1, 2025-07-07.
  13. D. Kube, S. Hadwiger, T. Meisen Robotic Foundation Models for Industrial Control: A Comprehensive Survey and Readiness Assessment Framework. arXiv 2603.06749 v1, 2026-03-06.
  14. Google DeepMind Introducing Gemini Robotics ER 2. blog.google, 2026-07-30.