Ch.15Part III · Cross-cutting engineering
Deployment, evaluation and industrial readiness
Latency budgets, compute placement, benchmarks and what a readiness assessment looks like for a real cell.
- Math
- L0
- Sources
- 14
- Figures
- 5
- Exercises
- 3
- Last checked
- 9 Oct 2026
No equations
4 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
- Pipeline latency
- 133 ms
- Worst-case reaction
- 1.13 s
Why this chapter exists
A model that works in a paper is not yet a model that works in a factory. Between the two lie three kinds of engineering. The model must run fast enough, on hardware the robot can carry or reach. Someone must measure whether it works, in a way that predicts how it will behave on the job. And it must fit into a safety case that a regulator, an insurer and a works council will accept. This chapter covers each in turn, at the level of what a deployment team needs to ask, not how to implement the answers.
After this chapter you will be able to:
- break a robot's reaction time into its parts with figure 15.1, and explain why long action chunks executed synchronously are dangerous;
- choose where to run each part of a model, on the robot or in the cloud, and say what each choice depends on;
- say what each common benchmark measures and does not, and read an industrial readiness assessment.
Advancing robotics research for real-world problems requires humanoids that can move, interact and manipulate with precision in dynamic environments.
The latency budget
Every pass of a VLA has a cost in time: capture an image, prepare it, run the vision-language backbone, run the action head, send the result to the motor controller. Figure 15.1 uses illustrative timings, about 130 milliseconds on the robot, but the structure is general. A cross-accelerator study from April 2026 found the same two phases in real VLAs: a compute-bound backbone followed by a memory-bound action expert, which leaves hardware underused. It also found that right-sized edge devices can meet control rates at lower cost and energy than flagship GPUs (Zhou et al. 2026, arXiv 2604.24447).
The number that matters for safety is not one pass, but how long the robot is blind to a change. With action chunking (chapter 4), a robot that stops to think executes a whole chunk before it looks again. If the chunk lasts two seconds, a person who steps into the workspace just after the robot looked is not seen for two seconds, plus the time to think about it. Asynchronous execution looks again as soon as the last pass finishes, so the worst case is about two passes, whatever the chunk length. Real-time chunking does this for flow-based VLAs without retraining (Black et al. 2025, arXiv 2506.07339), and LingBot-VA uses asynchronous inference for the same reason (Li et al. 2026, arXiv 2601.21998). The other tools that shrink the budget are fewer denoising steps, quantisation, and export to optimised runtimes. GR00T N1.7 adds full-pipeline ONNX and TensorRT export for desktop GPUs and edge platforms (NVIDIA 2026). One industry survey reports quantised VLAs running at 10 to 25 Hz on consumer-grade GPUsindustry survey (SVRC (Robotics Center of Silicon Valley) 2026).
Where to run the model
The controller never misses its deadline; only plan updates arrive late when the network stalls. Gemini Robotics' cloud backbone with an on-robot decoder is this split.
- Control cycles on time
- 100%
- Depends on the network for
- plans, not motion
There are three options. All on the robot makes timing predictable and removes the network as a failure mode. Google DeepMind's case for Gemini Robotics On-Device is exactly that it is "helpful for latency sensitive applications" and robust "in environments with intermittent or zero connectivity" (Google DeepMind 2025). NVIDIA's GR00T platform deploys post-trained policies to Jetson Thor for on-device inference (NVIDIA Technical Blog 2026). All in the cloud buys the biggest models and fastest GPUs, and makes every motion depend on the network. The split keeps the controller on the robot and puts the reasoning in the cloud. Gemini Robotics hosts its backbone in the cloud with a local action decoder on the robot (Gemini Robotics Team et al. 2025, arXiv 2503.20020), and the embodied reasoning models of chapter 10 sit above an on-robot controller. In the split, a network stall delays a plan, not a motion, which is usually the right failure to have.
Evaluation
Chapter 11 showed how many trials it takes to tell two policies apart. This section is about what the trials should be. Figure 15.3 places the common benchmarks by two properties: simulated or real, and short or long tasks.
LIBERO-Plus (2025)
LIBERO under perturbations of layout, camera, initial state, language, light, background and sensor noise.
What it does not measure: still simulated and short.
Two findings sharpen the picture. LIBERO-Plus perturbed the setups VLAs are usually tested on, changing camera viewpoints, initial robot states, lighting and more. Performance dropped "from 95% to below 30% under modest perturbations", and the models largely ignored their language instructions (Fei et al. 2025, arXiv 2510.13626). A study of the BEHAVIOR-1K household challenge argued that scoring only the final state of objects "can potentially exaggerate reported performance", and proposed counting safety violations along the way (Rasouli et al. 2026, arXiv 2604.21192). On the real-robot side, RoboArena crowd-sources double-blind pairwise comparisons across evaluators who choose their own tasks, which scales diversity (Atreya et al. 2025, arXiv 2506.18123). TRI's blind A/B trials show what rigour looks like inside one lab (TRI LBM Team et al. 2025, arXiv 2507.05331). None of these measures contact quality, or recovery over a long task on a real robot.
Industrial readiness
A March 2026 survey asked a different question: not how well models do on benchmarks, but how well their papers address what industry needs. It distilled the industrial literature into eleven implications, from adaptability and safety to real-time performance and data, and operationalised them as 149 criteria. It then assessed 324 manipulation-capable models through 48,276 criterion-level decisions, made by an LLM-assisted pipeline validated against expert judgements (Kube et al. 2026, arXiv 2603.06749). Figure 15.4 shows its five highest-rated models.
- Share of all 149 criteria fulfilled
- 12%
- Strongest implication
- Adaptability and flexibility
- Implications with nothing fulfilled
- 3 of 11
The highest-rated model, Gemini Robotics 1.5, fulfils 12 percent of the criteria overall, with 44 percent on adaptability and flexibility and none on precision, real-time performance or cost-effective integration (Kube et al. 2026, arXiv 2603.06749). The survey's reading is that recent work advances "one or two enabling dimensions at a time". Gemini's lead comes from combining a reasoning model with a VLA, which suggests that near-term deployments will be layered systems (Kube et al. 2026, arXiv 2603.06749). Its scores rate papers, not deployed systems, and a vendor may have solved in practice what it did not publish, so the numbers are a lower bound on readiness and an upper bound on public evidence.
Safety and certification
Industrial robots operate under functional-safety standards. The survey notes the role of ISO 10218 for the safe installation and operation of robots, and ISO/TS 15066 for collaborative operation and contact (Kube et al. 2026, arXiv 2603.06749). These standards were written for systems whose behaviour can be specified, analysed and tested part by part. A modular pipeline fits that pattern: perception, planning and control each have a specification, and much of the chain can sit inside the safety case. A learned end-to-end policy does not. It has no specification beyond its training data, and its failure modes cannot be enumerated in advance.
The practical answer today is the one in figure 15.5. Keep a conventional, certifiable safety layer, such as speed and separation monitoring and force limits, between the learned policy and the motors, independent of it and able to stop it. One of the survey's criteria describes exactly this: a system supervised "by independently certified controllers/modules" that can reject unsafe commands (Kube et al. 2026, arXiv 2603.06749). Model makers say the same in their own terms. Google's model card for Gemini Robotics ER 2 asks users not to use the robotics models for safety-critical applications such as healthcare and transportation (Google DeepMind 2026). This is the modular pipeline's advantage that learned policies gave up, and chapter 17 returns to it as an open problem.
In code
The crate's budget is a list of stages and a sum. The value is in making each stage's time explicit, so it can be measured and argued about.
/// One stage of the pipeline from camera to motor, with its time in milliseconds.
#[derive(Clone, Copy, Debug)]
pub struct Stage {
pub name: &'static str,
pub ms: f64,
}
/// An illustrative on-robot pipeline for a VLA with a flow-matching action expert.
pub fn on_device() -> Vec<Stage> {
vec![
Stage { name: "camera frame", ms: 33.0 },
Stage { name: "preprocess and transfer", ms: 8.0 },
Stage { name: "VLM backbone", ms: 60.0 },
Stage { name: "action expert, 10 steps", ms: 30.0 },
Stage { name: "send to controller", ms: 2.0 },
]
}
/// The same pipeline with the model in a data centre: a faster GPU, plus the network both ways.
pub fn in_the_cloud(round_trip_ms: f64) -> Vec<Stage> {
vec![
Stage { name: "camera frame", ms: 33.0 },
Stage { name: "preprocess and transfer", ms: 8.0 },
Stage { name: "network round trip", ms: round_trip_ms },
Stage { name: "VLM backbone", ms: 25.0 },
Stage { name: "action expert, 10 steps", ms: 12.0 },
Stage { name: "send to controller", ms: 2.0 },
]
}
/// Total time from a frame being captured to its actions reaching the controller.
pub fn latency_ms(stages: &[Stage]) -> f64 {
stages.iter().map(|s| s.ms).sum()
}
/// Round-trip times in milliseconds for a run of control cycles over a shared network:
/// mostly near `base`, with occasional long stalls. Deterministic, so runs can be compared.
pub fn network_samples(base: f64, n: usize) -> Vec<f64> {
(0..n)
.map(|i| {
let x = ((i as u64).wrapping_mul(0x9E37_79B9_7F4A_7C15) >> 40) as f64 / (1u64 << 24) as f64;
if x > 0.97 { base * 8.0 } else if x > 0.85 { base * 2.5 } else { base * (0.8 + 0.4 * x) }
})
.collect()
}Exercises
The stubs are in rust/ch15-deployment/src/exercises.rs. This chapter is equation-free; the exercises are arithmetic about time and coverage.
Exercise 15.1
How long until the robot reacts?
Implement worst_reaction_ms. Then find the longest chunk for which a synchronous robot reacts within half a second, and compare it with the chunk lengths reported in chapters 4 and 5.
Check your answer from the rust/ folder:
cargo test -p ch15-deployment --test exercises ex15_1Hint 1
Synchronous: the rest of the chunk, then one pass.
Hint 2
Asynchronous: the pass already running, then one more.
Exercise 15.2
What the network does to a deadline
Implement deadline_hit_rate. Then find the deadline the cloud placement needs to be on time 99 percent of the time with the crate's network, and say whether a motion controller could live with it.
Check your answer from the rust/ folder:
cargo test -p ch15-deployment --test exercises ex15_2Hint
Count the cycles whose compute plus round trip fits the deadline, and divide by the number of cycles.
Exercise 15.3
The weakest link
Implement weakest. Then argue, in three sentences, whether a readiness review should rank models by their total score or by their weakest implication.
Check your answer from the rust/ folder:
cargo test -p ch15-deployment --test exercises ex15_3Hint 1
Track the lowest value and its index in one pass.
Hint 2
Count the implications at or above the threshold separately.
What we are not sure about
What real deployments achieve. Latency, uptime and failure rates of deployed robot foundation models are rarely published. The figures in this chapter are illustrative or come from papers and surveys, not from production logs.
Whether a learned policy can ever sit inside the safety case. Today's answer is to wrap it in an independent safety layer. Whether formal methods, runtime verification or new standards will let some learned components inside the boundary is open, and chapter 17 discusses the candidates.
How to evaluate long tasks on real robots. The empty quadrant in figure 15.3 is the field's largest evaluation gap. Distributed efforts like RoboArena show one way to scale real-world evaluation; none yet covers long horizons.
Further reading
- The industrial readiness survey, the best checklist for a deployment review (Kube et al. 2026, arXiv 2603.06749).
- LIBERO-Plus, for what perturbations reveal about benchmark scores (Fei et al. 2025, arXiv 2510.13626).
- RoboArena, for distributed real-world evaluation (Atreya et al. 2025, arXiv 2506.18123).
- How VLAs (really) work in open-world environments, for safety-aware metrics (Rasouli et al. 2026, arXiv 2604.21192).
- Real-time chunking, for asynchronous execution without retraining (Black et al. 2025, arXiv 2506.07339).
- The cross-accelerator study, for choosing deployment hardware (Zhou et al. 2026, arXiv 2604.24447).
References
14 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- K. Zhou, Q. Chen, D. Peng and 3 others Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment. arXiv 2604.24447 v1, 2026-04-27.
- K. Black, M. Y. Galliker, S. Levine Real-Time Execution of Action Chunking Flow Policies. arXiv 2506.07339 v2, 2025-06-09.
- L. Li, Q. Zhang, Y. Luo and 9 others Causal World Modeling for Robot Control. arXiv 2601.21998 v2, 2026-01-29.
- NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
- SVRC (Robotics Center of Silicon Valley) State of Robotics 2026. roboticscenter.ai, 2026-03.
- Google DeepMind Gemini Robotics On-Device brings AI to local robotic devices. deepmind.google, 2025-06-24.
- NVIDIA Technical Blog Develop Humanoid Robot Policies End-to-End with NVIDIA Isaac GR00T. developer.nvidia.com, 2026-07-07.
- Gemini Robotics Team, S. Abeyruwan, J. Ainslie and 115 others Gemini Robotics: Bringing AI into the Physical World. arXiv 2503.20020 v1, 2025-03-25.
- S. Fei, S. Wang, J. Shi and 10 others LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models. arXiv 2510.13626 v3, 2025-10-15.
- A. Rasouli, Y. Wu, Z. Li and 4 others How VLAs (Really) Work In Open-World Environments. arXiv 2604.21192 v1, 2026-04-23.
- P. Atreya, K. Pertsch, T. Lee and 29 others RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies. arXiv 2506.18123 v2, 2025-06-22.
- TRI LBM Team, J. Barreiros, A. Beaulieu and 79 others A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation. arXiv 2507.05331 v1, 2025-07-07.
- D. Kube, S. Hadwiger, T. Meisen Robotic Foundation Models for Industrial Control: A Comprehensive Survey and Readiness Assessment Framework. arXiv 2603.06749 v1, 2026-03-06.
- Google DeepMind Introducing Gemini Robotics ER 2. blog.google, 2026-07-30.