Ch.17Part IV · Open problems and futures
Open problems
The ten problems the 2026 surveys converge on, each with a concrete failure example.
- Math
- L1
- Sources
- 20
- Figures
- 4
- Exercises
- 3
- Last checked
- 9 Oct 2026
One boxed equation per idea
13 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
9. Safety, verification and certification
Fitting learned policies into safety-rated systems built for components that can be specified and tested.
A failure you have already seen: A learned policy that sits outside the certifiable boundary of its own cell. See figure 15.5.
Hits hardest: every family.
Why this chapter exists
The previous sixteen chapters described what robot foundation models can do. This one collects what they cannot yet do, in the form the 2026 surveys converge on. The WAM survey names four major bottlenecks: action grounding, spatial consistency, closed-loop policy improvement and real-time inference (Lu et al. 2026, arXiv 2609.16074). The April 2026 comprehensive review and the industrial readiness survey add data, evaluation and safety (Psiris et al. 2026, arXiv 2604.15395; Kube et al. 2026, arXiv 2603.06749). The result is ten problems. Each is restated here in plain language, with one concrete failure taken from earlier in the book, so the problem is something you have seen rather than something you are told.
Two of the ten are under-researched relative to their importance: evaluation and safety. The source base for them is thin, and the chapter says so rather than padding.
After this chapter you will be able to:
- state each of the ten problems in a sentence and point to a failure that illustrates it;
- say which problems belong to particular families and which cut across all of them, using figure 17.1;
- explain why evaluation and safety sit in the severe but under-researched corner of figure 17.2.
There's a lot of great technological work happening, a lot of great talent working on these, but they are not yet well defined products.
How severe, how studied
9. Safety, verification and certification
Fitting learned policies into safety-rated systems built for components that can be specified and tested.
The placement is a judgement, made from how much recent work in this book's sources addresses each problem. Latency, action alignment and neural simulation receive steady attention, with many 2026 papers on each. Evaluation and safety receive little. A handful of benchmark critiques and one industrial survey stand against a deployment record that grows every quarter.
1. Action alignment
A video or latent model learns how the world moves; a robot needs joint commands. Turning one into the other without destroying what the model learned is the first bottleneck the WAM survey names. Its recommendations are staged training that aligns predicted motion to executable actions, and protection of the pretrained prior through frozen backbones, alignment layers and residual action heads (Lu et al. 2026, arXiv 2609.16074). The failure is figure 8.4's: a generated rollout in which every frame looks right and the implied motion is impossible for the robot (Wang et al. 2026, arXiv 2603.17808).
2. World-action factorisation
Should futures and actions be learned together, and should futures still be generated at test time? Fast-WAM's answer is to train with video and drop it at inference (Yuan et al. 2026, arXiv 2603.16666). LingBot-VLA 2.0 arrives at a similar design from the VLA side, with future prediction as a training-only proxy task (Wu et al. 2026, arXiv 2607.06403). The failure is paying for video generation at every step, the cost curve of figure 6.6, for a benefit a training-only objective might have given. A March 2026 robustness study complicates the picture. It found world-action models robust under perturbation, with LingBot-VA reaching 74.2 percent on RoboTwin 2.0-Plus and Cosmos Policy 82.2 percent on LIBERO-Plus, while VLAs such as π0.5 matched them on some tasks only with extensive and varied training (Zhang et al. 2026, arXiv 2603.22078).
3. Spatial and multi-view consistency
Robots usually have several cameras, and a model that generates the future generates several views of it. If those views disagree about where an object is, the plan built on them is wrong in a way no single view reveals. The WAM survey lists multi-view consistency among its open challenges (Lu et al. 2026, arXiv 2609.16074), and VLA-MBPO needed a dedicated decoding scheme to enforce it (Zhang et al. 2026, arXiv 2603.20607). Figure 17.3 shows the check that exposes the problem.
- Position error from views A and B
- 17 cm
- Camera C disagrees by
- 2.0°a check with a 1° tolerance would flag this future
The geometry is simple. Two bearings from cameras at known positions meet at one point . A third camera's predicted bearing should point at it, so the disagreement is
An error of 4 degrees in one generated view moves the located object by about 17 centimetres and makes the third view disagree by about 2 degrees: invisible frame by frame, obvious when compared. The failure generalises chapter 9's point: when futures are generated as images rather than as geometry, consistency is something to check, not something to assume.
4. Long-horizon memory
A policy that sees only the last moment forgets what it has done. MEM's combination of video short-term memory and text long-term memory reaches tasks of up to fifteen minutes (Torne et al. 2026, arXiv 2603.03596), and the WAM survey lists long-horizon memory among its open challenges (Lu et al. 2026, arXiv 2609.16074). The failure is figure 14.5's robot ten minutes into cooking that no longer knows whether the stove is on. The open question is architectural: memory inside the model, or an agentic harness around it that keeps state, as an embodied reasoning model does (chapter 10).
5. Neural simulators and closed-loop learning
Can a learned world model be trusted as a training environment? Errors in an imagined rollout compound: each predicted frame is the input to the next, so an error grows roughly as
where is how much the model amplifies an existing error and is the new error it adds each step. With , thirty steps turn a per-step error of 0.01 into more than 1.6, while showing the model a real observation every five steps keeps it near 0.01. Exercise 17.1 computes both. The failure is figure 8.3's policy that finds the model's blind spot, and figure 6.4 shows the drift itself. WorldArena found that visual quality is a poor guide to whether a world model is useful for training or evaluation (Shang et al. 2026, arXiv 2602.08971).
6. Inference latency and compute
Large models are slow and robots are impatient. The WAM survey lists inference efficiency among its four bottlenecks (Lu et al. 2026, arXiv 2609.16074), and a cross-accelerator study found VLA inference split between a compute-bound backbone and a memory-bound action expert, wasting hardware (Zhou et al. 2026, arXiv 2604.24447). The failure is figure 15.1's robot that is blind for two seconds while finishing a chunk. This is the problem with the most mature solutions (chapter 18), but world-action models that generate video at test time make it worse again.
7. Data scarcity and provenance
The top of the pyramid, action-labelled robot data, stays small and expensive, and the reported hours of different sources are not a common currency (chapter 13). The failure is figure 13.3's two curves that fit the same scaling points and disagree completely beyond them, and the unsourced survey claim that 40 percent synthetic data matches real dataindustry survey (SVRC (Robotics Center of Silicon Valley) 2026).
8. Evaluation
The field's benchmarks are fragmented and mostly simulated, its main public real-robot leaderboard covers short tasks that evaluators choose, and there is no shared protocol for long tasks on real robots (chapter 15). The evidence that this matters is now direct. LIBERO-Plus found success falling from 95 percent to below 30 percent under modest perturbations, and models ignoring their instructions (Fei et al. 2025, arXiv 2510.13626). A study of the BEHAVIOR-1K challenge argued that final-state metrics exaggerate performance and hide safety violations (Rasouli et al. 2026, arXiv 2604.21192). BeTTER found state-of-the-art VLAs failing "catastrophically" in dynamic scenarios that static benchmarks never test (Xu et al. 2026, arXiv 2604.18000). The failure is figure 11.4's: with the 20 trials common in papers, a real 10-point improvement is detected less than one time in five.
9. Safety, verification and certification
Functional-safety standards such as ISO 10218 and ISO/TS 15066 assume components that can be specified and tested (Kube et al. 2026, arXiv 2603.06749). A learned policy cannot be, so today it sits outside the certifiable boundary, wrapped by an independent monitor (figure 15.5). The industrial survey found the highest-rated model fulfilling 6 percent of its safety and compliance criteria (Kube et al. 2026, arXiv 2603.06749). Google's model card for Gemini Robotics ER 2 rules out safety-critical uses (Google DeepMind 2026), and Brooks notes that no one has yet solved the humanoid version of the emergency stop (Rodney Brooks 2026).
- Commands the monitor changed
- 69 of 120
- Time stopped for a person
- 11 steps
The monitor's rule fits in one line:
where is the policy's command, the distance to the nearest person and what reaches the motors. A rule this simple can be certified; the policy it supervises cannot. The failure that remains is everything inside the envelope. A monitor stops the robot from moving too fast near a person. It does not stop the robot from carefully putting a knife in the wrong place.
10. Embodiment transfer and the humanoid question
Does whole-body control belong inside the foundation model, or in a learned controller beneath it? GR00T N1.7 answers with latent action tokens decoded by SONIC (NVIDIA 2026; Luo et al. 2025, arXiv 2511.07820), and Helix with a fast controller under a slow reasoner (Figure AI 2025). The failure is a model that transfers from one arm to another but must relearn everything for a new body, or that leaves balance to a layer the foundation model cannot see (figure 16.2).
In code
The two geometric helpers behind figure 17.3 are a bearing and the smallest difference between two angles.
pub type P = [f64; 2];
/// A camera on the floor plan: where it is. It reports only the bearing, in radians, from its
/// position to an object, which is what a generated image pins down well.
#[derive(Clone, Copy, Debug)]
pub struct Camera {
pub at: P,
}
pub fn bearing(cam: Camera, p: P) -> f64 {
(p[1] - cam.at[1]).atan2(p[0] - cam.at[0])
}
/// Three cameras watching a workspace, as on a bimanual rig with a head and two wrist views.
pub fn rig() -> [Camera; 3] {
[Camera { at: [0.0, 0.0] }, Camera { at: [4.0, 0.0] }, Camera { at: [2.0, -1.5] }]
}
/// Smallest difference between two angles, in radians.
pub fn angle_diff(a: f64, b: f64) -> f64 {
let d = (a - b).rem_euclid(std::f64::consts::TAU);
d.min(std::f64::consts::TAU - d)
}Exercises
The stubs are in rust/ch17-open-problems/src/exercises.rs.
Exercise 17.1
Errors that compound
Implement rollout_error. Then find the longest imagination horizon for which the error stays below 0.1 with growth 0.1 and noise 0.01, and relate it to the chunk lengths of chapter 5.
Check your answer from the rust/ folder:
cargo test -p ch17-open-problems --test exercises ex17_1Hint 1
Update the error first, then reset it if this step is a regrounding step.
Hint 2
Steps count from 1.
Exercise 17.2
Do the views agree?
Implement consistency. Then find how large an error in one view must be before a 1 degree tolerance on the third view catches it, and say how the camera layout changes the answer.
Check your answer from the rust/ folder:
cargo test -p ch17-open-problems --test exercises ex17_2Hint 1
Write each ray as start + s × direction and solve the 2 × 2 system for s.
Hint 2
Measure the miss with the smallest difference between angles, not a raw subtraction.
Exercise 17.3
The monitor
Implement monitor. Then name two failures a monitor of this kind cannot prevent, and for each say which chapter of this book discusses a partial remedy.
Check your answer from the rust/ folder:
cargo test -p ch17-open-problems --test exercises ex17_3Hint 1
Check the distance first: inside the protective distance the command is zero.
Hint 2
Count a step as changed whenever the executed value differs from the command.
What we are not sure about
Whether the list is right. These ten problems are what the 2026 surveys converge on. Lists like this tend to miss the problem that turns out to matter, often something that looks like engineering rather than research.
How the problems interact. Latency limits how much world modelling a robot can afford at test time; data limits how good the world model can be; evaluation limits whether anyone can tell. Treating them one at a time, as this chapter does, understates how much progress on one depends on the others.
Whether evaluation and safety will stay neglected. Both are hard to publish and expensive to do. That may change quickly if a serious incident or a regulation forces it; this book cannot predict which comes first.
Further reading
- The WAM survey's section on open challenges (Lu et al. 2026, arXiv 2609.16074).
- The April 2026 comprehensive review's challenges and future directions (Psiris et al. 2026, arXiv 2604.15395).
- The robustness study comparing world-action models and VLAs (Zhang et al. 2026, arXiv 2603.22078).
- LIBERO-Plus and BeTTER, for what benchmarks hide (Fei et al. 2025, arXiv 2510.13626; Xu et al. 2026, arXiv 2604.18000).
- The industrial readiness survey, for safety and integration as research problems (Kube et al. 2026, arXiv 2603.06749).
References
20 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.
- A. Psiris, V. Argyriou, E. K. Markakis and 5 others Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions. Transactions on Machine Learning Research (TMLR), 07/2026. arXiv 2604.15395 v3, 2026-04-16.
- D. Kube, S. Hadwiger, T. Meisen Robotic Foundation Models for Industrial Control: A Comprehensive Survey and Readiness Assessment Framework. arXiv 2603.06749 v1, 2026-03-06.
- R. Wang, Q. Liu, Y. Deng and 3 others EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards. arXiv 2603.17808 v2, 2026-03-18.
- T. Yuan, Z. Dong, Y. Liu, H. Zhao Fast-WAM: Do World Action Models Need Test-time Future Imagination?. arXiv 2603.16666 v2, 2026-03-17.
- W. Wu, F. Wang, F. Lu and 21 others From Foundation to Application: Improving VLA Models in Practice. arXiv 2607.06403 v1, 2026-07-07.
- Z. Zhang, Z. Li, B. Rahmati and 11 others Do World Action Models Generalize Better than VLAs? A Robustness Study. arXiv 2603.22078 v5, 2026-03-23.
- Z. Zhang, H. Ren, Y. Sun and 6 others Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models. arXiv 2603.20607 v1, 2026-03-21.
- M. Torne, K. Pertsch, H. Walke and 14 others MEM: Multi-Scale Embodied Memory for Vision Language Action Models. arXiv 2603.03596 v2, 2026-03-04.
- Y. Shang, Z. Li, Y. Ma and 18 others WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models. arXiv 2602.08971 v2, 2026-02-09.
- K. Zhou, Q. Chen, D. Peng and 3 others Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment. arXiv 2604.24447 v1, 2026-04-27.
- SVRC (Robotics Center of Silicon Valley) State of Robotics 2026. roboticscenter.ai, 2026-03.
- S. Fei, S. Wang, J. Shi and 10 others LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models. arXiv 2510.13626 v3, 2025-10-15.
- A. Rasouli, Y. Wu, Z. Li and 4 others How VLAs (Really) Work In Open-World Environments. arXiv 2604.21192 v1, 2026-04-23.
- H. Xu, S. Zheng, H. Luo and 3 others Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models. arXiv 2604.18000 v1, 2026-04-20.
- Google DeepMind Introducing Gemini Robotics ER 2. blog.google, 2026-07-30.
- Rodney Brooks Predictions Scorecard, 2026 January 01. rodneybrooks.com, 2026-01-01.
- NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
- Z. Luo, Y. Yuan, T. Wang and 25 others SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control. Science Robotics 11 (117), eaed4592 (2026). arXiv 2511.07820 v4, 2025-11-11.
- Figure AI Helix: A Vision-Language-Action Model for Generalist Humanoid Control. figure.ai, 2025-02-20.