Ch.18Part IV · Open problems and futures
Candidate solutions and research bets
For each open problem: the approaches being tried, who is trying them, and how mature they are.
- Math
- L1
- Sources
- 43
- Figures
- 3
- Exercises
- 3
- Last checked
- 9 Oct 2026
One boxed equation per idea
23 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
A certified envelope around the learned policy
Problem: Safety, verification and certification. Maturity: established. Who: Industrial practice; survey criterion on independently certified modules.
Families it would change: Reactive VLAs, Large behavior models.
Why this chapter exists
Chapter 17 listed what is unsolved. This chapter lists what is being tried. For each problem it names the candidate solutions visible in the source base, who is trying them, and an assessment of how far along each is. The assessment uses three levels. Established means in shipped models or standard practice, with more than one group reporting it works. Promising means demonstrated by more than one group or on real robots, but not yet standard. Early means one or two papers, often only in simulation.
"Promising" is a judgement. Each assessment is dated October 2026, and the chapter expects to be revised.
After this chapter you will be able to:
- name, for each open problem, at least one candidate solution and say how mature it is;
- say which families of figure 1.1 each solution would change, using figure 18.2;
- explain with figure 18.3 why spending more compute at test time helps only as much as the verifier allows.
Learning manipulable representations of the world and its dynamics is central to AI.
Where the solutions land
Most solutions do not fix a problem in one family; they shift where families sit relative to each other. Video-only-in-training makes world-action models cost what VLAs cost at inference. Latent prediction pulls world-action models towards chapter 7. Explicit 3D state pulls them towards chapter 9. Figure 18.2 draws those moves on the family grid.
Action alignment and factorisation
Action alignment. The WAM survey's two recommendations are the main line. The first is staged training, which aligns a pretrained video model's predicted motion to executable actions on embodied data. The second is prior preservation through frozen parameters, alignment layers and residual action heads (Lu et al. 2026, arXiv 2609.16074). Both are promising. Latent action interfaces with small per-robot decoders (chapter 7) are a second promising route (Ye et al. 2024, arXiv 2410.11758; Routray et al. 2025, arXiv 2511.07732), and using an inverse dynamics model as a reward for executability is early (Wang et al. 2026, arXiv 2603.17808).
Factorisation. Training with futures and acting without them is the most visible bet. Fast-WAM keeps video co-training and drops future generation at test time (Yuan et al. 2026, arXiv 2603.16666), and LingBot-VLA 2.0 reaches the same design from the VLA side (Wu et al. 2026, arXiv 2607.06403). Privileged Foresight Distillation sharpens the idea. It treats what the future teaches as a correction to the action, measured by a teacher that sees the future, and distils it into a small adapter on a student that never generates video (Fang et al. 2026, arXiv 2604.25859). Predicting representations instead of pixels, as JEPA-WAM and LaWAM do, is the other promising bet (Lin et al. 2026, arXiv 2608.09381; Chen et al. 2026, arXiv 2606.15768).
Consistency, memory and neural simulators
Multi-view consistency. Explicit 3D state (chapter 9) is the structural answer, and is promising (Huang et al. 2026, arXiv 2601.03782; Lu et al. 2025, arXiv 2508.17600). A July 2026 paper turns chapter 17's consistency check into a selection rule. It samples several rollouts from a world-action model and keeps the one whose predicted futures agree best across views, measured by depth reprojection with a frozen geometry model (Zhao et al. 2026, arXiv 2607.17454). With eight samples it raised RoboCasa success from 66.3 to 68.4 percent with Cosmos Policy, and from 80.8 to 82.5 percent with another backbone. A gate that only samples when the first rollout looks inconsistent kept about three-quarters of that gain while sampling extra at about a quarter of decisions (Zhao et al. 2026, arXiv 2607.17454). Those modest gains are what figure 18.3 predicts for an imperfect verifier.
Memory. MEM's video-plus-text memory is promising, published by one lab and built into π0.7 (Torne et al. 2026, arXiv 2603.03596; Physical Intelligence et al. 2026, arXiv 2604.15483). The alternative is an agentic harness: a reasoning model that keeps state outside the policy, as Gemini Robotics ER 2 does (Google DeepMind 2026).
Neural simulators. Closing the loop, so the world model is retrained on the policy's own rollouts, is promising (Liu et al. 2026, arXiv 2602.06508; Wang et al. 2026, arXiv 2603.08546). Dense teacher supervision instead of sparse reward (Yang et al. 2026, arXiv 2608.22364) and unifying search with value learning (Cheng et al. 2026, arXiv 2605.26282) are early.
Spending compute at test time
A general bet that cuts across problems 3, 5 and 6 is to spend more computation at decision time. Sample several candidate plans, let a verifier check them, and execute one the verifier accepts. The verifier can be a world model imagining each plan's outcome, a geometric consistency check, or a reasoning model judging the plan. If one candidate succeeds with probability , the verifier accepts good plans with probability and bad ones with probability , and candidates are checked one at a time up to , then the chance of success is
Read it from the left. is the chance that any one candidate is accepted. is the chance that none of the is, in which case the robot executes the last, rejected one. The first fraction is the verifier's precision, the chance that an accepted plan is good. As grows, approaches that precision and never exceeds it.
- Success with 4 candidates
- 79%
- Expected time to decide
- 262 msat an illustrative 100 ms to generate and 50 ms to check each plan
With a policy that succeeds half the time and a verifier that accepts 90 percent of good plans and 20 percent of bad ones, the ceiling is about 82 percent, reached within about eight candidates. Cutting the verifier's false acceptances from 20 to 5 percent raises the ceiling to about 95 percent. The lesson for research bets is that test-time compute is worth as much as the verifier is good, which is why verification is becoming a research problem of its own. FFDC, from May 2026, uses a learned verifier the other way round: to decide when an imagined plan has stopped matching reality, so the robot replans only when needed. It reports up to 74 percent fewer planning calls on a real humanoid, with higher success (Wang et al. 2026, arXiv 2605.06222).
Latency, data, evaluation, safety, embodiment
Latency has the most mature answers. Asynchronous execution and real-time chunking, few-step denoising and parallel decoding are established (Black et al. 2025, arXiv 2506.07339; Kim et al. 2025, arXiv 2502.19645). Right-sized, quantised edge deployment is promising (Zhou et al. 2026, arXiv 2604.24447; NVIDIA 2026), and adaptive replanning is early (Wang et al. 2026, arXiv 2605.06222).
Data. Egocentric human video at scale is promising and heavily funded (Zheng et al. 2026, arXiv 2602.16710; Gao et al. 2026, arXiv 2602.06949; Dyna Robotics 2026; Generalist AI 2026). Synthetic rewriting is promising, with its exchange rate against real data unsettled (Jang et al. 2025, arXiv 2505.12705; NVIDIA et al. 2025, arXiv 2503.14492). Learning from failures through labelled prompts is early (Physical Intelligence et al. 2026, arXiv 2604.15483).
Evaluation. Distributed double-blind real-robot comparison and blind A/B trials with statistical confidence are promising (Atreya et al. 2025, arXiv 2506.18123; TRI LBM Team et al. 2025, arXiv 2507.05331). Perturbation and safety-aware benchmarks are early (Fei et al. 2025, arXiv 2510.13626; Xu et al. 2026, arXiv 2604.18000; Rasouli et al. 2026, arXiv 2604.21192). Readiness frameworks are early too, with one detailed instance so far (Kube et al. 2026, arXiv 2603.06749).
Safety. The only established answer is architectural: a certified envelope around the learned policy (chapter 15). Reasoning models as runtime monitors, judging whether a step succeeded or a plan is unsafe, are early (Google DeepMind 2026). Physical fallbacks for balancing robots are an open question (Rodney Brooks 2026).
Embodiment. Latent tokens to a learned whole-body controller is promising, shipped in GR00T N1.7 (NVIDIA 2026; Luo et al. 2025, arXiv 2511.07820). Cross-embodiment pretraining is promising (Open X-Embodiment Collaboration et al. 2023, arXiv 2310.08864; Octo Model Team et al. 2024, arXiv 2405.12213). Embodiment-agnostic action spaces, such as PointWorld's 3D point flows, are early (Huang et al. 2026, arXiv 2601.03782).
In code
The crate's verifier has two numbers, and the procedure is a short loop.
/// A verifier that checks candidate plans before one is executed.
#[derive(Clone, Copy, Debug)]
pub struct Verifier {
/// Chance it accepts a plan that would succeed.
pub accept_good: f64,
/// Chance it accepts a plan that would fail.
pub accept_bad: f64,
}
/// The policy's chance that any one sampled plan succeeds, and the verifier that screens them.
/// Candidates are checked one at a time; the first accepted plan is executed. If none of `n` is
/// accepted, the last one is executed anyway.
#[derive(Clone, Copy, Debug)]
pub struct Setup {
pub p_good: f64,
pub verifier: Verifier,
}
impl Setup {
/// Chance that a single candidate is accepted.
pub fn accept_rate(&self) -> f64 {
self.p_good * self.verifier.accept_good + (1.0 - self.p_good) * self.verifier.accept_bad
}
}
/// A small deterministic generator, so simulated runs can be compared with the exact answers.
pub struct Rng(u64);
impl Rng {
pub fn new(seed: u64) -> Rng {
Rng(seed)
}
pub fn next_f64(&mut self) -> f64 {
self.0 = self.0.wrapping_add(0x9E37_79B9_7F4A_7C15);
let mut z = self.0;
z = (z ^ (z >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9);
z = (z ^ (z >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB);
((z ^ (z >> 31)) >> 11) as f64 / (1u64 << 53) as f64
}
}
/// Run the procedure once: returns (success, candidates examined).
pub fn run_once(s: &Setup, n: usize, rng: &mut Rng) -> (bool, usize) {
for k in 1..=n {
let good = rng.next_f64() < s.p_good;
let accept = rng.next_f64() < if good { s.verifier.accept_good } else { s.verifier.accept_bad };
if accept || k == n {
return (good, k);
}
}
unreachable!("n must be at least 1")
}Exercises
The stubs are in rust/ch18-research-bets/src/exercises.rs.
Exercise 18.1
What best-of-N buys
Implement success. The test compares it with 200,000 simulated runs. Then explain in one sentence why the answer with one candidate must equal the policy's own success rate.
Check your answer from the rust/ folder:
cargo test -p ch18-research-bets --test exercises ex18_1Hint 1
Split on whether any candidate is accepted: if one is, it is good with the verifier's precision.
Hint 2
If none is, the executed plan was rejected, which changes its chance of being good.
Exercise 18.2
The ceiling
Implement ceiling and n_for. Then find which matters more for the ceiling, raising acceptance of good plans from 0.9 to 0.99 or lowering acceptance of bad ones from 0.2 to 0.1, and say what that implies for building verifiers.
Check your answer from the rust/ folder:
cargo test -p ch18-research-bets --test exercises ex18_2Hint 1
The ceiling is the limit as N grows: the verifier's precision.
Hint 2
Search N upwards until the target is reached, and give up at 1,000.
Exercise 18.3
What it costs
Implement expected_latency_ms. Then combine it with chapter 15's latency budget: how many candidates can a robot afford if it must decide within 300 ms?
Check your answer from the rust/ folder:
cargo test -p ch18-research-bets --test exercises ex18_3Hint 1
The expected number of candidates checked is a geometric sum that stops at N.
Hint 2
Multiply by the cost of generating and checking one candidate.
What we are not sure about
Which bets will pay off. The assessments are a snapshot. Several "early" entries are a single paper from mid-2026; some will become standard within a year and others will disappear.
Whether the families converge. Figure 18.2 shows solutions pulling world-action models towards VLAs, latent models and geometry. Convergence on one hybrid design is one reading; another is that each family absorbs the others' best ideas and stays distinct. Chapter 19 treats these as scenarios.
Where the hard problems get their solutions. Evaluation and safety have the fewest and least mature candidates. They may be solved by research, by regulation, or by industry practice that never appears in a paper.
Further reading
- The WAM survey's section on open challenges and future directions (Lu et al. 2026, arXiv 2609.16074).
- The world-model-to-WAM tutorial, for a compact map of the approaches (Zhang et al. 2026, arXiv 2607.00836).
- Test-time scaling for WAMs, for verification by geometric consistency (Zhao et al. 2026, arXiv 2607.17454).
- When to Trust Imagination, for verifying a plan against reality as it executes (Wang et al. 2026, arXiv 2605.06222).
- Privileged Foresight Distillation, for keeping the benefit of futures without generating them (Fang et al. 2026, arXiv 2604.25859).
- LeJEPA, for the representation-learning bet (Balestriero et al. 2025, arXiv 2511.08544).
References
43 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.
- S. Ye, J. Jang, B. Jeon and 13 others Latent Action Pretraining from Videos. arXiv 2410.11758 v2, 2024-10-15.
- S. Routray, H. Pan, U. Jain and 2 others ViPRA: Video Prediction for Robot Actions. arXiv 2511.07732 v2, 2025-11-11.
- R. Wang, Q. Liu, Y. Deng and 3 others EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards. arXiv 2603.17808 v2, 2026-03-18.
- T. Yuan, Z. Dong, Y. Liu, H. Zhao Fast-WAM: Do World Action Models Need Test-time Future Imagination?. arXiv 2603.16666 v2, 2026-03-17.
- W. Wu, F. Wang, F. Lu and 21 others From Foundation to Application: Improving VLA Models in Practice. arXiv 2607.06403 v1, 2026-07-07.
- P. Fang, H. Chen, X. Cai Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models. arXiv 2604.25859 v2, 2026-04-28.
- Y. Lin, J. He, S. Bao and 6 others JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling. arXiv 2608.09381 v1, 2026-08-10.
- J. Chen, K. Wang, K. Chen and 9 others LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies. arXiv 2606.15768 v1, 2026-06-14.
- W. Huang, Y. Chao, A. Mousavian and 4 others PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation. arXiv 2601.03782 v1, 2026-01-07.
- G. Lu, B. Jia, P. Li and 4 others GWM: Towards Scalable Gaussian World Models for Robotic Manipulation. arXiv 2508.17600 v2, 2025-08-25.
- Z. Zhao, M. Cho, H. shen and 4 others Test-Time Scaling for World Action Models via Zero-Shot Geometric Evaluation. arXiv 2607.17454 v1, 2026-07-20.
- M. Torne, K. Pertsch, H. Walke and 14 others MEM: Multi-Scale Embodied Memory for Vision Language Action Models. arXiv 2603.03596 v2, 2026-03-04.
- Physical Intelligence, B. Ai, A. Amin and 85 others π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities. arXiv 2604.15483 v2, 2026-04-16.
- Google DeepMind Introducing Gemini Robotics ER 2. blog.google, 2026-07-30.
- X. Liu, Z. Bai, H. Ci and 2 others World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy. arXiv 2602.06508 v2, 2026-02-06.
- Y. Wang, R. Syed, F. Wu and 7 others Interactive World Simulator for Robot Policy Training and Evaluation. arXiv 2603.08546 v1, 2026-03-09.
- L. Yang, Z. Jiang, C. Sheng, Z. Tang WAM-OPD: On-Policy Distillation for World Action Models. arXiv 2608.22364 v1, 2026-08-23.
- X. Cheng, W. Yuan, Z. Mu and 5 others Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization. arXiv 2605.26282 v1, 2026-05-25.
- R. Wang, Y. Zhang, J. Lin and 4 others When to Trust Imagination: Adaptive Action Execution for World Action Models. arXiv 2605.06222 v2, 2026-05-07.
- K. Black, M. Y. Galliker, S. Levine Real-Time Execution of Action Chunking Flow Policies. arXiv 2506.07339 v2, 2025-06-09.
- M. J. Kim, C. Finn, P. Liang Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success. arXiv 2502.19645 v2, 2025-02-27.
- K. Zhou, Q. Chen, D. Peng and 3 others Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment. arXiv 2604.24447 v1, 2026-04-27.
- NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
- R. Zheng, D. Niu, Y. Xie and 12 others EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data. arXiv 2602.16710 v1, 2026-02-18.
- S. Gao, W. Liang, K. Zheng and 27 others DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos. arXiv 2602.06949 v1, 2026-02-06.
- Dyna Robotics Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models. dyna.co, 2026-08.
- Generalist AI GEN-1: Scaling Embodied Foundation Models to Mastery. generalistai.com, 2026-04-02.
- J. Jang, S. Ye, Z. Lin and 25 others DreamGen: Unlocking Generalization in Robot Learning through Video World Models. arXiv 2505.12705 v2, 2025-05-19.
- NVIDIA, H. A. Alhaija, J. Alvarez and 37 others Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control. arXiv 2503.14492 v2, 2025-03-18.
- P. Atreya, K. Pertsch, T. Lee and 29 others RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies. arXiv 2506.18123 v2, 2025-06-22.
- TRI LBM Team, J. Barreiros, A. Beaulieu and 79 others A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation. arXiv 2507.05331 v1, 2025-07-07.
- S. Fei, S. Wang, J. Shi and 10 others LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models. arXiv 2510.13626 v3, 2025-10-15.
- H. Xu, S. Zheng, H. Luo and 3 others Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models. arXiv 2604.18000 v1, 2026-04-20.
- A. Rasouli, Y. Wu, Z. Li and 4 others How VLAs (Really) Work In Open-World Environments. arXiv 2604.21192 v1, 2026-04-23.
- D. Kube, S. Hadwiger, T. Meisen Robotic Foundation Models for Industrial Control: A Comprehensive Survey and Readiness Assessment Framework. arXiv 2603.06749 v1, 2026-03-06.
- Google DeepMind Gemini Robotics-ER 1.6. blog.google, 2026-04-14.
- Rodney Brooks Predictions Scorecard, 2026 January 01. rodneybrooks.com, 2026-01-01.
- Z. Luo, Y. Yuan, T. Wang and 25 others SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control. Science Robotics 11 (117), eaed4592 (2026). arXiv 2511.07820 v4, 2025-11-11.
- Open X-Embodiment Collaboration, A. O'Neill, A. Rehman and 291 others Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv 2310.08864 v9, 2023-10-13.
- Octo Model Team, D. Ghosh, H. Walke and 16 others Octo: An Open-Source Generalist Robot Policy. arXiv 2405.12213 v2, 2024-05-20.
- X. Zhang, X. Zeng, W. Zhang From World Models to World Action Models: A Concise Tutorial for Robotics. arXiv 2607.00836 v8, 2026-07-01.
- R. Balestriero, Y. LeCun LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics. arXiv 2511.08544 v3, 2025-11-11.