Skip to content

Ch.19Part IV · Open problems and futures

Three scenarios for 2027 to 2030

Turn the open problems into futures the reader can watch for, without picking one.

Math
L0

No equations

Sources
20

15 papers from 2026

Figures
3

Most are interactive

Exercises
3

Rust, with tests

Last checked
10 Oct 2026

The field moves monthly

Papers cited, by half-year of first arXiv version

so farwhat each scenario would predict20262027202820292030VLA consolidationWAM convergenceLatent and geometricπ0.7, MEMRL from deployment is routineNo robustness gapWAMs a nicheGR00T N2 preview, Fast-WAMWAMs ship at control ratesRobustness replicatesPrediction inside every policyLaWAM, PointWorldLatent beats pixelA geometry benchmarkRGB generators as simulators

RGB generators as simulators 2029-12

If this scenario holds: pixel generation survives as a data engine and simulator (chapter 8), not as the policy.

Figure 19.1Three different futures are each consistent with what has happened so far; they diverge in what they predict for the next three years. Three scenarios on one time axis. Milestones up to October 2026 happened and are cited in earlier chapters; later milestones are what each scenario would predict, not forecasts. Select a milestone.schematic, not measured Source: past milestones from chapters 5 to 14; future milestones are the authors' construction from Book.md chapter 19.

Why this chapter exists

A book about a field that changes monthly should not end with a prediction. It should end by giving the reader a way to recognise which future is arriving. This chapter turns the open problems of chapter 17 and the bets of chapter 18 into three scenarios for 2027 to 2030. It names the evidence that would tell them apart, and closes with a checklist for reading the next model release in about ten minutes.

The scenarios are a device for organising uncertainty, not forecasts. Real outcomes will mix them, and the authors do not pick one.

After this chapter you will be able to:

  • describe the three scenarios and the evidence that already points towards each;
  • use figure 19.2 to keep score as new results arrive, and say which observations would change your mind;
  • read a launch post or paper with the checklist of figure 19.3: place the model on the family grid, find its data on the pyramid, and flag what its evidence does not show.

In the next ten years deployable dexterity may improve markedly, but not in the way the current hype for humanoid robots suggests.

Rodney Brooks, roboticist, co-founder of iRobot and Rethink Robotics. Predictions Scorecard, 2026 January 01, 1 January 2026.

Scenario 1: consolidation around reactive VLAs

The story. Reactive VLAs absorb the fixes for their weaknesses without changing what they are. Steerable prompts let them learn from failures and select behaviours by description (Physical Intelligence et al. 2026, arXiv 2604.15483). Multi-scale memory carries them through quarter-hour tasks (Torne et al. 2026, arXiv 2603.03596). RL from deployment makes them faster and more reliable than their demonstrators (Physical Intelligence et al. 2025, arXiv 2511.14759; Xu et al. 2026, arXiv 2604.23073). Embodied reasoning models above them handle planning and success detection (chapter 10). World-action models stay a research niche for contact-rich tasks, where predicting consequences is worth its cost.

Why it is plausible. The family has the largest open ecosystem, the cheapest inference and the longest deployment record (chapters 5 and 15). The March 2026 robustness study found that VLAs such as π0.5 can match world-action models on some tasks when trained extensively on diverse data (Zhang et al. 2026, arXiv 2603.22078).

What to watch. Independent real-robot studies at matched data that find no robustness gap between the families. VLA inference staying cheap on edge hardware while video-generating models do not get cheaper.

Scenario 2: convergence on world-action models

The story. Predicting the future becomes a standard part of every deployed policy, mostly through designs that cost little at inference: train with video, act without generating it. The words "VLA" and "WAM" stop meaning different things, because every serious model is trained on both actions and futures.

Why it is plausible. The same robustness study found world-action models strong under perturbation, with LingBot-VA at 74.2 percent on RoboTwin 2.0-Plus and Cosmos Policy at 82.2 percent on LIBERO-Plus (Zhang et al. 2026, arXiv 2603.22078). On RoboArena a world-action model led the VLAs it was compared with (chapter 6, figure 6.7)leaderboard claim (Atreya et al. 2025, arXiv 2506.18123). The cost argument is being dismantled from several sides: Fast-WAM drops generation at test time, Privileged Foresight Distillation keeps the future's benefit without the future, and LingBot-VLA 2.0 adds future prediction to a VLA as a training-only task (Yuan et al. 2026, arXiv 2603.16666; Fang et al. 2026, arXiv 2604.25859; Wu et al. 2026, arXiv 2607.06403). NVIDIA previewed GR00T N2 as a world-action model for humanoidsvendor claim (NVIDIA Newsroom 2026).

What to watch. The robustness advantage replicating outside the labs that reported it. A GR00T N2-class model shipping at control rates on robot hardware. Training with video becoming the default in new releases from the VLA side.

Scenario 3: resurgence of latent and geometric prediction

The story. Generating pixels turns out to be the wrong target for control. Policies predict in learned feature spaces or explicit geometry, pretrain on latent actions from human video, and keep RGB video generators only as simulators and data engines (chapter 8).

Why it is plausible. Latent world-action models already run far faster than pixel ones: LaWAM reports up to 24 times lower latency (Chen et al. 2026, arXiv 2606.15768). JEPA-style objectives avoid spending capacity on what cannot be predicted (Sun et al. 2026, arXiv 2602.10098; Lin et al. 2026, arXiv 2608.09381). Geometry makes the variables of contact explicit, and touch shows what cameras miss (Huang et al. 2026, arXiv 2601.03782; Higuera et al. 2026, arXiv 2602.06001). Even the test-time scaling work on pixel world-action models selects among rollouts using geometric consistency (Zhao et al. 2026, arXiv 2607.17454).

What to watch. Matched-compute comparisons that favour latent targets over pixel prediction. A shared benchmark for geometry-first models. Latent-action pretraining on human video beating teleoperation-only pretraining in an independent study.

What would tell them apart

Each scenario makes predictions the others do not. Figure 19.2 lists ten observable results and what each would mean for each scenario. Tick the ones you have seen reported, preferably by someone other than the model's maker, and watch the scores.

SeenObservable resultVLA consolidationWAM convergenceLatent and geometric
An independent real-robot study finds no robustness gap between VLAs and WAMs at matched data▲ supports▼ undercuts·
WAMs' robustness advantage replicates outside the labs that reported it▼ undercuts▲ supports▼ undercuts
A GR00T N2-class world-action model ships at control rates on robot hardware▼ undercuts▲ supports▼ undercuts
VLA inference stays cheap while video-generating WAMs do not get cheaper▲ supports▼ undercuts·
Steerability, memory and RL post-training close the long-horizon gap in VLAs▲ supports▼ undercuts·
Training with video and acting without it becomes the default in new releases· ▲ supports·
A matched-compute comparison favours latent targets over pixel prediction· ▼ undercuts▲ supports
A shared benchmark for geometry-first models appears· · ▲ supports
Latent-action pretraining on human video beats teleoperation-only pretraining, independently▼ undercuts· ▲ supports
Pixel world models end up used mainly as simulators, not as policies· ▼ undercuts▲ supports
VLA consolidation0WAM convergence0Latent and geometric0

Tick the results you have seen reported, by someone other than the model's maker.

Figure 19.2No single result settles the question, but a handful of independent results would separate the scenarios clearly; most of them are about evaluation, not new models. Ten observable results, each marked as supporting or undercutting each scenario. Tick the results you have seen; the bars keep score. Source: the authors' construction from Book.md chapter 19; the same table and scoring are in rust/ch19-scenarios.

Two things stand out in the table. Most of the discriminating results are evaluations, not new models: independent replications, matched-compute comparisons, shared benchmarks. Chapter 17 rated evaluation as severe and under-researched, and the scenarios inherit that. The future of the field will be decided partly by who builds the experiments that compare its families fairly. Second, several results move two scenarios at once in opposite directions, which is why one result can reverse the lead. Exercise 19.2 asks what it takes to be ahead.

How to read the next model release

New models will arrive faster than this book can be revised. The tools of the earlier chapters make a ten-minute reading possible. Figure 19.3 is the worksheet.

1. Place it on the family grid. What does it predict, and in what space?

Actions onlyActions plus futuresFutures onlyPlan in languagePixels or video latentsLearned feature latentsExplicit geometryLanguage or action tokensContinuous actions, no language backboneReactive VLAsWorld-action modelsWorld simulatorsLatent predictionLatent predictionLatent predictionGeometry-firstGeometry-firstGeometry-firstReactive VLAsEmbodied reasoningLarge behavior modelsSelect a cell.

2. Which layers of the data pyramid does it use?

3. Ten questions about the evidence.

7 red flags

no latency on named hardware; too few trials; no intervals; no independent evaluation; own numbers not flagged; baselines not matched; no failures reported.

A model trained on robot trajectories. Several flags are a reason to wait for evidence, not a reason to dismiss the model.

Figure 19.3Placing a release on the family grid and the data pyramid takes a minute; checking its evidence against ten questions takes ten, and usually raises several flags. The release-reading worksheet: select the release's cell on the family grid, tick the data layers it uses, and answer ten questions about its evidence. Source: the authors' checklist, drawing on the family grid (chapter 1), the data pyramid (chapters 2 and 13), the vendor-claim rules of Book.md and the evaluation practice of chapters 11 and 15.

The checklist in words:

  1. Family. What does it predict: actions only, actions and futures, futures only, or a plan in language? In what space: pixels, learned features, geometry, or tokens? That places it in chapters 5 to 11.
  2. Data. Which layers of the pyramid does it learn from, how many hours of each, and are those hours measured the same way (chapter 13)?
  3. Release. Are the weights released, under what licence, and does the backbone carry its own terms (chapter 5)?
  4. Footprint. How large is it, and what hardware does it need?
  5. Latency. What control rate does it reach, on which hardware, synchronous or asynchronous (chapter 15)?
  6. Evaluation. Real robots or simulation? Which benchmark, and is it saturated?
  7. Trials. How many per condition, with what intervals (chapter 11)?
  8. Who measured. The maker, or someone else?
  9. Claims. Are numbers about the maker's own data and model presented as claims?
  10. Comparison and failures. Are baselines matched for data and compute, and are failure cases shown?

Several flags are normal, and are a reason to wait for evidence rather than to dismiss the model. What the checklist protects against is the commonest mistake in reading this field: treating a maker's demonstration as a measurement.

In code

The crate's evidence table is the data behind figure 19.2, with the release questions as a struct.

rust/ch19-scenarios/src/lib.rsThe ten questions of the release checklist
/// What a launch post or paper says about a new model, answered yes or no by the reader.
#[derive(Clone, Copy, Debug, Default)]
pub struct Release {
    pub weights_and_license_stated: bool,
    pub size_and_hardware_stated: bool,
    pub latency_on_named_hardware: bool,
    pub real_robot_evaluation: bool,
    pub trials_per_condition: Option<u32>,
    pub intervals_reported: bool,
    pub independent_evaluation: bool,
    pub own_numbers_flagged: bool,
    pub baselines_matched: bool,
    pub failures_reported: bool,
}

Exercises

The stubs are in rust/ch19-scenarios/src/exercises.rs. This chapter is equation-free; the exercises are bookkeeping for judgement.

Exercise 19.1

Keep score

Implement tally. Then pick the two rows you think most likely to be observed by 2028, and say which scenario they would favour together.

Check your answer from the rust/ folder:

cargo test -p ch19-scenarios --test exercises ex19_1
Hint 1

Skip rows that have not been observed.

Hint 2

Supports adds one, undercuts subtracts one, neutral does nothing.

Exercise 19.2

Is anything ahead?

Implement leading. Then find the smallest set of observations that puts each scenario ahead on its own, and say which set you expect to see first.

Check your answer from the rust/ folder:

cargo test -p ch19-scenarios --test exercises ex19_2
Hint

Find the highest score, then count how many scenarios have it.

Exercise 19.3

Read a release

Implement red_flags. Then apply it to the announcement of a model in this book, chapter 5's π0.7 or chapter 10's Gemini Robotics ER 2, and list the flags you find.

Check your answer from the rust/ folder:

cargo test -p ch19-scenarios --test exercises ex19_3
Hint 1

Check each question in the order given, and add its flag when the answer is no.

Hint 2

Unstated trial counts count as too few.

What we are not sure about

Whether three scenarios are enough. A fourth is easy to imagine: progress stalls on safety and evaluation, and deployments stay narrow whatever the architecture. The scenarios in this chapter assume the technical race continues; the industrial readiness survey and Brooks' scorecard both suggest that deployment speed may be set by other things (Kube et al. 2026, arXiv 2603.06749; Rodney Brooks 2026).

Whether the scenarios are independent. They mix easily: a world-action model trained in latent space with a geometry-aware verifier belongs to all three. The table scores them separately because separating them makes evidence easier to read, not because the world will.

What the authors believe. The book has tried not to pick a winner. The evidence in October 2026 is thin enough that any confident pick would be a guess, and the field's own history (chapter 3) is full of confident picks that aged badly.

Further reading

References

20 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.

  1. Physical Intelligence, B. Ai, A. Amin and 85 others π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities. arXiv 2604.15483 v2, 2026-04-16.
  2. M. Torne, K. Pertsch, H. Walke and 14 others MEM: Multi-Scale Embodied Memory for Vision Language Action Models. arXiv 2603.03596 v2, 2026-03-04.
  3. Physical Intelligence, A. Amin, R. Aniceto and 53 others π*0.6: a VLA That Learns From Experience. arXiv 2511.14759 v2, 2025-11-18.
  4. C. Xu, J. T. Springenberg, M. Equi and 4 others RL Token: Bootstrapping Online RL with Vision-Language-Action Models. arXiv 2604.23073 v2, 2026-04-24.
  5. Z. Zhang, Z. Li, B. Rahmati and 11 others Do World Action Models Generalize Better than VLAs? A Robustness Study. arXiv 2603.22078 v5, 2026-03-23.
  6. P. Atreya, K. Pertsch, T. Lee and 29 others RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies. arXiv 2506.18123 v2, 2025-06-22.
  7. T. Yuan, Z. Dong, Y. Liu, H. Zhao Fast-WAM: Do World Action Models Need Test-time Future Imagination?. arXiv 2603.16666 v2, 2026-03-17.
  8. P. Fang, H. Chen, X. Cai Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models. arXiv 2604.25859 v2, 2026-04-28.
  9. W. Wu, F. Wang, F. Lu and 21 others From Foundation to Application: Improving VLA Models in Practice. arXiv 2607.06403 v1, 2026-07-07.
  10. NVIDIA Newsroom NVIDIA and Global Robotics Leaders Take Physical AI to the Real World (GTC 2026, GR00T N2 preview). nvidianews.nvidia.com, 2026-03-16.
  11. J. Chen, K. Wang, K. Chen and 9 others LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies. arXiv 2606.15768 v1, 2026-06-14.
  12. J. Sun, W. Zhang, Z. Qi and 6 others VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model. arXiv 2602.10098 v2, 2026-02-10.
  13. Y. Lin, J. He, S. Bao and 6 others JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling. arXiv 2608.09381 v1, 2026-08-10.
  14. W. Huang, Y. Chao, A. Mousavian and 4 others PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation. arXiv 2601.03782 v1, 2026-01-07.
  15. C. Higuera, S. Arnaud, B. Boots and 3 others Visuo-Tactile World Models. arXiv 2602.06001 v1, 2026-02-05.
  16. Z. Zhao, M. Cho, H. shen and 4 others Test-Time Scaling for World Action Models via Zero-Shot Geometric Evaluation. arXiv 2607.17454 v1, 2026-07-20.
  17. D. Kube, S. Hadwiger, T. Meisen Robotic Foundation Models for Industrial Control: A Comprehensive Survey and Readiness Assessment Framework. arXiv 2603.06749 v1, 2026-03-06.
  18. Rodney Brooks Predictions Scorecard, 2026 January 01. rodneybrooks.com, 2026-01-01.
  19. Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.
  20. NVIDIA Technical Blog Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models. developer.nvidia.com, 2026-06-15.