Skip to content

Sources

Sources

Every paper, report and page in the book’s reference collection, newest first, with the chapters that cite it.

In the collection
190

local copies in reference/

Papers
159

linked to arXiv

Web and reports
31

archived pages

Cited in chapters
170

the rest is background

190 of 190

2026 88

  1. Latent evolving World Action Model

    X. Fang, B. Duan, H. Wu and 2 others · 2026-09-23 · arXiv 2609.27455

    CH.06

  2. ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

    S. Lian, B. Yu, Z. Shen and 5 others · 2026-09-16 · arXiv 2609.18487

    CH.04

  3. World-Action Models for Robot Learning and Control: A Survey

    Z. Lu, H. Zhai, G. Wang and 13 others · 2026-09-13 · arXiv 2609.16074

    CH.01CH.03CH.04CH.05CH.06CH.09CH.12CH.13CH.14CH.16CH.17CH.18CH.19

  4. Awesome-World-Action-Models (project page and paper list of the WAM survey)

    RCL-Robotics · 2026-09-13 · rcl-robotics.github.io

  5. VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models

    C. Su, Z. Shen, Y. Qian and 8 others · 2026-09-03 · arXiv 2609.04355

    CH.14

  6. Motus2: A Self-Evolving General World Model for Dexterous Manipulation

    H. Bi, Z. Zhou, Y. Tang and 16 others · 2026-08-31 · arXiv 2608.30237

    CH.06

  7. 4DGS-WAM: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting

    Y. Ma, Z. Xu, I. King · 2026-08-26 · arXiv 2608.25956

    CH.09

  8. WAM-OPD: On-Policy Distillation for World Action Models

    L. Yang, Z. Jiang, C. Sheng, Z. Tang · 2026-08-23 · arXiv 2608.22364

    CH.08CH.18

  9. GEN-1.5: Embodied Foundation Models are One-Shot Learners

    Generalist AI · 2026-08-19 · generalistai.com

    CH.01CH.05CH.16

  10. JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

    Y. Lin, J. He, S. Bao and 6 others · 2026-08-10 · arXiv 2608.09381

    CH.07CH.18CH.19

  11. 4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields

    L. Yang, W. Song, X. Wang and 14 others · 2026-08-08 · arXiv 2608.08023

  12. DreamWAM: Beyond RGB Future Prediction for World Action Models

    S. Yuan, W. Zhao, X. Shi and 6 others · 2026-08-05 · arXiv 2608.04996

    CH.06

  13. JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations

    T. Liu, J. Zhu, T. Su and 4 others · 2026-08-04 · arXiv 2609.20277

    CH.07

  14. Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control?

    F. Lazzati, K. Stachowicz, W. Chen and 3 others · 2026-08-03 · arXiv 2608.02547

    CH.04

  15. DreamTrajectory: Trajectory-Guided Action Generation with World Model Alignment for Mobile Manipulation

    Z. Yang, W. Zhang, X. Chen and 9 others · 2026-08-02 · arXiv 2608.01381

    CH.06CH.16

  16. Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models

    Dyna Robotics · 2026-08 · dyna.co

    CH.02CH.07CH.13CH.18

  17. Introducing Gemini Robotics ER 2

    Google DeepMind · 2026-07-30 · blog.google

    CH.10CH.15CH.16CH.17CH.18

  18. Test-Time Scaling for World Action Models via Zero-Shot Geometric Evaluation

    Z. Zhao, M. Cho, H. shen and 4 others · 2026-07-20 · arXiv 2607.17454

    CH.06CH.18CH.19

  19. GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

    GigaWorld Team, A. Ye, A. Ma and 26 others · 2026-07-15 · arXiv 2607.13960

    CH.06

  20. Native Video-Action Pretraining for Generalizable Robot Control

    Q. Zhang, L. Li, L. Zhang and 26 others · 2026-07-09 · arXiv 2607.08639

    CH.06

  21. From Foundation to Application: Improving VLA Models in Practice

    W. Wu, F. Wang, F. Lu and 21 others · 2026-07-07 · arXiv 2607.06403

    CH.05CH.13CH.17CH.18CH.19

  22. Develop Humanoid Robot Policies End-to-End with NVIDIA Isaac GR00T

    NVIDIA Technical Blog · 2026-07-07 · developer.nvidia.com

    CH.02CH.15

  23. From World Models to World Action Models: A Concise Tutorial for Robotics

    X. Zhang, X. Zeng, W. Zhang · 2026-07-01 · arXiv 2607.00836

    CH.04CH.06CH.18

  24. SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors

    P. Rao, W. Zhang, R. Balestriero and 2 others · 2026-06-22 · arXiv 2606.23444

    CH.07

  25. World Action Models: A Survey

    Q. Shen, S. Zhang, Y. Liao and 5 others · 2026-06-18 · arXiv 2606.20781

    CH.01CH.06

  26. ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

    Y. Zhang, W. Zhang, Z. Qi and 7 others · 2026-06-17 · arXiv 2606.19531

    CH.06

  27. Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models

    NVIDIA Technical Blog · 2026-06-15 · developer.nvidia.com

    CH.02CH.06CH.19

  28. LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies

    J. Chen, K. Wang, K. Chen and 9 others · 2026-06-14 · arXiv 2606.15768

    CH.07CH.18CH.19

  29. NVIDIA Announces NVIDIA Isaac GR00T Reference Humanoid Robot for Academic Research

    NVIDIA Newsroom · 2026-05-31 · nvidianews.nvidia.com

    CH.02CH.16

  30. Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization

    X. Cheng, W. Yuan, Z. Mu and 5 others · 2026-05-25 · arXiv 2605.26282

    CH.08CH.18

  31. Humanoid Robotics in 2026: The Race From Pilot To Platform

    KraneShares · 2026-05-12 · kraneshares.com

  32. When to Trust Imagination: Adaptive Action Execution for World Action Models

    R. Wang, Y. Zhang, J. Lin and 4 others · 2026-05-07 · arXiv 2605.06222

    CH.06CH.18

  33. MolmoAct2: Action Reasoning Models for Real-world Deployment

    H. Fang, J. Duan, D. Clay and 26 others · 2026-05-04 · arXiv 2605.02881

    CH.05CH.10

  34. World Model for Robot Learning: A Comprehensive Survey

    B. Hou, G. Li, J. Jia and 15 others · 2026-04-30 · arXiv 2605.00080

  35. Motubrain: An Advanced World Action Model for Robot Control

    Motubrain Team, C. Xiang, F. Bao and 17 others · 2026-04-30 · arXiv 2604.27792

    CH.02CH.06

  36. Top 10 Physical AI Models Powering Real-World Robots in 2026

    MarkTechPost · 2026-04-28 · marktechpost.com

    CH.05

  37. Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models

    P. Fang, H. Chen, X. Cai · 2026-04-28 · arXiv 2604.25859

    CH.06CH.18CH.19

  38. Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment

    K. Zhou, Q. Chen, D. Peng and 3 others · 2026-04-27 · arXiv 2604.24447

    CH.02CH.05CH.15CH.17CH.18

  39. RL Token: Bootstrapping Online RL with Vision-Language-Action Models

    C. Xu, J. T. Springenberg, M. Equi and 4 others · 2026-04-24 · arXiv 2604.23073

    CH.14CH.19

  40. Three Teams, Three Robot Brains: A Head-to-Head Comparison of GR00T, Gemini, and pi Architectures

    Pebblous · 2026-04-24 · blog.pebblous.ai

  41. How VLAs (Really) Work In Open-World Environments

    A. Rasouli, Y. Wu, Z. Li and 4 others · 2026-04-23 · arXiv 2604.21192

    CH.05CH.15CH.17CH.18

  42. Lecture 24: Embodied Intelligence (CS 7643 Deep Learning, Spring 2026)

    Georgia Tech CS 7643 · 2026-04-20 · faculty.cc.gatech.edu

  43. Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models

    H. Xu, S. Zheng, H. Luo and 3 others · 2026-04-20 · arXiv 2604.18000

    CH.10CH.17CH.18

  44. NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository)

    NVIDIA · 2026-04-17 · github.com

    CH.01CH.02CH.03CH.05CH.07CH.10CH.13CH.15CH.16CH.17CH.18

  45. Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions

    A. Psiris, V. Argyriou, E. K. Markakis and 5 others · 2026-04-16 · arXiv 2604.15395

    CH.01CH.03CH.12CH.17

  46. π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

    Physical Intelligence, B. Ai, A. Amin and 85 others · 2026-04-16 · arXiv 2604.15483

    CH.01CH.02CH.03CH.05CH.06CH.14CH.16CH.18CH.19

  47. Gemini Robotics-ER 1.6

    Google DeepMind · 2026-04-14 · blog.google

    CH.10CH.16CH.18

  48. Gemini Robotics: Advancing Physical AI with Vision-Language-Action Models

    Encord · 2026-04-05 · encord.com

  49. GEN-1: Scaling Embodied Foundation Models to Mastery

    Generalist AI · 2026-04-02 · generalistai.com

    CH.05CH.11CH.13CH.18

  50. Best VLA Models 2026: Complete Vision-Language-Action Guide

    Robotics Center (SVRC) · 2026-04 · roboticscenter.ai

    CH.11

  51. Do World Action Models Generalize Better than VLAs? A Robustness Study

    Z. Zhang, Z. Li, B. Rahmati and 11 others · 2026-03-23 · arXiv 2603.22078

    CH.06CH.08CH.17CH.19

  52. Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models

    Z. Zhang, H. Ren, Y. Sun and 6 others · 2026-03-21 · arXiv 2603.20607

    CH.08CH.17

  53. GigaWorld-Policy: An Efficient Action-Centered World--Action Model

    A. Ye, B. Wang, C. Ni and 21 others · 2026-03-18 · arXiv 2603.17240

    CH.06

  54. EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards

    R. Wang, Q. Liu, Y. Deng and 3 others · 2026-03-18 · arXiv 2603.17808

    CH.08CH.17CH.18

  55. Fast-WAM: Do World Action Models Need Test-time Future Imagination?

    T. Yuan, Z. Dong, Y. Liu, H. Zhao · 2026-03-17 · arXiv 2603.16666

    CH.03CH.06CH.17CH.18CH.19

  56. DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models

    E. Y. Jia, W. Yuan, T. Shi and 3 others · 2026-03-17 · arXiv 2603.16860

    CH.08

  57. NVIDIA and Global Robotics Leaders Take Physical AI to the Real World (GTC 2026, GR00T N2 preview)

    NVIDIA Newsroom · 2026-03-16 · nvidianews.nvidia.com

    CH.03CH.05CH.06CH.16CH.19

  58. Interactive World Simulator for Robot Policy Training and Evaluation

    Y. Wang, R. Syed, F. Wu and 7 others · 2026-03-09 · arXiv 2603.08546

    CH.08CH.18

  59. Robotic Foundation Models for Industrial Control: A Comprehensive Survey and Readiness Assessment Framework

    D. Kube, S. Hadwiger, T. Meisen · 2026-03-06 · arXiv 2603.06749

    CH.01CH.12CH.15CH.16CH.17CH.18CH.19

  60. MEM: Multi-Scale Embodied Memory for Vision Language Action Models

    M. Torne, K. Pertsch, H. Walke and 14 others · 2026-03-04 · arXiv 2603.03596

    CH.14CH.17CH.18CH.19

  61. State of Robotics 2026

    SVRC (Robotics Center of Silicon Valley) · 2026-03 · roboticscenter.ai

    CH.02CH.05CH.08CH.13CH.14CH.15CH.16CH.17

  62. LeRobot: An Open-Source Library for End-to-End Robot Learning

    R. Cadene, S. Aliberts, F. Capuano and 14 others · 2026-02-26 · arXiv 2602.22818

    CH.05CH.11

  63. EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data

    R. Zheng, D. Niu, Y. Xie and 12 others · 2026-02-18 · arXiv 2602.16710

    CH.02CH.13CH.18

  64. Learning to unfold cloth: Scaling up world models to deformable object manipulation

    J. Rome, S. James, S. Ramamoorthy · 2026-02-18 · arXiv 2602.16675

    CH.09

  65. World Action Models are Zero-shot Policies

    S. Ye, Y. Ge, K. Zheng and 33 others · 2026-02-17 · arXiv 2602.15922

    CH.03CH.06

  66. ActionCodec: What Makes for Good Action Tokenizers

    Z. Dong, Y. Liu, S. Zhang and 8 others · 2026-02-17 · arXiv 2602.15397

    CH.04

  67. Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution

    R. Cai, J. Guo, X. He and 20 others · 2026-02-13 · arXiv 2602.12684

    CH.05

  68. Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models

    L. Shi, S. Chen, F. Gao and 8 others · 2026-02-13 · arXiv 2602.12628

    CH.14

  69. JEPA-VLA: Video Predictive Embedding is Needed for VLA Models

    S. Miao, N. Feng, J. Wu and 4 others · 2026-02-12 · arXiv 2602.11832

    CH.07

  70. GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning

    GigaBrain Team, B. Wang, B. Li and 23 others · 2026-02-12 · arXiv 2602.12099

    CH.14

  71. VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model

    J. Sun, W. Zhang, Z. Qi and 6 others · 2026-02-10 · arXiv 2602.10098

    CH.07CH.19

  72. WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models

    Y. Shang, Z. Li, Y. Ma and 18 others · 2026-02-09 · arXiv 2602.08971

    CH.08CH.17

  73. DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

    S. Gao, W. Liang, K. Zheng and 27 others · 2026-02-06 · arXiv 2602.06949

    CH.02CH.07CH.08CH.18

  74. World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy

    X. Liu, Z. Bai, H. Ci and 2 others · 2026-02-06 · arXiv 2602.06508

    CH.08CH.14CH.18

  75. Visuo-Tactile World Models

    C. Higuera, S. Arnaud, B. Boots and 3 others · 2026-02-05 · arXiv 2602.06001

    CH.09CH.13CH.19

  76. Causal World Modeling for Robot Control

    L. Li, Q. Zhang, Y. Luo and 9 others · 2026-01-29 · arXiv 2601.21998

    CH.03CH.06CH.15

  77. A Pragmatic VLA Foundation Model

    W. Wu, F. Lu, Y. Wang and 22 others · 2026-01-26 · arXiv 2601.18692

    CH.02CH.14

  78. Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

    M. J. Kim, Y. Gao, T. Lin and 8 others · 2026-01-22 · arXiv 2601.16163

    CH.03CH.06CH.08

  79. PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation

    W. Huang, Y. Chao, A. Mousavian and 4 others · 2026-01-07 · arXiv 2601.03782

    CH.09CH.13CH.18CH.19

  80. SOP: A Scalable Online Post-Training System for Vision-Language-Action Models

    M. Pan, S. Feng, Q. Zhang and 9 others · 2026-01-06 · arXiv 2601.03044

    CH.14

  81. RoboReward: General-Purpose Vision-Language Reward Models for Robotics

    T. Lee, A. Wagenmaker, K. Pertsch and 3 others · 2026-01-02 · arXiv 2601.00675

    CH.10

  82. Predictions Scorecard, 2026 January 01

    Rodney Brooks · 2026-01-01 · rodneybrooks.com

    CH.12CH.16CH.17CH.18CH.19

  83. Awesome-WAM: papers, explainers and resources on World Action Models

    OpenMOSS · 2026 · github.com

  84. Physical Intelligence research posts

    Physical Intelligence · 2026 · pi.website

  85. NVIDIA GR00T N2 Explained; VLA Tutorial 2026

    RoboCloud Hub · 2026 · robocloudhub.tech

    CH.12

  86. Robot Foundation Models: 2026 landscape

    Humanoid Hub · 2026 · humanoid-world.com

  87. Vision-Language-Action (VLA) Models 2026: Robotics Foundation Models and General-Purpose Robot AI

    Internet Pros · 2026 · internet-pros.com

  88. Gemini Robotics (Wikipedia article)

    Wikipedia · 2026 · en.wikipedia.org

2025 57

  1. Act2Goal: From World Model To General Goal-conditioned Policy

    P. Zhou, L. Chen, S. Chen and 5 others · 2025-12-29 · arXiv 2512.23541

    CH.06

  2. Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation

    T. Zou, H. Zeng, Y. Nong and 6 others · 2025-12-23 · arXiv 2512.20188

    CH.05

  3. mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs

    J. Pai, L. Achenbach, V. Montesinos and 3 others · 2025-12-17 · arXiv 2512.15692

    CH.06

  4. Motus: A Unified Latent Action World Model

    H. Bi, H. Tan, S. Xie and 13 others · 2025-12-15 · arXiv 2512.13030

    CH.06

  5. An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges

    C. Xu, S. Zhang, Y. Liu and 11 others · 2025-12-12 · arXiv 2512.11362

    CH.01CH.03CH.05

  6. Learning Robot Manipulation from Audio World Models

    F. Zhang, M. Gienger · 2025-12-09 · arXiv 2512.08405

    CH.09CH.13

  7. Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation

    M. Wei, C. Wan, J. Peng and 8 others · 2025-12-09 · arXiv 2512.08186

    CH.16

  8. GigaWorld-0: World Models as Data Engine to Empower Embodied AI

    GigaWorld Team, A. Ye, B. Wang and 22 others · 2025-11-25 · arXiv 2511.19861

    CH.08

  9. π*0.6: a VLA That Learns From Experience

    Physical Intelligence, A. Amin, R. Aniceto and 53 others · 2025-11-18 · arXiv 2511.14759

    CH.03CH.14CH.19

  10. WMPO: World Model-based Policy Optimization for Vision-Language-Action Models

    F. Zhu, Z. Yan, Z. Hong and 3 others · 2025-11-12 · arXiv 2511.09515

    CH.08

  11. LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics

    R. Balestriero, Y. LeCun · 2025-11-11 · arXiv 2511.08544

    CH.07CH.18

  12. ViPRA: Video Prediction for Robot Actions

    S. Routray, H. Pan, U. Jain and 2 others · 2025-11-11 · arXiv 2511.07732

    CH.07CH.18

  13. SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

    Z. Luo, Y. Yuan, T. Wang and 25 others · 2025-11-11 · arXiv 2511.07820

    CH.07CH.13CH.16CH.17CH.18

  14. Robot Learning from a Physical World Model

    J. Mao, S. He, H. Wu and 9 others · 2025-11-10 · arXiv 2511.07416

    CH.09

  15. GEN-0: Embodied Foundation Models That Scale with Physical Interaction

    Generalist AI · 2025-11-04 · generalistai.com

    CH.03CH.11CH.13

  16. World Simulation with Video Foundation Models for Physical AI

    NVIDIA, A. Ali, J. Bai and 86 others · 2025-10-28 · arXiv 2511.00062

    CH.08

  17. GigaBrain-0: A World Model-Powered Vision-Language-Action Model

    GigaBrain Team, A. Ye, B. Wang and 24 others · 2025-10-22 · arXiv 2510.19430

    CH.08

  18. LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

    S. Fei, S. Wang, J. Shi and 10 others · 2025-10-15 · arXiv 2510.13626

    CH.15CH.16CH.17CH.18

  19. Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

    Gemini Robotics Team, A. Abdolmaleki, S. Abeyruwan and 169 others · 2025-10-02 · arXiv 2510.03342

    CH.10

  20. Video models are zero-shot learners and reasoners

    T. Wiedemer, Y. Li, P. Vicol and 6 others · 2025-09-24 · arXiv 2509.20328

    CH.02

  21. World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation

    Z. Jiang, K. Liu, Y. Qin and 6 others · 2025-09-23 · arXiv 2509.19080

    CH.08

  22. GWM: Towards Scalable Gaussian World Models for Robotic Manipulation

    G. Lu, B. Jia, P. Li and 4 others · 2025-08-25 · arXiv 2508.17600

    CH.09CH.18

  23. GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

    L. Sun, B. Xie, Y. Liu and 3 others · 2025-08-12 · arXiv 2508.09071

    CH.09

  24. A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation

    TRI LBM Team, J. Barreiros, A. Beaulieu and 79 others · 2025-07-07 · arXiv 2507.05331

    CH.11CH.12CH.13CH.15CH.18

  25. A Survey on Vision-Language-Action Models: An Action Tokenization Perspective

    Y. Zhong, F. Bai, S. Cai and 11 others · 2025-07-02 · arXiv 2507.01925

    CH.04CH.05

  26. ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation

    S. Huang, Q. Chen, X. Zhang and 2 others · 2025-06-29 · arXiv 2506.23126

    CH.09

  27. WorldVLA: Towards Autoregressive Action World Model

    J. Cen, C. Yu, H. Yuan and 9 others · 2025-06-26 · arXiv 2506.21539

    CH.06

  28. Gemini Robotics On-Device brings AI to local robotic devices

    Google DeepMind · 2025-06-24 · deepmind.google

    CH.05CH.14CH.15

  29. ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model

    T. Yu, G. Lu, Z. Yang and 7 others · 2025-06-24 · arXiv 2506.19842

    CH.09

  30. RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

    T. Chen, Z. Chen, B. Chen and 23 others · 2025-06-22 · arXiv 2506.18088

  31. RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies

    P. Atreya, K. Pertsch, T. Lee and 29 others · 2025-06-22 · arXiv 2506.18123

    CH.06CH.15CH.18CH.19

  32. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

    M. Assran, A. Bardes, D. Fan and 26 others · 2025-06-11 · arXiv 2506.09985

    CH.07CH.08CH.13

  33. Real-Time Execution of Action Chunking Flow Policies

    K. Black, M. Y. Galliker, S. Levine · 2025-06-09 · arXiv 2506.07339

    CH.02CH.04CH.05CH.11CH.15CH.18

  34. SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

    M. Shukor, D. Aubakirova, F. Capuano and 11 others · 2025-06-02 · arXiv 2506.01844

    CH.01CH.05

  35. Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review

    M. Lisondra, B. Benhabib, G. Nejat · 2025-05-26 · arXiv 2505.20503

    CH.16

  36. DreamGen: Unlocking Generalization in Robot Learning through Video World Models

    J. Jang, S. Ye, Z. Lin and 25 others · 2025-05-19 · arXiv 2505.12705

    CH.08CH.18

  37. EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video

    R. Hoque, P. Huang, D. J. Yoon and 2 others · 2025-05-16 · arXiv 2505.11709

    CH.02CH.13

  38. OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation

    C. Cui, P. Ding, W. Song and 10 others · 2025-05-06 · arXiv 2505.03912

    CH.05

  39. PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation

    W. Li, H. Zhao, Z. Yu and 4 others · 2025-04-23 · arXiv 2504.16693

    CH.09

  40. π0.5: a Vision-Language-Action Model with Open-World Generalization

    Physical Intelligence, K. Black, N. Brown and 33 others · 2025-04-22 · arXiv 2504.16054

    CH.02CH.03CH.05CH.10

  41. Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets

    C. Zhu, R. Yu, S. Feng and 3 others · 2025-04-03 · arXiv 2504.02792

    CH.04CH.06

  42. Wan: Open and Advanced Large-Scale Video Generative Models

    Team Wan, A. Wang, B. Ai and 59 others · 2025-03-26 · arXiv 2503.20314

    CH.02CH.06

  43. GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving

    L. Russell, A. Hu, L. Bertoni and 4 others · 2025-03-26 · arXiv 2503.20523

    CH.16

  44. Gemini Robotics: Bringing AI into the Physical World

    Gemini Robotics Team, S. Abeyruwan, J. Ainslie and 115 others · 2025-03-25 · arXiv 2503.20020

    CH.03CH.05CH.10CH.15

  45. Gemma 3 Technical Report

    Gemma Team, A. Kamath, J. Ferret and 212 others · 2025-03-25 · arXiv 2503.19786

    CH.02

  46. AdaWorld: Learning Adaptable World Models with Latent Actions

    S. Gao, S. Zhou, Y. Du and 2 others · 2025-03-24 · arXiv 2503.18938

  47. GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

    NVIDIA, J. Bjorck, F. Castañeda and 39 others · 2025-03-18 · arXiv 2503.14734

    CH.03CH.05

  48. Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control

    NVIDIA, H. A. Alhaija, J. Alvarez and 37 others · 2025-03-18 · arXiv 2503.14492

    CH.08CH.18

  49. Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

    NVIDIA, A. Azzolini, J. Bai and 50 others · 2025-03-18 · arXiv 2503.15558

    CH.10

  50. NVIDIA Announces Isaac GR00T N1, the World's First Open Humanoid Robot Foundation Model, and Simulation Frameworks to Speed Robot Development

    NVIDIA Newsroom · 2025-03-18 · nvidianews.nvidia.com

    CH.03CH.08CH.13

  51. Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

    M. J. Kim, C. Finn, P. Liang · 2025-02-27 · arXiv 2502.19645

    CH.05CH.18

  52. Helix: A Vision-Language-Action Model for Generalist Humanoid Control

    Figure AI · 2025-02-20 · figure.ai

    CH.03CH.05CH.16CH.17

  53. SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

    D. Qu, H. Song, Q. Chen and 8 others · 2025-01-27 · arXiv 2501.15830

    CH.09

  54. FAST: Efficient Action Tokenization for Vision-Language-Action Models

    K. Pertsch, K. Stachowicz, B. Ichter and 6 others · 2025-01-16 · arXiv 2501.09747

    CH.04CH.05

  55. Cosmos World Foundation Model Platform for Physical AI

    NVIDIA, N. Agarwal, A. Ali and 75 others · 2025-01-07 · arXiv 2501.03575

    CH.02CH.13

  56. EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation

    S. Huang, L. Chen, P. Zhou and 8 others · 2025-01-03 · arXiv 2501.01895

    CH.09

  57. GR00T-Dreams (DreamGen) repository

    NVIDIA · 2025 · github.com

2024 20

  1. Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

    Y. Hu, Y. Guo, P. Wang and 6 others · 2024-12-19 · arXiv 2412.14803

    CH.06

  2. Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

    Y. Tian, S. Yang, J. Zeng and 4 others · 2024-12-19 · arXiv 2412.15109

    CH.06

  3. Navigation World Models

    A. Bar, G. Zhou, D. Tran and 2 others · 2024-12-04 · arXiv 2412.03572

    CH.16

  4. Prediction with Action: Visual Policy Learning via Joint Denoising Process

    Y. Guo, Y. Hu, J. Zhang and 4 others · 2024-11-27 · arXiv 2411.18179

    CH.06

  5. VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation

    Y. Wen, J. Lin, Y. Zhu and 4 others · 2024-11-14 · arXiv 2411.09153

  6. DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

    G. Zhou, H. Pan, Y. LeCun, L. Pinto · 2024-11-07 · arXiv 2411.04983

    CH.07

  7. π0: A Vision-Language-Action Flow Model for General Robot Control

    K. Black, N. Brown, D. Driess and 21 others · 2024-10-31 · arXiv 2410.24164

    CH.03CH.04CH.05CH.11

  8. X-MOBILITY: End-To-End Generalizable Navigation via World Modeling

    W. Liu, H. Zhao, C. Li and 5 others · 2024-10-23 · arXiv 2410.17491

    CH.16

  9. Latent Action Pretraining from Videos

    S. Ye, J. Jang, B. Jeon and 13 others · 2024-10-15 · arXiv 2410.11758

    CH.04CH.07CH.18

  10. ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI

    S. Tao, F. Xiang, A. Shukla and 20 others · 2024-10-01 · arXiv 2410.00425

  11. PaliGemma: A versatile 3B VLM for transfer

    L. Beyer, A. Steiner, A. S. Pinto and 32 others · 2024-07-10 · arXiv 2407.07726

    CH.02

  12. NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking

    D. Dauner, M. Hallgarten, T. Li and 9 others · 2024-06-21 · arXiv 2406.15349

    CH.16

  13. OpenVLA: An Open-Source Vision-Language-Action Model

    M. J. Kim, K. Pertsch, S. Karamcheti and 15 others · 2024-06-13 · arXiv 2406.09246

    CH.03CH.04CH.05CH.14

  14. Octo: An Open-Source Generalist Robot Policy

    Octo Model Team, D. Ghosh, H. Walke and 16 others · 2024-05-20 · arXiv 2405.12213

    CH.01CH.03CH.11CH.18

  15. What Foundation Models can Bring for Robot Learning in Manipulation : A Survey

    D. Li, Y. Jin, Y. Sun and 11 others · 2024-04-28 · arXiv 2404.18201

    CH.01

  16. DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

    A. Khazatsky, K. Pertsch, S. Nair and 98 others · 2024-03-19 · arXiv 2403.12945

    CH.02CH.06CH.13

  17. Covariant introduces RFM-1 to give robots the human-like ability to reason

    Covariant · 2024-03-11 · covariant.ai

    CH.03CH.16

  18. Genie: Generative Interactive Environments

    J. Bruce, M. Dennis, A. Edwards and 22 others · 2024-02-23 · arXiv 2402.15391

    CH.07CH.08

  19. Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots

    C. Chi, Z. Xu, C. Pan and 5 others · 2024-02-15 · arXiv 2402.10329

    CH.02CH.13

  20. SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning

    J. Luo, Z. Hu, C. Xu and 7 others · 2024-01-29 · arXiv 2401.16013

    CH.14

2023 14

  1. Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis

    Y. Hu, Q. Xie, V. Jain and 20 others · 2023-12-14 · arXiv 2312.08782

    CH.01

  2. OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving

    W. Zheng, W. Chen, Y. Huang and 3 others · 2023-11-27 · arXiv 2311.16038

    CH.16

  3. Robot Learning in the Era of Foundation Models: A Survey

    X. Xiao, J. Liu, Z. Wang and 5 others · 2023-11-24 · arXiv 2311.14379

    CH.01

  4. Vision-Language Foundation Models as Effective Robot Imitators

    X. Li, M. Liu, H. Zhang and 9 others · 2023-11-02 · arXiv 2311.01378

    CH.03

  5. Open X-Embodiment: Robotic Learning Datasets and RT-X Models

    Open X-Embodiment Collaboration, A. O'Neill, A. Rehman and 291 others · 2023-10-13 · arXiv 2310.08864

    CH.02CH.03CH.05CH.13CH.18

  6. GAIA-1: A Generative World Model for Autonomous Driving

    A. Hu, L. Russell, H. Yeo and 5 others · 2023-09-29 · arXiv 2309.17080

    CH.16

  7. DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

    X. Wang, Z. Zhu, G. Huang and 3 others · 2023-09-18 · arXiv 2309.09777

    CH.16

  8. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

    J. Bai, S. Bai, S. Yang and 6 others · 2023-08-24 · arXiv 2308.12966

    CH.02

  9. BridgeData V2: A Dataset for Robot Learning at Scale

    H. Walke, K. Black, A. Lee and 11 others · 2023-08-24 · arXiv 2308.12952

    CH.02CH.13

  10. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

    A. Brohan, N. Brown, J. Carbajal and 51 others · 2023-07-28 · arXiv 2307.15818

    CH.03CH.04CH.05

  11. LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

    B. Liu, Y. Zhu, C. Gao and 4 others · 2023-06-05 · arXiv 2306.03310

  12. Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

    T. Z. Zhao, V. Kumar, S. Levine, C. Finn · 2023-04-23 · arXiv 2304.13705

    CH.02CH.03CH.04CH.11CH.16

  13. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

    C. Chi, Z. Xu, S. Feng and 5 others · 2023-03-07 · arXiv 2303.04137

    CH.03CH.04CH.11CH.16

  14. Mastering Diverse Domains through World Models

    D. Hafner, J. Pasukonis, J. Ba, T. Lillicrap · 2023-01-10 · arXiv 2301.04104

    CH.03CH.08

Before 2023 11

  1. RT-1: Robotics Transformer for Real-World Control at Scale

    A. Brohan, N. Brown, J. Carbajal and 48 others · 2022-12-13 · arXiv 2212.06817

    CH.03

  2. Temporal Difference Learning for Model Predictive Control

    N. Hansen, X. Wang, H. Su · 2022-03-09 · arXiv 2203.04955

    CH.03

  3. CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks

    O. Mees, L. Hermann, E. Rosete-Beas, W. Burgard · 2021-12-06 · arXiv 2112.03227

  4. Ego4D: Around the World in 3,000 Hours of Egocentric Video

    K. Grauman, A. Westbury, E. Byrne and 82 others · 2021-10-13 · arXiv 2110.07058

    CH.13

  5. ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

    T. Mu, Z. Ling, F. Xiang and 6 others · 2021-07-30 · arXiv 2107.14483

  6. LoRA: Low-Rank Adaptation of Large Language Models

    E. J. Hu, Y. Shen, P. Wallis and 5 others · 2021-06-17 · arXiv 2106.09685

    CH.14

  7. Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments

    J. Krantz, E. Wijmans, A. Majumdar and 2 others · 2020-04-06 · arXiv 2004.02857

    CH.16

  8. Dream to Control: Learning Behaviors by Latent Imagination

    D. Hafner, T. Lillicrap, J. Ba, M. Norouzi · 2019-12-03 · arXiv 1912.01603

    CH.03CH.04CH.08

  9. nuScenes: A multimodal dataset for autonomous driving

    H. Caesar, V. Bankiti, A. H. Lang and 7 others · 2019-03-26 · arXiv 1903.11027

    CH.16

  10. Learning Latent Dynamics for Planning from Pixels

    D. Hafner, T. Lillicrap, I. Fischer and 4 others · 2018-11-12 · arXiv 1811.04551

    CH.03CH.04CH.08

  11. Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments

    P. Anderson, Q. Wu, D. Teney and 6 others · 2017-11-20 · arXiv 1711.07280

    CH.16