Sources
Sources
Every paper, report and page in the book’s reference collection, newest first, with the chapters that cite it.
- In the collection
- 190
- Papers
- 159
- Web and reports
- 31
- Cited in chapters
- 170
local copies in reference/
linked to arXiv
archived pages
the rest is background
2026 88
ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
World-Action Models for Robot Learning and Control: A Survey
CH.01CH.03CH.04CH.05CH.06CH.09CH.12CH.13CH.14CH.16CH.17CH.18CH.19
Awesome-World-Action-Models (project page and paper list of the WAM survey)
Motus2: A Self-Evolving General World Model for Dexterous Manipulation
JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling
4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields
DreamWAM: Beyond RGB Future Prediction for World Action Models
Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control?
Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models
Test-Time Scaling for World Action Models via Zero-Shot Geometric Evaluation
GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch
Native Video-Action Pretraining for Generalizable Robot Control
From Foundation to Application: Improving VLA Models in Practice
Develop Humanoid Robot Policies End-to-End with NVIDIA Isaac GR00T
From World Models to World Action Models: A Concise Tutorial for Robotics
SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
NVIDIA Announces NVIDIA Isaac GR00T Reference Humanoid Robot for Academic Research
Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization
When to Trust Imagination: Adaptive Action Execution for World Action Models
MolmoAct2: Action Reasoning Models for Real-world Deployment
Top 10 Physical AI Models Powering Real-World Robots in 2026
Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
Three Teams, Three Robot Brains: A Head-to-Head Comparison of GR00T, Gemini, and pi Architectures
Lecture 24: Embodied Intelligence (CS 7643 Deep Learning, Spring 2026)
Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models
π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
Gemini Robotics: Advancing Physical AI with Vision-Language-Action Models
Do World Action Models Generalize Better than VLAs? A Robustness Study
Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models
GigaWorld-Policy: An Efficient Action-Centered World--Action Model
EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards
Fast-WAM: Do World Action Models Need Test-time Future Imagination?
DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models
NVIDIA and Global Robotics Leaders Take Physical AI to the Real World (GTC 2026, GR00T N2 preview)
Interactive World Simulator for Robot Policy Training and Evaluation
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
LeRobot: An Open-Source Library for End-to-End Robot Learning
EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data
Learning to unfold cloth: Scaling up world models to deformable object manipulation
Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models
JEPA-VLA: Video Predictive Embedding is Needed for VLA Models
GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy
Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
SOP: A Scalable Online Post-Training System for Vision-Language-Action Models
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
Awesome-WAM: papers, explainers and resources on World Action Models
Vision-Language-Action (VLA) Models 2026: Robotics Foundation Models and General-Purpose Robot AI
2025 57
Act2Goal: From World Model To General Goal-conditioned Policy
Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
GigaWorld-0: World Models as Data Engine to Empower Embodied AI
WMPO: World Model-based Policy Optimization for Vision-Language-Action Models
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
GEN-0: Embodied Foundation Models That Scale with Physical Interaction
World Simulation with Video Foundation Models for Physical AI
GigaBrain-0: A World Model-Powered Vision-Language-Action Model
LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation
Gemini Robotics On-Device brings AI to local robotic devices
ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
DreamGen: Unlocking Generalization in Robot Learning through Video World Models
EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation
π0.5: a Vision-Language-Action Model with Open-World Generalization
Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
AdaWorld: Learning Adaptable World Models with Latent Actions
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
Helix: A Vision-Language-Action Model for Generalist Humanoid Control
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
FAST: Efficient Action Tokenization for Vision-Language-Action Models
EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
2024 20
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation
Prediction with Action: Visual Policy Learning via Joint Denoising Process
VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
π0: A Vision-Language-Action Flow Model for General Robot Control
X-MOBILITY: End-To-End Generalizable Navigation via World Modeling
ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI
What Foundation Models can Bring for Robot Learning in Manipulation : A Survey
Covariant introduces RFM-1 to give robots the human-like ability to reason
Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots
SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning
2023 14
Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis
OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving
Vision-Language Foundation Models as Effective Robot Imitators
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
Before 2023 11
ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations
Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments