Ch.05Part II · The seven familiesReactive VLAs
Reactive vision-language-action models
A pretrained VLM backbone plus an action head: observations and language in, an action chunk out, no modelling of the future.
- Math
- L1
- Sources
- 31
- Figures
- 4
- Exercises
- 3
- Last checked
- 3 Oct 2026
One boxed equation per idea
8 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
How the head works. a separate, smaller network turns noise into the whole chunk in K steps. The cost depends on the number of denoising steps, not on the chunk length. Continuous actions, many modes, smooth motion.
Examples: π0 (October 2024), GR00T N1 (March 2025), SmolVLA (June 2025), Xiaomi-Robotics-0 (February 2026), π0.7 and GR00T N1.7 (April 2026).
- Sequential passes for one chunk
- 10 expert passes50 steps × 14 dimensions, after the backbone reads the inputs once; π0 uses 10 steps
- What crosses into the head
- the backbone's features, read by the expert through attention
Why this chapter exists
The models that made "robot foundation model" a household phrase in robotics labs are almost all of one kind. Take a vision-language model (VLM) that has learned from the web what a mug is and what "put it in the sink" means, and teach it to output robot actions. Observations and an instruction go in; a short chunk of actions comes out. Nothing in the model predicts what the world will look like next. This book calls that family reactive vision-language-action models, or VLAs.
They are the reference point for everything in Part II. World-action models (chapter 6) were invented to fix what VLAs get wrong; large behavior models (chapter 11) drop the language backbone VLAs depend on; embodied reasoning models (chapter 10) sit above a VLA and tell it what to do. Knowing exactly what a VLA does, and does not do, makes the other families easy to place.
After this chapter you will be able to:
- name the three sub-families of VLA, explain with figure 5.1 what each costs per action chunk, and place a new model in one of them from its paper;
- explain with the simulation in figure 5.2 why the fastest humanoid VLAs split into a slow reasoner and a fast controller, and what that split cannot fix;
- read an open-model announcement critically: what exactly is released, under what terms, and on what hardware it runs.
At Ai2, our mission is to accelerate robotics through open science.
The mechanism
In the notation of chapter 4, a reactive VLA is a function from the recent observations and the instruction to a chunk of future actions:
Read it as: "given what I have just seen and what I was asked, here is what to do for the next steps." The observation is camera images plus the robot's joint angles. The output is a chunk because, as chapter 4 showed, a model that is slower than the robot's control loop has to hand over several actions at once. And there is no on the left: a reactive VLA never says what it expects to happen. That single absence is what separates this chapter from the next.
Inside are always two parts: a pretrained VLM backbone that reads the images, the instruction and the robot state as one sequence of tokens, and a head that turns what the backbone computed into numbers a motor controller can use. The backbones are ordinary open VLMs: PaLI-X and PaLM-E for RT-2, Llama 2 for OpenVLA, PaliGemma for π0 and π0.5, Gemma 3 for π0.7, and Cosmos-Reason2 for GR00T N1.7 (Brohan et al. 2023, arXiv 2307.15818; Kim et al. 2024, arXiv 2406.09246; Black et al. 2024, arXiv 2410.24164; Physical Intelligence et al. 2025, arXiv 2504.16054; Physical Intelligence et al. 2026, arXiv 2604.15483; NVIDIA 2026). The heads are where the families differ.
Three ways to get actions out
Action tokens
The first VLAs reused the language model's own output. RT-2 and OpenVLA divide each action dimension into 256 bins and treat each bin as a word, so "move the wrist 2 cm left" becomes a short string of tokens the model writes one after another (Brohan et al. 2023, arXiv 2307.15818; Kim et al. 2024, arXiv 2406.09246). The appeal is that nothing in the language model changes. The cost is in the sequence: every token needs its own pass through the decoder, so a chunk of steps for a robot with action dimensions needs
For one second of a two-armed robot at 50 Hz, that is 50 × 14 = 700 passes, which is why the first token VLAs predicted a single step at a time and ran at a few hertz. RT-2's largest, 55-billion-parameter model ran at 1 to 3 Hz from a cloud service (Brohan et al. 2023, arXiv 2307.15818).
Two fixes followed. FAST compresses a chunk before tokenizing it, using the discrete cosine transform also used in JPEG images: smooth motion needs only a few low-frequency coefficients, so the chunk becomes a much shorter token sequence. Its authors report matching diffusion-based VLAs while training up to five times faster, and released FAST+, a tokenizer trained on one million real robot trajectories (Pertsch et al. 2025, arXiv 2501.09747). Exercise 5.2 builds the compression step. OpenVLA-OFT went further: it decodes all actions in one parallel pass and switches to continuous outputs, raising OpenVLA's average success on the LIBERO simulation benchmark from 76.5 to 97.1 percent while generating actions 26 times faster (Kim et al. 2025, arXiv 2502.19645). At that point the "token" family has quietly moved towards the next one. A 2025 survey of action tokenization maps the full range of options (Zhong et al. 2025, arXiv 2507.01925).
An action expert that denoises
The second sub-family keeps the backbone for understanding and adds a separate, smaller network for acting. π0 attached a 300-million-parameter "action expert" to a 3-billion-parameter VLM; the expert reads the backbone's features through attention and produces a whole chunk by flow matching (Black et al. 2024, arXiv 2410.24164). Chapter 4 described flow matching operationally: start from random noise, and move it towards a plausible chunk in small steps,
where is the chunk at denoising time , is the expert's prediction of which way to move it, and is the step size. π0 uses 10 steps (Black et al. 2024, arXiv 2410.24164). The cost is passes of the small expert, independent of how long the chunk is or how many joints the robot has. That is why this design reached 50 Hz control on dexterous tasks such as laundry folding (Black et al. 2024, arXiv 2410.24164), and why most VLAs released since use it: GR00T N1 and N1.7, SmolVLA, Xiaomi-Robotics-0, π0.5 and π0.7 (NVIDIA et al. 2025, arXiv 2503.14734; Shukor et al. 2025, arXiv 2506.01844; Cai et al. 2026, arXiv 2602.12684; Physical Intelligence et al. 2025, arXiv 2504.16054; Physical Intelligence et al. 2026, arXiv 2604.15483). Continuous output also avoids the averaging trap from chapter 4: when a demonstration set contains two good ways to do something, a denoising head can produce either, not their useless mean.
Two models on two clocks
The third sub-family splits the model in two. Figure's Helix runs a 7-billion-parameter VLM as "System 2" at 7 to 9 Hz, and an 80-million-parameter "System 1" at 200 Hz that turns System 2's latest latent vector and fresh camera images into motor commands (Figure AI 2025). GR00T N1 uses the same vocabulary: its System 2 VLM runs at 10 Hz on an NVIDIA L40 and its System 1 diffusion transformer produces closed-loop actions at 120 Hz (NVIDIA et al. 2025, arXiv 2503.14734). Gemini Robotics splits along a different line, with a backbone in the cloud and an action decoder on the robot's own computer (Gemini Robotics Team et al. 2025, arXiv 2503.20020). Academic versions followed: OpenHelix surveyed and released an open dual-system model (Cui et al. 2025, arXiv 2505.03912), and DuoCore-FS bridges a 1 to 3 Hz VLM and a 25 to 30 Hz action generator through a buffer of latent representations (Zou et al. 2025, arXiv 2512.20188).
What crosses the interface matters. In all of these designs it is not a word or a plan but a latent vector: a compressed summary of "what to do" that the fast model can condition on many times before the slow model updates it. Figure 5.2 shows what the split buys and what it does not.
- Tracking error, before the switch
- 0.004mean distance to the named cup; set by the fast loop
- Reaction to the new instruction
- 0.54 suntil the hand stays within 0.1 of cup B; set by the slow loop
The simulation makes two separate costs visible. With one model at 8 Hz, the hand chases where the cup was when the model last looked, and it is always behind: the tracking error before the switch is around 0.1 table units. Add a 200 Hz controller that re-reads the cup's position every 5 ms, and the error drops about twenty-five-fold. But look at what happens when the instruction changes. The fast controller cannot know which cup to follow until the slow model has looked at the new instruction and thought about it, so the reaction time is set by the slow rate. At 2 Hz the robot keeps confidently tracking the wrong cup for up to a second. Exercise 5.3 measures it.
That is the honest summary of dual-system designs: they make the robot reactive to the world, not to new meaning. A real System 2 can of course be faster or slower than in the toy, and real System 1 models do more than track a point. The structure of the trade-off is what carries over.
Why it exists
Two things made this family the first to work at scale. The first is transfer: a VLM that already knows objects, spatial words and some common sense brings that knowledge to robot control. RT-2's central result was that co-training on web-scale vision-language data and robot trajectories markedly improved generalization to new objects and to commands absent from the robot data (Brohan et al. 2023, arXiv 2307.15818). The second is supervision: teleoperated demonstrations are the cheapest signal that reaches real hardware. They are expensive by web standards, about $118 an hour in early 2026 according to one industry surveyindustry survey (SVRC (Robotics Center of Silicon Valley) 2026), but they are labelled with exactly the actions a VLA needs. Datasets such as Open X-Embodiment pooled them (Open X-Embodiment Collaboration et al. 2023, arXiv 2310.08864), and a VLA is the simplest model that can learn from them.
Representative models
| Model | First released | Sub-family | Notes |
|---|---|---|---|
| RT-2 | July 2023 | tokens | up to 55B parameters, PaLI-X backbone (Brohan et al. 2023, arXiv 2307.15818) |
| OpenVLA | June 2024 | tokens | 7B, Llama 2 backbone, trained on 970,000 episodes (Kim et al. 2024, arXiv 2406.09246) |
| π0 | October 2024 | action expert | 3.3B, 50 Hz control (Black et al. 2024, arXiv 2410.24164) |
| FAST and π0-FAST | January 2025 | tokens, compressed | DCT tokenizer; FAST+ released (Pertsch et al. 2025, arXiv 2501.09747) |
| Helix | February 2025 | fast-slow | 7B over 80M, 7 to 9 Hz over 200 Hz (Figure AI 2025) |
| OpenVLA-OFT | February 2025 | parallel decoding | 26× faster action generation than OpenVLA (Kim et al. 2025, arXiv 2502.19645) |
| GR00T N1 | March 2025 | fast-slow, action expert | 2.2B released checkpoint (NVIDIA et al. 2025, arXiv 2503.14734) |
| Gemini Robotics | March 2025 | fast-slow, cloud and robot | on-device version June 2025 (Gemini Robotics Team et al. 2025, arXiv 2503.20020; Google DeepMind 2025) |
| π0.5 | April 2025 | action expert | co-training for open-world generalization (Physical Intelligence et al. 2025, arXiv 2504.16054) |
| SmolVLA | June 2025 | action expert | under 0.5B, trains on one GPU (Shukor et al. 2025, arXiv 2506.01844) |
| Xiaomi-Robotics-0 | February 2026 | action expert | 4.7B, asynchronous execution, open (Cai et al. 2026, arXiv 2602.12684) |
| GEN-1 | April 2026 | no VLM backbone (chapter 11) | 99% on simple tasks where earlier models reach 64%vendor claim (Generalist AI 2026) |
| π0.7 | April 2026 | action expert | about 5B, steerable prompt (figure 5.4) (Physical Intelligence et al. 2026, arXiv 2604.15483) |
| GR00T N1.7 | April 2026 | fast-slow, action expert | 3B, Apache 2.0 (NVIDIA 2026) |
| MolmoAct2 | May 2026 | action reasoning, flow expert | fully open, including datasets (Fang et al. 2026, arXiv 2605.02881) |
| LingBot-VLA 2.0 | July 2026 | action expert | about 60,000 hours of pretraining data (Wu et al. 2026, arXiv 2607.06403) |
| GEN-1.5 | August 2026 | no VLM backbone (chapter 11) | 59% one-shot success across 10 tasksvendor claim (Generalist AI 2026) |
Two entries show where the family is heading. LingBot-VLA 2.0 adds future prediction as a training-only proxy task, helped by a video representation model for semantics and a depth estimator for geometry (Wu et al. 2026, arXiv 2607.06403). That is close to the "video only in training" style of chapter 6's figure 6.2, arriving inside a model its makers still call a VLA. GEN-1.5 reports learning a new task in context from a single 3 to 12 second demonstration, at 59 percent average success across 10 tasks, and 83 percent after 10 gradient steps on five minutes of datavendor claim (Generalist AI 2026). The GEN models sit in this table because they are usually compared with VLAs, but they are trained from scratch on their makers' own data rather than on a pretrained VLM; chapter 11 discusses them with the other large behavior models.
The open ecosystem
No other family has as much released code and weights. LeRobot provides an open library and training stack (Cadene et al. 2026, arXiv 2602.22818), NVIDIA publishes GR00T checkpoints with fine-tuning scripts (NVIDIA 2026), and Physical Intelligence open-sourced π0 in February 2025 (MarkTechPost 2026). "Open" covers very different things, though, and figure 5.3 separates them.
GR00T N1.7
3B parameters, on a Cosmos-Reason2-2B backbone. Open weights: weights under Apache 2.0, fully commercially licensable; the Cosmos-Reason2-2B backbone it loads is gated and needs an access request.
April 2026. Source: NVIDIA Isaac-GR00T repository.
Three readings of the figure are worth making explicit. First, open models now cover the whole range from SmolVLA, designed to be deployed on consumer GPUs or even CPUs, to 7-billion-parameter OpenVLA (Shukor et al. 2025, arXiv 2506.01844; Kim et al. 2024, arXiv 2406.09246). Second, "open" is layered. GR00T N1.7's weights are Apache 2.0 and commercially licensable, but the Cosmos-Reason2 backbone it loads is gated behind an access request (NVIDIA 2026). MolmoAct2's authors argue that the few open frontier models "release weights alone", and release datasets and recipes as well (Fang et al. 2026, arXiv 2605.02881). Third, the most capable commercial models sit in the bottom row. Helix, Gemini Robotics and π0.7 are described in posts and papers, not downloadable.
Key paper: steering with the prompt
how: segment metadata. Whether a segment contains a mistake, such as a failed grasp or the wrong subtask. Annotated coarsely by humans.
Toy model: 100 collected episodes, some with mistakes. A generative policy imitates the episodes that match its prompt.
- Episodes trained on
- 100 of 100
- Imitated behaviour that is a mistake
- 4.5%among episodes matching the prompt
The toy states the argument in its simplest form. Throwing away flawed episodes gives the same low mistake rate as prompting with "Mistake: false", but it trains on less data. In real models the kept failures still teach something: what the objects look like, how the scene responds to a fumble, how to recover. The toy also shows the dependency nobody can prompt away: when annotators label mistakes inaccurately, the mistakes leak into the "good" part of the data, whatever the recipe.
The π0.7 prompt also blurs this chapter's boundary. Subgoal images are pictures of the future. The WAM survey reads goal-image conditioning as a goal-conditioned world-action model (Lu et al. 2026, arXiv 2609.16074); this book keeps π0.7 here because at inference the model itself does not generate those futures. Chapter 1's family grid calls it a boundary case for that reason.
What it is good at
Deployment. VLAs are the most documented, most packaged family. A team can download a model, fine-tune it on a few hundred demonstrations and run it on a single GPU. One industry survey reports that fine-tuning a pretrained VLA on 200 to 500 demonstrations now beats training a task-specific policy from scratch on more than 1,000industry survey (SVRC (Robotics Center of Silicon Valley) 2026).
Speed. Chunking, action experts and asynchronous execution make control rates achievable. Real-time chunking generates the next chunk while the current one runs, without retraining (Black et al. 2025, arXiv 2506.07339), and a 2026 study across accelerators finds that right-sized edge devices can meet control rates at lower cost and energy than flagship GPUs (Zhou et al. 2026, arXiv 2604.24447).
Language and objects. Instructions about unseen objects and places transfer from the backbone, which is the reason the family exists.
What it is bad at
Consequences. A reactive VLA has no internal model of what its actions will do. It cannot notice that a locally sensible action makes a later step impossible, which is the gap chapter 6 opened with.
The largest data. A VLA learns from action-labelled robot data. Internet video and most human video carry no robot actions, so a VLA cannot learn from the two largest layers of chapter 2's data pyramid without extra machinery: latent actions (chapter 7), hand tracking, or the training-time video prediction LingBot-VLA 2.0 now adds (Wu et al. 2026, arXiv 2607.06403).
Composition. Combining known skills into new sequences was historically weak. π0.7 is among the first to claim strong signs of it (Physical Intelligence et al. 2026, arXiv 2604.15483), a vendor claim that independent evaluation has not yet tested.
Measured performance. Success rates on benchmarks say less than they appear to. A 2026 study of state-of-the-art VLAs on the BEHAVIOR-1K household challenge argues that metrics based only on final object states "say little about safety aspects of operation and can potentially exaggerate reported performance", and proposes protocols that count safety violations along the way (Rasouli et al. 2026, arXiv 2604.21192).
In code
The chapter's crate models the three things this chapter claims: what a head costs, why cosine compression suits smooth motion, and what the fast-slow split buys. The cost model is short enough to read in full.
/// The three ways a reactive VLA turns backbone features into an action chunk.
#[derive(Clone, Copy, Debug, PartialEq)]
pub enum Head {
/// Autoregressive action tokens, one per action dimension per step, each needing its own
/// sequential decoder pass (RT-2, OpenVLA).
Tokens,
/// Every action token decoded at once in a single pass (OpenVLA-OFT's parallel decoding).
Parallel,
/// A flow-matching or diffusion action expert: `steps` passes of a small network that turn
/// noise into the whole chunk at once (π0, GR00T N1).
Denoising { steps: usize },
}
/// Illustrative costs of one pass through each part of the model, in milliseconds.
#[derive(Clone, Copy, Debug)]
pub struct PassCost {
/// Reading the images and the instruction once (the prefill).
pub backbone_ms: f64,
/// One decoder pass of the language model.
pub decoder_ms: f64,
/// One pass of the action expert.
pub expert_ms: f64,
}The loop of figure 5.2 is the same few lines in Rust and in the page's TypeScript: each model reads its inputs at the start of its period and delivers its output one period later.
/// Two cups slide back and forth on a table. The instruction names cup 0, then changes to cup 1.
pub const SWITCH_S: f64 = 1.45;
pub const DURATION_S: f64 = 3.0;
/// Fastest the hand can move, in table units per second.
pub const HAND_SPEED: f64 = 3.0;
/// Position of cup `which` (0 or 1) at time `t` in seconds.
pub fn cup(which: usize, t: f64) -> f64 {
which as f64 + 0.25 * (std::f64::consts::TAU * 0.6 * t + 1.3 * which as f64).sin()
}
/// The cup the instruction names at time `t`.
pub fn instructed(t: f64) -> usize {
usize::from(t >= SWITCH_S)
}
#[derive(Clone, Copy, Debug, PartialEq)]
pub enum Loop {
/// One large model does everything: it looks, thinks for one period, then sends the hand to
/// where the named cup was when it looked.
Single { hz: f64 },
/// A slow reasoner decides which cup; a fast controller follows that cup with fresh images.
Dual { slow_hz: f64, fast_hz: f64 },
}
#[derive(Clone, Debug, Default)]
pub struct Trace {
pub t: Vec<f64>,
pub target: Vec<f64>,
pub hand: Vec<f64>,
}
/// Simulate three seconds at 1 kHz. Each model reads its inputs at the start of a period and its
/// output arrives one period later, so a model's latency equals its period.
pub fn simulate(l: Loop) -> Trace {
let period = |hz: f64| ((1000.0 / hz).round() as usize).max(1);
let (slow_p, fast_p) = match l {
Loop::Single { hz } => (period(hz), 0),
Loop::Dual { slow_hz, fast_hz } => (period(slow_hz), period(fast_hz)),
};
let n = (DURATION_S * 1000.0) as usize;
let mut hand = cup(0, 0.0);
let mut command = hand;
let mut belief = 0; // the cup the fast controller has been told to follow
let (mut slow_snapshot, mut fast_snapshot): (Option<(usize, f64)>, Option<f64>) = (None, None);
let mut tr = Trace::default();
for k in 0..n {
let t = k as f64 / 1000.0;
if k % slow_p == 0 {
if let Some((which, pos)) = slow_snapshot.take() {
match l {
Loop::Single { .. } => command = pos,
Loop::Dual { .. } => belief = which,
}
}
let which = instructed(t);
slow_snapshot = Some((which, cup(which, t)));
}
if fast_p > 0 && k % fast_p == 0 {
if let Some(pos) = fast_snapshot.take() {
command = pos;
}
fast_snapshot = Some(cup(belief, t));
}
let step = HAND_SPEED / 1000.0;
hand += (command - hand).clamp(-step, step);
tr.t.push(t);
tr.target.push(cup(instructed(t), t));
tr.hand.push(hand);
}
tr
}
/// Mean distance between hand and the named cup over the window [from, to) seconds.
pub fn tracking_error(tr: &Trace, from: f64, to: f64) -> f64 {
let (mut sum, mut n) = (0.0, 0);
for i in 0..tr.t.len() {
if tr.t[i] >= from && tr.t[i] < to {
sum += (tr.hand[i] - tr.target[i]).abs();
n += 1;
}
}
sum / n.max(1) as f64
}Exercises
The stubs are in rust/ch05-reactive-vla/src/exercises.rs.
Exercise 5.1
What a chunk costs
Implement sequential_passes and chunk_latency_ms for the three heads. Then compute how many passes a token head would need for a 25-step chunk on a 7-dimensional arm, and explain in one sentence why parallel decoding changed what the token family could do.
Check your answer from the rust/ folder:
cargo test -p ch05-reactive-vla --test exercises ex5_1Hint 1
A token head needs one pass per token; how many tokens are in a chunk?
Hint 2
Denoising passes are paid by the expert, not by the language model's decoder.
Exercise 5.2
FAST in miniature
Implement the orthonormal discrete cosine transform dct and compress, which keeps only the first coefficients. The tests check that six of fifty coefficients describe a smooth reach almost exactly, and that jittery motion does not compress. What does that predict about using FAST on high-frequency, noisy teleoperation data?
Check your answer from the rust/ folder:
cargo test -p ch05-reactive-vla --test exercises ex5_2Hint 1
Write the sum directly; the length is small, so an O(n²) transform is fine.
Hint 2
To compress, zero every coefficient from index keep onwards, then call idct.
Exercise 5.3
How long does a new instruction take?
Implement reaction_time: how long after the instruction switch the hand settles on the new cup. Compare a 2 Hz and an 8 Hz reasoner over the same 200 Hz controller, then use figure 5.2 to check whether a faster controller changes the answer.
Check your answer from the rust/ folder:
cargo test -p ch05-reactive-vla --test exercises ex5_3Hint 1
Walk backwards from the end of the trace while the hand stays within tolerance.
Hint 2
The reaction time is measured from SWITCH_S, not from zero.
What we are not sure about
Where the family ends. The line between a reactive VLA and a world-action model is blurring from both sides. π0.7 conditions on images of the future; LingBot-VLA 2.0 trains with future prediction as a proxy task; Fast-WAM in chapter 6 trains like a WAM and runs like a VLA. This book keeps a model here if it does not generate futures at inference, but that is a convention, not a law.
Which head wins. Action experts dominate releases since 2025, but OpenVLA-OFT reports that L1 regression with parallel decoding matches diffusion-based fine-tuning on its benchmarks while training faster (Kim et al. 2025, arXiv 2502.19645), and FAST reports token VLAs matching diffusion VLAs (Pertsch et al. 2025, arXiv 2501.09747). The comparisons use different robots and data, so the question is open.
How much the dual-system split matters. Helix, GR00T and Gemini Robotics all split the model, but no public ablation separates the benefit of the split from the benefit of each company's data and hardware. The toy in figure 5.2 shows the mechanism, not its size on a real robot.
Vendor numbers. GEN-1's 99 percent, GEN-1.5's one-shot learning and π0.7's emergent composition are each reported by the model's maker on its own tasks. Until independent evaluations exist, they are claims about a direction, not measurements to compare.
Further reading
- π0, the paper that introduced the flow-matching action expert (Black et al. 2024, arXiv 2410.24164).
- π0.7, for the steerable prompt and what it lets a model learn from (Physical Intelligence et al. 2026, arXiv 2604.15483).
- OpenVLA-OFT, a careful study of which design choices in a VLA actually matter (Kim et al. 2025, arXiv 2502.19645).
- FAST, for action tokenization done well (Pertsch et al. 2025, arXiv 2501.09747).
- OpenHelix, a short survey and empirical analysis of dual-system VLAs (Cui et al. 2025, arXiv 2505.03912).
- An anatomy of VLA models, a modules-to-challenges survey with a maintained project page (Xu et al. 2025, arXiv 2512.11362).
- How VLAs (really) work in open-world environments, for reading benchmark claims critically (Rasouli et al. 2026, arXiv 2604.21192).
- MolmoAct2, the most open recent release: weights, data and recipes (Fang et al. 2026, arXiv 2605.02881).
References
31 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- A. Brohan, N. Brown, J. Carbajal and 51 others RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv 2307.15818 v1, 2023-07-28.
- M. J. Kim, K. Pertsch, S. Karamcheti and 15 others OpenVLA: An Open-Source Vision-Language-Action Model. arXiv 2406.09246 v3, 2024-06-13.
- K. Black, N. Brown, D. Driess and 21 others π0: A Vision-Language-Action Flow Model for General Robot Control. arXiv 2410.24164 v4, 2024-10-31.
- Physical Intelligence, K. Black, N. Brown and 33 others π0.5: a Vision-Language-Action Model with Open-World Generalization. arXiv 2504.16054 v1, 2025-04-22.
- Physical Intelligence, B. Ai, A. Amin and 85 others π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities. arXiv 2604.15483 v2, 2026-04-16.
- NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
- K. Pertsch, K. Stachowicz, B. Ichter and 6 others FAST: Efficient Action Tokenization for Vision-Language-Action Models. arXiv 2501.09747 v1, 2025-01-16.
- M. J. Kim, C. Finn, P. Liang Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success. arXiv 2502.19645 v2, 2025-02-27.
- Y. Zhong, F. Bai, S. Cai and 11 others A Survey on Vision-Language-Action Models: An Action Tokenization Perspective. arXiv 2507.01925 v1, 2025-07-02.
- NVIDIA, J. Bjorck, F. Castañeda and 39 others GR00T N1: An Open Foundation Model for Generalist Humanoid Robots. arXiv 2503.14734 v2, 2025-03-18.
- M. Shukor, D. Aubakirova, F. Capuano and 11 others SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics. arXiv 2506.01844 v1, 2025-06-02.
- R. Cai, J. Guo, X. He and 20 others Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution. arXiv 2602.12684 v2, 2026-02-13.
- Figure AI Helix: A Vision-Language-Action Model for Generalist Humanoid Control. figure.ai, 2025-02-20.
- Gemini Robotics Team, S. Abeyruwan, J. Ainslie and 115 others Gemini Robotics: Bringing AI into the Physical World. arXiv 2503.20020 v1, 2025-03-25.
- C. Cui, P. Ding, W. Song and 10 others OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation. arXiv 2505.03912 v1, 2025-05-06.
- T. Zou, H. Zeng, Y. Nong and 6 others Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation. arXiv 2512.20188 v1, 2025-12-23.
- SVRC (Robotics Center of Silicon Valley) State of Robotics 2026. roboticscenter.ai, 2026-03.
- Open X-Embodiment Collaboration, A. O'Neill, A. Rehman and 291 others Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv 2310.08864 v9, 2023-10-13.
- Google DeepMind Gemini Robotics On-Device brings AI to local robotic devices. deepmind.google, 2025-06-24.
- Generalist AI GEN-1: Scaling Embodied Foundation Models to Mastery. generalistai.com, 2026-04-02.
- H. Fang, J. Duan, D. Clay and 26 others MolmoAct2: Action Reasoning Models for Real-world Deployment. arXiv 2605.02881 v2, 2026-05-04.
- W. Wu, F. Wang, F. Lu and 21 others From Foundation to Application: Improving VLA Models in Practice. arXiv 2607.06403 v1, 2026-07-07.
- Generalist AI GEN-1.5: Embodied Foundation Models are One-Shot Learners. generalistai.com, 2026-08-19.
- R. Cadene, S. Aliberts, F. Capuano and 14 others LeRobot: An Open-Source Library for End-to-End Robot Learning. arXiv 2602.22818 v1, 2026-02-26.
- MarkTechPost Top 10 Physical AI Models Powering Real-World Robots in 2026. marktechpost.com, 2026-04-28.
- Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.
- K. Black, M. Y. Galliker, S. Levine Real-Time Execution of Action Chunking Flow Policies. arXiv 2506.07339 v2, 2025-06-09.
- K. Zhou, Q. Chen, D. Peng and 3 others Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment. arXiv 2604.24447 v1, 2026-04-27.
- NVIDIA Newsroom NVIDIA and Global Robotics Leaders Take Physical AI to the Real World (GTC 2026, GR00T N2 preview). nvidianews.nvidia.com, 2026-03-16.
- A. Rasouli, Y. Wu, Z. Li and 4 others How VLAs (Really) Work In Open-World Environments. arXiv 2604.21192 v1, 2026-04-23.
- C. Xu, S. Zhang, Y. Liu and 11 others An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges. arXiv 2512.11362 v3, 2025-12-12.