Ch.14Part III · Cross-cutting engineering
Post-training: adaptation, RL, steerability, memory
Fine-tuning economics, three routes to reinforcement learning, steerability and memory.
- Math
- L1
- Sources
- 16
- Figures
- 5
- Exercises
- 3
- Last checked
- 9 Oct 2026
One boxed equation per idea
10 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
2. Fine-tune: tens to hundreds of demonstrations per task
Supervised fine-tuning on demonstrations of the target task, often with low-rank adapters so only a sliver of the weights change.
Examples: 200 to 500 demonstrations (industry survey); 50 to 100 for Gemini Robotics On-Device (vendor claim).
Why this chapter exists
A pretrained robot foundation model is a generalist. It can attempt many tasks and does none of them as well as a specialist trained for one. Getting from the generalist to a robot that folds this laundry, in this home, quickly and without dropping it, is post-training. It is where most of the engineering effort in a deployment goes, and where the economics of chapter 2 come from: if a few hundred demonstrations and a few hours of practice turn a shared base into a specialist, the base is worth building.
This chapter covers the four tools the field uses: fine-tuning on demonstrations, reinforcement learning (RL) from practice, steering through the prompt, and memory for tasks longer than a model's context. It leans on Physical Intelligence's published line of work, the most complete public record of post-training a VLA, and flags where that record comes from a single lab.
After this chapter you will be able to:
- explain why low-rank adaptation lets a team fine-tune a billion-parameter model on one GPU, and count what it trains;
- compare the three RL routes of figure 14.2 by what each costs and what each risks, and say what an anchor to the pretrained policy buys;
- describe how steering through the prompt and memory at two time scales extend what a model can do without new demonstrations.
It is acknowledged among practitioners that the particular implementation details of these algorithms are often just as important (if not more so) for performance as the choice of algorithm.
Fine-tuning
The first step is supervised fine-tuning: more imitation, on demonstrations of the target task. The economics are what made foundation models attractive. One industry survey reports that fine-tuning a pretrained VLA on 200 to 500 demonstrations now beats a task-specific policy trained from scratch on more than 1,000industry survey (SVRC (Robotics Center of Silicon Valley) 2026). Model makers report smaller numbers for their own models: 50 to 100 demonstrations for Gemini Robotics On-Devicevendor claim (Google DeepMind 2025), and 130 post-training episodes per task in LingBot-VLA's evaluation (Wu et al. 2026, arXiv 2601.18692).
Fine-tuning every weight of a multi-billion-parameter model needs a lot of GPU memory. Low-rank adaptation (LoRA) avoids it. It freezes each weight matrix and learns a small correction made of two thin matrices:
The rank is small, typically 4 to 64, so has far fewer parameters than . In the language-model setting where it was introduced, LoRA cut the trainable parameters by 10,000 times and GPU memory by 3 times relative to full fine-tuning, with no loss in quality (Hu et al. 2021, arXiv 2106.09685). For VLAs, OpenVLA's authors fine-tuned on consumer GPUs this way (Kim et al. 2024, arXiv 2406.09246). Exercise 14.1 counts the savings for a 2-billion-parameter backbone: at rank 16, LoRA trains about 1 percent of the weights.
Reinforcement learning: three routes
Fine-tuning can only teach what the demonstrations show. To be faster or more reliable than the demonstrator, a policy has to practise: try, observe the outcome, and improve. That is reinforcement learning, and for a robot the question is where the practice happens. The WAM survey frames three routes (Lu et al. 2026, arXiv 2609.16074), drawn in figure 14.2.
Real-world RL
Costs: robot time, resets, human supervision and corrections. Risks: safety during learning, and results from only a few labs.
SERL learns tasks in 25 to 50 minutes per policy (2024); π*0.6 with RECAP more than doubles throughput on its hardest tasks (vendor report, November 2025); RL Token speeds up the hardest part of a task up to 3× within hours (April 2026); VLA-Precision targets high-precision tasks with large VLAs (September 2026).
Simulation RL trains in a physics simulator, with privileged state and unlimited resets, and pays at transfer. A February 2026 study fine-tuned OpenVLA and π0.5 with RL in simulation, while anchoring them with a supervised loss on real data so they would not forget. It reports 24 and 20 points of real-world success over real-only fine-tuning (Shi et al. 2026, arXiv 2602.12628).
Real-world RL practises on the robot itself. It used to be impractical; SERL showed in 2024 that a careful implementation learns contact-rich tasks such as PCB assembly and cable routing in 25 to 50 minutes of training per policy (Luo et al. 2024, arXiv 2401.16013). Physical Intelligence's RECAP method trains π*0.6 from demonstrations, its own autonomous attempts and human corrections during execution. It reports more than doubling throughput and roughly halving failures on its hardest tasks, such as folding laundry and making espressovendor claim (Physical Intelligence et al. 2025, arXiv 2511.14759). Its RL Token work exposes a compact readout from the VLA and trains a small actor-critic head on it, anchored to the VLA. It reports up to 3 times faster execution on the hardest part of tasks such as Ethernet insertion within minutes to a few hours of practice (Xu et al. 2026, arXiv 2604.23073). At fleet scale, SOP streams experience and interventions from many robots to a cloud learner and sends back updated policies (Pan et al. 2026, arXiv 2601.03044). VLA-Precision, from September 2026, targets high-precision tasks with large VLAs (Su et al. 2026, arXiv 2609.04355).
World-model RL practises in a learned model, chapter 8's RL environment role. GigaBrain-0.5M* reports about 30 percent gains over RECAP on laundry folding, box packing and espresso preparation by training with a world modelvendor claim (GigaBrain Team et al. 2026, arXiv 2602.12099). World-VLA-Loop co-trains the world model and the policy so the model stays accurate where the policy goes (Liu et al. 2026, arXiv 2602.06508).
Staying close to what was learned
Every route faces the same danger: a policy that chases reward can forget what pretraining taught it, or exploit a flaw in the reward or the model. The common defence is an anchor that keeps the improved policy close to the pretrained one. RL Token anchors its actor to the VLA, and RL-Co anchors to real data. The simplest form of the idea is reward-weighted improvement with a pull back towards the original:
Read it as: try a few actions around the current aim , move towards the ones that earned more reward , and then pull back by a fraction towards the pretrained aim . The temperature sets how strongly reward decides. Figure 14.3 runs it.
- Reward at the current aim
- 0.89pretrained policy: 0.02
- Distance moved from the pretrained aim
- 0.50
Without an anchor, the aim reaches the best setting in about ten rounds. An anchor of 0.1 settles at 0.50 with a reward of 0.89, and an anchor of 0.3 stops at 0.31. Shrink the exploration and nothing moves at all, because a policy that never tries anything different has nothing to learn from. Real RL for VLAs is far more complex, but the trade-off between improvement and staying close is the same one.
Steering through the prompt
Chapter 5 introduced π0.7's prompt, which says how to do a task as well as what to do: speed, quality, whether a mistake occurred, a subgoal image (Physical Intelligence et al. 2026, arXiv 2604.15483). For post-training this is an alternative to collecting new data. A behaviour that exists somewhere in the training data can be selected by describing it, rather than taught again. Figure 14.4 shows the idea on a toy dataset of 24 demonstrations.
Task: go to the goal. Speed: fast. Quality: 5. Mistake: false.
- Demonstrations matching
- 6 of 24
- With a fumble
- 0%
- Slow
- 0%
- Passing above
- 50%
With the task alone, the model imitates the whole dataset, fumbles and slow runs included. Adding the metadata narrows it to fast, clean demonstrations, on either route. Adding a subgoal image of passing above narrows it to one route. Nothing was retrained between the three prompts. The limits are the same as in chapter 5: the behaviour must be in the data, and the labels must be right.
Memory for long tasks
A VLA that sees only the last moment cannot remember that the stove is on or that it has already flipped the sandwich. Simply feeding it more past frames is expensive and does not scale to minutes. Physical Intelligence's MEM combines two memories: a short-term memory of recent video, compressed by a video encoder, which covers occlusion; and a long-term memory in text, which records what has been done (Torne et al. 2026, arXiv 2603.03596). Its authors report policies performing tasks of up to fifteen minutes, such as cleaning a kitchen or making a grilled cheese sandwich, and adapting manipulation strategies in context (Torne et al. 2026, arXiv 2603.03596). π0.7 includes a MEM-style video history encoder (Physical Intelligence et al. 2026, arXiv 2604.15483).
| Question the policy must answer | Current frame only | Short-term video | Video plus text, as in MEM |
|---|---|---|---|
| Is the stove on? (yes) | cannot tell | has forgotten | knows |
| Has the sandwich been flipped? (yes) | cannot tell | has forgotten | knows |
| Where is the spatula the arm is now hiding? (where it was seen a few seconds ago) | cannot tell | knows | knows |
In code
The crate's refinement loop is the boxed equation, with deterministic attempts so runs can be compared.
/// A one-dimensional skill, such as how far to push a connector. The pretrained policy aims at
/// `BASE`; the task rewards aiming at `BEST`.
pub const BASE: f64 = 0.0;
pub const BEST: f64 = 0.6;
/// Reward of an attempt at `a`: 1 at the best setting, falling off with distance.
pub fn reward(a: f64) -> f64 {
(-((a - BEST) / 0.3).powi(2)).exp()
}
/// Deterministic standard-normal quantiles for `k` attempts, so runs are reproducible.
pub fn quantiles(k: usize) -> Vec<f64> {
(0..k)
.map(|i| {
let p = (i as f64 + 0.5) / k as f64;
let q = if p < 0.5 { p } else { 1.0 - p };
let t = (-2.0 * q.ln()).sqrt();
let z = t - (2.515517 + 0.802853 * t + 0.010328 * t * t) / (1.0 + 1.432788 * t + 0.189269 * t * t + 0.001308 * t * t * t);
if p < 0.5 { -z } else { z }
})
.collect()
}Exercises
The stubs are in rust/ch14-post-training/src/exercises.rs.
Exercise 14.1
What LoRA trains
Implement trainable. Then estimate, at two bytes per parameter, how much memory the adapter weights need at rank 16, and compare with the full model.
Check your answer from the rust/ folder:
cargo test -p ch14-post-training --test exercises ex14_1Hint 1
A full matrix has d_in × d_out entries; a rank-r adapter has r × (d_in + d_out).
Hint 2
Sum over every adapted matrix.
Exercise 14.2
Improve from reward, stay anchored
Implement refine. Then find the smallest exploration spread for which the unanchored policy reaches the best setting within 30 rounds, and explain why exploration and anchoring pull in opposite directions.
Check your answer from the rust/ folder:
cargo test -p ch14-post-training --test exercises ex14_2Hint 1
For each round, compute the attempts from the current aim before moving it.
Hint 2
Apply the anchor after computing the reward-weighted mean.
Exercise 14.3
Two memories
Implement recall. Then list three questions about a task you know whose answers need long-term memory, and one that needs short-term memory, and say which modality, video or text, suits each.
Check your answer from the rust/ folder:
cargo test -p ch14-post-training --test exercises ex14_3Hint 1
Keep only events at or before now.
Hint 2
The short-term memory holds events newer than now minus the window.
What we are not sure about
How general real-world RL is. The most impressive real-world RL results for large VLAs come from one or two labs, on their own robots and tasks. Whether the recipes transfer to other teams, robots and task types is not yet known.
Whether world-model RL beats real-world RL. GigaBrain-0.5M*'s comparison against RECAP is the clearest head-to-head, and it is the maker's own. Chapter 8's warnings about model bias apply in full.
How much steering replaces data. Prompting for speed and quality selects behaviour the data already contains. How far that extends to behaviour the data never showed, and how robust it is when labels are noisy, is open.
What memory should hold. MEM's split into video and text is one answer. Others, such as learned latent memories, have not been compared on the same tasks.
Further reading
- LoRA, for the adaptation method most VLA fine-tuning uses (Hu et al. 2021, arXiv 2106.09685).
- SERL, for practical real-world RL and why implementation matters (Luo et al. 2024, arXiv 2401.16013).
- π*0.6 and RECAP, for RL on a large VLA from deployment data (Physical Intelligence et al. 2025, arXiv 2511.14759).
- RL Token, for fast online RL anchored to a VLA (Xu et al. 2026, arXiv 2604.23073).
- RL-Co, for simulation RL that does not forget the real world (Shi et al. 2026, arXiv 2602.12628).
- MEM, for memory over fifteen-minute tasks (Torne et al. 2026, arXiv 2603.03596).
References
16 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- SVRC (Robotics Center of Silicon Valley) State of Robotics 2026. roboticscenter.ai, 2026-03.
- Google DeepMind Gemini Robotics On-Device brings AI to local robotic devices. deepmind.google, 2025-06-24.
- W. Wu, F. Lu, Y. Wang and 22 others A Pragmatic VLA Foundation Model. arXiv 2601.18692 v4, 2026-01-26.
- E. J. Hu, Y. Shen, P. Wallis and 5 others LoRA: Low-Rank Adaptation of Large Language Models. arXiv 2106.09685 v2, 2021-06-17.
- M. J. Kim, K. Pertsch, S. Karamcheti and 15 others OpenVLA: An Open-Source Vision-Language-Action Model. arXiv 2406.09246 v3, 2024-06-13.
- Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.
- L. Shi, S. Chen, F. Gao and 8 others Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models. arXiv 2602.12628 v4, 2026-02-13.
- J. Luo, Z. Hu, C. Xu and 7 others SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning. arXiv 2401.16013 v4, 2024-01-29.
- Physical Intelligence, A. Amin, R. Aniceto and 53 others π*0.6: a VLA That Learns From Experience. arXiv 2511.14759 v2, 2025-11-18.
- C. Xu, J. T. Springenberg, M. Equi and 4 others RL Token: Bootstrapping Online RL with Vision-Language-Action Models. arXiv 2604.23073 v2, 2026-04-24.
- M. Pan, S. Feng, Q. Zhang and 9 others SOP: A Scalable Online Post-Training System for Vision-Language-Action Models. arXiv 2601.03044 v1, 2026-01-06.
- C. Su, Z. Shen, Y. Qian and 8 others VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models. arXiv 2609.04355 v3, 2026-09-03.
- GigaBrain Team, B. Wang, B. Li and 23 others GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning. arXiv 2602.12099 v2, 2026-02-12.
- X. Liu, Z. Bai, H. Ci and 2 others World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy. arXiv 2602.06508 v2, 2026-02-06.
- Physical Intelligence, B. Ai, A. Amin and 85 others π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities. arXiv 2604.15483 v2, 2026-04-16.
- M. Torne, K. Pertsch, H. Walke and 14 others MEM: Multi-Scale Embodied Memory for Vision Language Action Models. arXiv 2603.03596 v2, 2026-03-04.