Ch.11Part II · The seven familiesLarge behavior models
Large behavior models and scaled imitation
Imitation learning at scale without a language backbone: Diffusion Policy, ACT and their descendants.
- Math
- L1
- Sources
- 10
- Figures
- 4
- Exercises
- 3
- Last checked
- 9 Oct 2026
One boxed equation per idea
1 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
Why this chapter exists
Before vision-language-action models, robot learning had its own workhorses: policies trained by imitation on teleoperated demonstrations, with no language model anywhere. Two of them, Diffusion Policy and ACT, did so well in 2023 that their ideas spread into everything that came after. Figure 11.1 shows where those ideas ended up. π0's action expert generates chunks by flow matching, which its authors describe as a variant of diffusion, and the chunking that every VLA in chapter 5 relies on is ACT's (Black et al. 2024, arXiv 2410.24164).
Large behavior models (LBMs) are what happens when those workhorses are scaled up directly: trained on thousands of hours of robot demonstrations, without the vision-language backbone. Some teams want the data-scaling benefits of a foundation model without a VLM's inference cost or its language interface. This chapter is about that path, about the two primitives everyone borrowed from it, and about the evaluation discipline its most careful study brought with it.
After this chapter you will be able to:
- say what distinguishes a large behavior model from a VLA, and why the distinction is about inputs rather than the action head;
- explain with figure 11.3 how temporal ensembling trades reaction speed for smoothness;
- work out with figure 11.4 roughly how many real-robot trials it takes to show that one policy beats another.
While these models have garnered significant enthusiasm and investment, meaningful evaluation of real-world performance remains a challenge, limiting both the pace of development and inhibiting a nuanced understanding of current capabilities.
The mechanism
A large behavior model is a policy trained by behaviour cloning: given the observation , reproduce the chunk of actions the demonstrator took. Diffusion Policy writes the policy as a denoising process over the action chunk. Training teaches a network to recognise the noise that was added to a demonstrated chunk:
where is the noise level. At run time the network removes noise step by step, starting from pure noise, which is the operational picture of chapter 4: noise goes in, an action chunk comes out, in steps (Chi et al. 2023, arXiv 2303.04137). The authors found this formulation handles several valid ways of doing a task gracefully, suits high-dimensional actions, and trains stably, with an average improvement of 46.9 percent over prior methods across 12 tasks (Chi et al. 2023, arXiv 2303.04137). TRI's LBMs use this exact loss with a diffusion transformer that predicts 16 steps of 20-dimensional actions and runs at 10 Hz (TRI LBM Team et al. 2025, arXiv 2507.05331).
What is missing is the pretrained vision-language model. Inputs pass through image encoders, and in TRI's case a CLIP text encoder turns the task description into a feature vector (TRI LBM Team et al. 2025, arXiv 2507.05331). That is language conditioning, but not language understanding in the sense of chapter 5. A CLIP embedding tells the policy which of its trained tasks is meant; it does not let the policy follow an instruction it has never seen.
Large behavior model (Diffusion Policy, ACT, Octo, TRI LBM). Only the policy is learned and run. The instruction, if there is one, comes through a small text encoder such as CLIP, not a language model, so the task is mostly specified by the demonstrations themselves.
Action chunks, and how to execute them
ACT, introduced with the ALOHA bimanual teleoperation system, made action chunking standard. Its authors report learning six difficult real-world tasks, such as opening a translucent condiment cup and slotting a battery, at 80 to 90 percent success from only 10 minutes of demonstrations (Zhao et al. 2023, arXiv 2304.13705). They also proposed a way to execute chunks smoothly. Query the policy at every step, so that several overlapping chunks each contain a prediction for the current step, and average them:
Here is the -th prediction for step , and belongs to the oldest one (Zhao et al. 2023, arXiv 2304.13705). Figure 11.3 compares this temporal ensemble with the two simpler options.
- Mean tracking error
- 0.022
- Jerk
- 8.8mean second difference, ×1000
- Error, 20 steps after the nudge
- 0.189how fast it reacts
The three ways of executing fail differently. Executing whole chunks open loop drifts within each chunk and reacts to the nudge only at the next chunk boundary. Executing only the newest action reacts at once, but every query commits to a slightly different plan, so the motion jerks: in the toy it is about six times jerkier than the ensemble. The ensemble is smooth and accurate, and it lags after the nudge because predictions made before the nudge still point at the old position. Raising trusts those old predictions more and makes the lag worse. Real-time chunking and the asynchronous execution of chapters 4 and 5 are later answers to the same problem (Black et al. 2025, arXiv 2506.07339).
Scaling up: the large behavior models
Octo (May 2024) brought the recipe to a cross-embodiment scale: a transformer policy with a diffusion action head, trained on 800,000 trajectories from Open X-Embodiment, at 27 or 93 million parameters, instructable by language or goal images and fine-tunable on consumer GPUs (Octo Model Team et al. 2024, arXiv 2405.12213). It sits on the boundary with chapter 5: it takes language, through a small encoder, but has no pretrained VLM.
TRI's Large Behavior Models (July 2025) are the family's most careful study. TRI trained Diffusion Policy-style models on about 1,700 hours of robot demonstrations, covering more than 500 internally collected tasks plus public data. They evaluated them in 1,800 blind, randomised real-world trials and more than 47,000 simulation rollouts (TRI LBM Team et al. 2025, arXiv 2507.05331). Pretraining made policies more successful and more robust under distribution shift. To match a from-scratch policy in simulation, a fine-tuned LBM needed less than 30 percent of the data, and performance rose predictably with pretraining scale and diversity (TRI LBM Team et al. 2025, arXiv 2507.05331).
GEN-0 (November 2025) scaled the same bet much further. Generalist AI reports pretraining on more than 270,000 hours of real-world manipulation data, growing by 10,000 hours a week, and a "phase transition" at 7 billion parameters: smaller models stop improving with more data while larger ones keep goingvendor claim (Generalist AI 2025). Its successor GEN-1 is "trained from scratch" on half a million hours, with a base model that "uses data from low-cost wearable devices on humans" rather than robot datavendor claim (Generalist AI 2026). The GEN models are usually compared with VLAs, and chapter 5 lists them for that reason. By this book's test of what feeds the policy, they belong here: there is no pretrained vision-language model underneath.
How many trials?
Most robot learning papers report success rates from 10 to 50 real-world trials per condition. Figure 11.4 shows what that can and cannot tell you.
Chance that a one-sided test at the 5% level finds B better, against the number of trials per policy:
- Chance of detecting the difference
- 16%with 20 trials each
- Trials each for 80% power
- about 280
The uncertainty in a success rate is easy to underestimate. Eight successes out of ten has a 95 percent Wilson interval from about 49 to 94 percent:
where is the observed rate, the number of trials, and for 95 percent. The arithmetic of comparison is harsher. With 20 trials of each policy, a real 10-point difference between 60 and 70 percent is detected about 16 percent of the time; reaching the conventional 80 percent chance takes about 280 trials of each. A 20-point difference, 40 against 60 percent, needs about 80. This is why TRI's 1,800 real-world trials matter, and why chapter 15 treats evaluation as an engineering problem in its own right. Exercises 11.2 and 11.3 compute these numbers.
Why it exists, good at, bad at
Why it exists. Some teams want the data-scaling benefits without a VLM's inference cost or language interface, and the imitation-learning primitives became the action heads inside every other family anyway.
What it is good at. LBMs are small, fast and well understood: Octo's larger model has 93 million parameters (Octo Model Team et al. 2024, arXiv 2405.12213), against billions for a VLA. Pretrained broadly, they are robust under distribution shift (TRI LBM Team et al. 2025, arXiv 2507.05331). And they remain the workhorse of bimanual teleoperation research, where a few hours of demonstrations and an afternoon of training produce a working policy.
What it is bad at. No web knowledge, and at most a narrow form of language grounding. The task is specified by demonstration: to teach a new behaviour, you show it. That rules out the "do something you have never seen, described in words" generalisation that is the point of chapter 5.
In code
The chunk model behind figure 11.2 is a policy whose every query commits to a slightly different plan, plus the two simple ways of executing it.
/// The motion the demonstrations show.
pub fn target(t: usize) -> f64 {
let t = t as f64;
0.6 * (0.07 * t).sin() + 0.25 * (0.19 * t + 1.0).sin()
}
/// At step `NUDGE_AT` someone moves the object by `NUDGE`: the right motion shifts from then on.
pub const NUDGE_AT: usize = 150;
pub const NUDGE: f64 = 0.3;
/// Where the robot should be at step t, nudge included.
pub fn world(t: usize) -> f64 {
target(t) + if t >= NUDGE_AT { NUDGE } else { 0.0 }
}
/// Deterministic noise in [-1, 1].
pub fn noise(a: u64, b: u64) -> f64 {
let mut x = a.wrapping_mul(0x9E37_79B9_7F4A_7C15) ^ b.wrapping_mul(0xC2B2_AE3D_27D4_EB4F) ^ 0x1656_67B1_9E37_79F9;
x ^= x >> 33;
x = x.wrapping_mul(0xFF51_AFD7_ED55_8CCD);
x ^= x >> 33;
(x % 20_001) as f64 / 10_000.0 - 1.0
}
/// A learned policy queried at step `t` predicts a chunk of `horizon` actions for t, t+1, ...
/// Each query commits to a slightly different plan (`offset`), and its error grows the further
/// ahead it looks (`drift` per step), as a real imitation policy's does.
#[derive(Clone, Copy, Debug)]
pub struct Policy {
pub horizon: usize,
pub offset: f64,
pub drift: f64,
pub seed: u64,
}
impl Policy {
/// The chunk predicted at step t. It sees the nudge only if it has already happened.
pub fn predict(&self, t: usize) -> Vec<f64> {
let bias = self.offset * noise(self.seed, t as u64);
let slope = self.drift * noise(self.seed + 1, t as u64);
let seen = if t >= NUDGE_AT { NUDGE } else { 0.0 };
(0..self.horizon).map(|i| target(t + i) + seen + bias + slope * i as f64).collect()
}
}
/// Query once per chunk and execute it to the end, then query again.
pub fn run_open_loop(p: &Policy, steps: usize) -> Vec<f64> {
let mut out = Vec::with_capacity(steps);
while out.len() < steps {
let t = out.len();
out.extend(p.predict(t).into_iter().take(steps - t));
}
out
}
/// Query every step and execute only the first action of the newest chunk.
pub fn run_latest(p: &Policy, steps: usize) -> Vec<f64> {
(0..steps).map(|t| p.predict(t)[0]).collect()
}
/// Mean distance from where the robot should be.
pub fn tracking_error(actions: &[f64]) -> f64 {
actions.iter().enumerate().map(|(t, a)| (a - world(t)).abs()).sum::<f64>() / actions.len() as f64
}
/// Steps after the nudge until the robot is within `tol` of where it should be, and stays there
/// for the rest of `actions`.
pub fn reaction_steps(actions: &[f64], tol: f64) -> Option<usize> {
let mut settle = None;
for t in (NUDGE_AT..actions.len()).rev() {
if (actions[t] - world(t)).abs() > tol {
break;
}
settle = Some(t - NUDGE_AT);
}
settle
}
/// Mean size of the second difference: how jerky the executed motion is.
pub fn jerk(actions: &[f64]) -> f64 {
let n = actions.len().saturating_sub(2).max(1);
actions.windows(3).map(|w| (w[2] - 2.0 * w[1] + w[0]).abs()).sum::<f64>() / n as f64
}Exercises
The stubs are in rust/ch11-behavior/src/exercises.rs.
Exercise 11.1
Temporal ensembling
Implement run_ensemble. The tests check that it is smoother than re-planning every step and more accurate than open-loop chunks. Then explain why a very large m makes the ensemble behave like executing old chunks, and what that does after the nudge.
Check your answer from the rust/ folder:
cargo test -p ch11-behavior --test exercises ex11_1Hint 1
Keep one buffer per future step; each query appends its prediction to the buffers it covers.
Hint 2
Weights are exp(−m·i) with i = 0 for the oldest entry in the buffer.
Exercise 11.2
An honest error bar
Implement wilson. Then take three success rates quoted in this book's chapters 5 to 10 that report a trial count, and compute their intervals.
Check your answer from the rust/ folder:
cargo test -p ch11-behavior --test exercises ex11_2Hint 1
Use the formula in the chapter; clamp the result to [0, 1].
Hint 2
With zero trials, the honest answer is the whole interval.
Exercise 11.3
How many trials?
Implement power. Use it to find how many trials per policy you would need to detect a 5-point improvement from 85 to 90 percent, and say what that implies for papers that report a few points of gain from 50 real-world trials.
Check your answer from the rust/ folder:
cargo test -p ch11-behavior --test exercises ex11_3Hint 1
Sum the probabilities of every pair of outcomes (successes of A, successes of B) for which the test statistic exceeds the critical value.
Hint 2
Each outcome's probability is the product of two binomial probabilities.
What we are not sure about
Where the family ends. The boundary with chapter 5 is porous. Octo takes language through a small encoder; GEN-1 describes its architecture as building on the strengths of vision and language models without starting from one. This book keeps LBMs separate because the absence of a pretrained language backbone changes what data they can use and what instructions they can follow, but some models can be argued either way.
Whether scale replaces the VLM. GEN-0 and GEN-1 are the strongest bet that robot and human-activity data at sufficient scale can do without web pretraining. The evidence is the makers' own, on their own tasks, so the question is open.
How much the primitives matter now. Diffusion heads, flow matching and L1 regression all appear in the same model families, and OpenVLA-OFT reports regression matching diffusion in its setting (chapter 5). Which head is best may depend more on data than on the head.
Further reading
- Diffusion Policy, the paper that made generative action heads standard (Chi et al. 2023, arXiv 2303.04137).
- ACT and ALOHA, for chunking, temporal ensembling and low-cost bimanual teleoperation (Zhao et al. 2023, arXiv 2304.13705).
- Octo, an open generalist policy with a diffusion head (Octo Model Team et al. 2024, arXiv 2405.12213).
- TRI's careful examination of LBMs, for evaluation with statistical confidence (TRI LBM Team et al. 2025, arXiv 2507.05331).
- GEN-0, read as a vendor report, for the case for scale without a VLM (Generalist AI 2025).
- LeRobot, for an open library that implements most of these policies (Cadene et al. 2026, arXiv 2602.22818).
References
10 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- K. Black, N. Brown, D. Driess and 21 others π0: A Vision-Language-Action Flow Model for General Robot Control. arXiv 2410.24164 v4, 2024-10-31.
- C. Chi, Z. Xu, S. Feng and 5 others Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. arXiv 2303.04137 v5, 2023-03-07.
- TRI LBM Team, J. Barreiros, A. Beaulieu and 79 others A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation. arXiv 2507.05331 v1, 2025-07-07.
- T. Z. Zhao, V. Kumar, S. Levine, C. Finn Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv 2304.13705 v1, 2023-04-23.
- K. Black, M. Y. Galliker, S. Levine Real-Time Execution of Action Chunking Flow Policies. arXiv 2506.07339 v2, 2025-06-09.
- Octo Model Team, D. Ghosh, H. Walke and 16 others Octo: An Open-Source Generalist Robot Policy. arXiv 2405.12213 v2, 2024-05-20.
- Generalist AI GEN-0: Embodied Foundation Models That Scale with Physical Interaction. generalistai.com, 2025-11-04.
- Generalist AI GEN-1: Scaling Embodied Foundation Models to Mastery. generalistai.com, 2026-04-02.
- Robotics Center (SVRC) Best VLA Models 2026: Complete Vision-Language-Action Guide. roboticscenter.ai, 2026-04.
- R. Cadene, S. Aliberts, F. Capuano and 14 others LeRobot: An Open-Source Library for End-to-End Robot Learning. arXiv 2602.22818 v1, 2026-02-26.