Skip to content

Ch.11Part II · The seven familiesLarge behavior models

Large behavior models and scaled imitation

Imitation learning at scale without a language backbone: Diffusion Policy, ACT and their descendants.

Math
L1

One boxed equation per idea

Sources
10

1 papers from 2026

Figures
4

Most are interactive

Exercises
3

Rust, with tests

Last checked
9 Oct 2026

The field moves monthly

Papers cited, by half-year of first arXiv version

Diffusion PolicyMarch 2023, chapter 11
a vision encoder, nothing else
denoising diffusion over actions
ACTApril 2023, chapter 11
a transformer trained as a variational autoencoder
not used
OctoMay 2024, chapter 11
a transformer trained on 800,000 Open X-Embodiment trajectories
a diffusion action head
TRI LBMJuly 2025, chapter 11
CLIP image and text encoders, no language model
a diffusion transformer
π0October 2024, chapter 5
a 3B vision-language model
a 300M flow-matching action expert
GR00T N1March 2025, chapter 5
a VLM as System 2
a flow-matching diffusion transformer as System 1
DreamZeroFebruary 2026, chapter 6
a 14B video generation model
video and action tokens denoised together
ViPRANovember 2025, chapter 7
a video-language model predicting latent actions
a flow-matching decoder
Figure 11.1The two workhorses of imitation learning, a denoising action head and action chunks, began in models with no language backbone and now sit inside nearly every other family. Eight models from four families. Switch between the two primitives to see which models use each, and what surrounds it.schematic, not measured Source: Diffusion Policy (arXiv 2303.04137), ACT (2304.13705), Octo (2405.12213), TRI LBM (2507.05331), π0 (2410.24164), GR00T N1 (2503.14734), DreamZero (chapter 6) and ViPRA (2511.07732); the 'layered stack' reading follows the Robotics Center's 2026 model guide.

Why this chapter exists

Before vision-language-action models, robot learning had its own workhorses: policies trained by imitation on teleoperated demonstrations, with no language model anywhere. Two of them, Diffusion Policy and ACT, did so well in 2023 that their ideas spread into everything that came after. Figure 11.1 shows where those ideas ended up. π0's action expert generates chunks by flow matching, which its authors describe as a variant of diffusion, and the chunking that every VLA in chapter 5 relies on is ACT's (Black et al. 2024, arXiv 2410.24164).

Large behavior models (LBMs) are what happens when those workhorses are scaled up directly: trained on thousands of hours of robot demonstrations, without the vision-language backbone. Some teams want the data-scaling benefits of a foundation model without a VLM's inference cost or its language interface. This chapter is about that path, about the two primitives everyone borrowed from it, and about the evaluation discipline its most careful study brought with it.

After this chapter you will be able to:

  • say what distinguishes a large behavior model from a VLA, and why the distinction is about inputs rather than the action head;
  • explain with figure 11.3 how temporal ensembling trades reaction speed for smoothness;
  • work out with figure 11.4 roughly how many real-robot trials it takes to show that one policy beats another.

While these models have garnered significant enthusiasm and investment, meaningful evaluation of real-world performance remains a challenge, limiting both the pace of development and inhibiting a nuanced understanding of current capabilities.

TRI LBM Team, Toyota Research Institute, authors of A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation. arXiv 2507.05331, 7 July 2025.

The mechanism

A large behavior model is a policy trained by behaviour cloning: given the observation oto_t, reproduce the chunk of actions At=at:t+HA_t = a_{t:t+H} the demonstrator took. Diffusion Policy writes the policy as a denoising process over the action chunk. Training teaches a network ϵθ\epsilon_\theta to recognise the noise ϵ\epsilon that was added to a demonstrated chunk:

L(θ)=∥ϵ−ϵθ(ot, At+noisek, k)∥2\boxed{\mathcal{L}(\theta) = \big\lVert \epsilon - \epsilon_\theta(o_t,\ A_t + \text{noise}_k,\ k) \big\rVert^2}

where kk is the noise level. At run time the network removes noise step by step, starting from pure noise, which is the operational picture of chapter 4: noise goes in, an action chunk comes out, in KK steps (Chi et al. 2023, arXiv 2303.04137). The authors found this formulation handles several valid ways of doing a task gracefully, suits high-dimensional actions, and trains stably, with an average improvement of 46.9 percent over prior methods across 12 tasks (Chi et al. 2023, arXiv 2303.04137). TRI's LBMs use this exact loss with a diffusion transformer that predicts 16 steps of 20-dimensional actions and runs at 10 Hz (TRI LBM Team et al. 2025, arXiv 2507.05331).

What is missing is the pretrained vision-language model. Inputs pass through image encoders, and in TRI's case a CLIP text encoder turns the task description into a feature vector (TRI LBM Team et al. 2025, arXiv 2507.05331). That is language conditioning, but not language understanding in the sense of chapter 5. A CLIP embedding tells the policy which of its trained tasks is meant; it does not let the policy follow an instruction it has never seen.

Past observationsand instructiono<t, ℓFuture observationso t:t+HActionsa t:t+HPolicyp(a | o<t, ℓ)Visual planningp(o | o<t, ℓ)Forward dynamicsp(o | o<t, a)Inverse dynamicsp(a | o<t, o)

Large behavior model (Diffusion Policy, ACT, Octo, TRI LBM). Only the policy is learned and run. The instruction, if there is one, comes through a small text encoder such as CLIP, not a language model, so the task is mostly specified by the demonstrations themselves.

Figure 11.2On the four-component template a large behavior model and a reactive VLA look identical; the difference is what feeds the policy arrow. Chapter 4's template for this family, with a reactive VLA for comparison. Both learn and run the policy only. Source: placement from Diffusion Policy (arXiv 2303.04137) and TRI LBM (2507.05331).

Action chunks, and how to execute them

ACT, introduced with the ALOHA bimanual teleoperation system, made action chunking standard. Its authors report learning six difficult real-world tasks, such as opening a translucent condiment cup and slotting a battery, at 80 to 90 percent success from only 10 minutes of demonstrations (Zhao et al. 2023, arXiv 2304.13705). They also proposed a way to execute chunks smoothly. Query the policy at every step, so that several overlapping chunks each contain a prediction for the current step, and average them:

at=∑iwi a^t(i)∑iwi,wi=e−m i\boxed{a_t = \frac{\sum_i w_i\, \hat a_t^{(i)}}{\sum_i w_i}, \qquad w_i = e^{-m\, i}}

Here a^t(i)\hat a_t^{(i)} is the ii-th prediction for step tt, and w0w_0 belongs to the oldest one (Zhao et al. 2023, arXiv 2304.13705). Figure 11.3 compares this temporal ensemble with the two simpler options.

object nudged0100200300stepInside the ensemble at step 160
Every chunk predicted in the last 20 steps contains an action for step 160. Each dot is one of those predictions, oldest on the left; its size is its weight exp(−m·i). The executed action is their weighted average. Predictions made before the nudge still point at the old position.
oldest predictionnewestgoalexecuted
Mean tracking error
0.022
Jerk
8.8mean second difference, ×1000
Error, 20 steps after the nudge
0.189how fast it reacts
Figure 11.3Averaging overlapping chunks gives smooth motion close to the demonstrations; the price is a slower reaction when the world changes. A policy predicts 20-step chunks of a reaching motion, each slightly different, and someone nudges the object at step 150. Compare executing whole chunks, executing only the newest action, and ACT's temporal ensemble; then look inside the ensemble at one step.toy simulation Source: this book's toy, in rust/ch11-behavior; the ensemble rule is ACT's (Zhao et al. 2023, arXiv 2304.13705).

The three ways of executing fail differently. Executing whole chunks open loop drifts within each chunk and reacts to the nudge only at the next chunk boundary. Executing only the newest action reacts at once, but every query commits to a slightly different plan, so the motion jerks: in the toy it is about six times jerkier than the ensemble. The ensemble is smooth and accurate, and it lags after the nudge because predictions made before the nudge still point at the old position. Raising mm trusts those old predictions more and makes the lag worse. Real-time chunking and the asynchronous execution of chapters 4 and 5 are later answers to the same problem (Black et al. 2025, arXiv 2506.07339).

Scaling up: the large behavior models

Octo (May 2024) brought the recipe to a cross-embodiment scale: a transformer policy with a diffusion action head, trained on 800,000 trajectories from Open X-Embodiment, at 27 or 93 million parameters, instructable by language or goal images and fine-tunable on consumer GPUs (Octo Model Team et al. 2024, arXiv 2405.12213). It sits on the boundary with chapter 5: it takes language, through a small encoder, but has no pretrained VLM.

TRI's Large Behavior Models (July 2025) are the family's most careful study. TRI trained Diffusion Policy-style models on about 1,700 hours of robot demonstrations, covering more than 500 internally collected tasks plus public data. They evaluated them in 1,800 blind, randomised real-world trials and more than 47,000 simulation rollouts (TRI LBM Team et al. 2025, arXiv 2507.05331). Pretraining made policies more successful and more robust under distribution shift. To match a from-scratch policy in simulation, a fine-tuned LBM needed less than 30 percent of the data, and performance rose predictably with pretraining scale and diversity (TRI LBM Team et al. 2025, arXiv 2507.05331).

GEN-0 (November 2025) scaled the same bet much further. Generalist AI reports pretraining on more than 270,000 hours of real-world manipulation data, growing by 10,000 hours a week, and a "phase transition" at 7 billion parameters: smaller models stop improving with more data while larger ones keep goingvendor claim (Generalist AI 2025). Its successor GEN-1 is "trained from scratch" on half a million hours, with a base model that "uses data from low-cost wearable devices on humans" rather than robot datavendor claim (Generalist AI 2026). The GEN models are usually compared with VLAs, and chapter 5 lists them for that reason. By this book's test of what feeds the policy, they belong here: there is no pretrained vision-language model underneath.

How many trials?

Most robot learning papers report success rates from 10 to 50 real-world trials per condition. Figure 11.4 shows what that can and cannot tell you.

0%25%50%75%100%policy A: 12/20policy B: 14/20

Chance that a one-sided test at the 5% level finds B better, against the number of trials per policy:

Chance of detecting the difference
16%with 20 trials each
Trials each for 80% power
about 280
Figure 11.4Telling a 70 percent policy from a 60 percent one takes hundreds of trials per policy; with the 20 trials common in papers, the better policy is detected less than one time in five. Two policies with true success rates you choose. Top: typical observed results and their 95% Wilson intervals. Bottom: the chance that a standard test detects that B is better, against the number of trials per policy.toy simulation Source: exact binomial calculations in rust/ch11-behavior; the evaluation practice follows TRI LBM (arXiv 2507.05331).

The uncertainty in a success rate is easy to underestimate. Eight successes out of ten has a 95 percent Wilson interval from about 49 to 94 percent:

p^+z22n±zp^(1−p^)n+z24n21+z2n\boxed{\frac{\hat p + \frac{z^2}{2n} \pm z\sqrt{\frac{\hat p(1-\hat p)}{n} + \frac{z^2}{4n^2}}}{1 + \frac{z^2}{n}}}

where p^\hat p is the observed rate, nn the number of trials, and z=1.96z = 1.96 for 95 percent. The arithmetic of comparison is harsher. With 20 trials of each policy, a real 10-point difference between 60 and 70 percent is detected about 16 percent of the time; reaching the conventional 80 percent chance takes about 280 trials of each. A 20-point difference, 40 against 60 percent, needs about 80. This is why TRI's 1,800 real-world trials matter, and why chapter 15 treats evaluation as an engineering problem in its own right. Exercises 11.2 and 11.3 compute these numbers.

Why it exists, good at, bad at

Why it exists. Some teams want the data-scaling benefits without a VLM's inference cost or language interface, and the imitation-learning primitives became the action heads inside every other family anyway.

What it is good at. LBMs are small, fast and well understood: Octo's larger model has 93 million parameters (Octo Model Team et al. 2024, arXiv 2405.12213), against billions for a VLA. Pretrained broadly, they are robust under distribution shift (TRI LBM Team et al. 2025, arXiv 2507.05331). And they remain the workhorse of bimanual teleoperation research, where a few hours of demonstrations and an afternoon of training produce a working policy.

What it is bad at. No web knowledge, and at most a narrow form of language grounding. The task is specified by demonstration: to teach a new behaviour, you show it. That rules out the "do something you have never seen, described in words" generalisation that is the point of chapter 5.

In code

The chunk model behind figure 11.2 is a policy whose every query commits to a slightly different plan, plus the two simple ways of executing it.

rust/ch11-behavior/src/lib.rsA chunking policy, open-loop and newest-action execution, and the two metrics
/// The motion the demonstrations show.
pub fn target(t: usize) -> f64 {
    let t = t as f64;
    0.6 * (0.07 * t).sin() + 0.25 * (0.19 * t + 1.0).sin()
}

/// At step `NUDGE_AT` someone moves the object by `NUDGE`: the right motion shifts from then on.
pub const NUDGE_AT: usize = 150;
pub const NUDGE: f64 = 0.3;

/// Where the robot should be at step t, nudge included.
pub fn world(t: usize) -> f64 {
    target(t) + if t >= NUDGE_AT { NUDGE } else { 0.0 }
}

/// Deterministic noise in [-1, 1].
pub fn noise(a: u64, b: u64) -> f64 {
    let mut x = a.wrapping_mul(0x9E37_79B9_7F4A_7C15) ^ b.wrapping_mul(0xC2B2_AE3D_27D4_EB4F) ^ 0x1656_67B1_9E37_79F9;
    x ^= x >> 33;
    x = x.wrapping_mul(0xFF51_AFD7_ED55_8CCD);
    x ^= x >> 33;
    (x % 20_001) as f64 / 10_000.0 - 1.0
}

/// A learned policy queried at step `t` predicts a chunk of `horizon` actions for t, t+1, ...
/// Each query commits to a slightly different plan (`offset`), and its error grows the further
/// ahead it looks (`drift` per step), as a real imitation policy's does.
#[derive(Clone, Copy, Debug)]
pub struct Policy {
    pub horizon: usize,
    pub offset: f64,
    pub drift: f64,
    pub seed: u64,
}

impl Policy {
    /// The chunk predicted at step t. It sees the nudge only if it has already happened.
    pub fn predict(&self, t: usize) -> Vec<f64> {
        let bias = self.offset * noise(self.seed, t as u64);
        let slope = self.drift * noise(self.seed + 1, t as u64);
        let seen = if t >= NUDGE_AT { NUDGE } else { 0.0 };
        (0..self.horizon).map(|i| target(t + i) + seen + bias + slope * i as f64).collect()
    }
}

/// Query once per chunk and execute it to the end, then query again.
pub fn run_open_loop(p: &Policy, steps: usize) -> Vec<f64> {
    let mut out = Vec::with_capacity(steps);
    while out.len() < steps {
        let t = out.len();
        out.extend(p.predict(t).into_iter().take(steps - t));
    }
    out
}

/// Query every step and execute only the first action of the newest chunk.
pub fn run_latest(p: &Policy, steps: usize) -> Vec<f64> {
    (0..steps).map(|t| p.predict(t)[0]).collect()
}

/// Mean distance from where the robot should be.
pub fn tracking_error(actions: &[f64]) -> f64 {
    actions.iter().enumerate().map(|(t, a)| (a - world(t)).abs()).sum::<f64>() / actions.len() as f64
}

/// Steps after the nudge until the robot is within `tol` of where it should be, and stays there
/// for the rest of `actions`.
pub fn reaction_steps(actions: &[f64], tol: f64) -> Option<usize> {
    let mut settle = None;
    for t in (NUDGE_AT..actions.len()).rev() {
        if (actions[t] - world(t)).abs() > tol {
            break;
        }
        settle = Some(t - NUDGE_AT);
    }
    settle
}

/// Mean size of the second difference: how jerky the executed motion is.
pub fn jerk(actions: &[f64]) -> f64 {
    let n = actions.len().saturating_sub(2).max(1);
    actions.windows(3).map(|w| (w[2] - 2.0 * w[1] + w[0]).abs()).sum::<f64>() / n as f64
}

Exercises

The stubs are in rust/ch11-behavior/src/exercises.rs.

Exercise 11.1

Temporal ensembling

Implement run_ensemble. The tests check that it is smoother than re-planning every step and more accurate than open-loop chunks. Then explain why a very large m makes the ensemble behave like executing old chunks, and what that does after the nudge.

Check your answer from the rust/ folder:

cargo test -p ch11-behavior --test exercises ex11_1
Hint 1

Keep one buffer per future step; each query appends its prediction to the buffers it covers.

Hint 2

Weights are exp(−m·i) with i = 0 for the oldest entry in the buffer.

Exercise 11.2

An honest error bar

Implement wilson. Then take three success rates quoted in this book's chapters 5 to 10 that report a trial count, and compute their intervals.

Check your answer from the rust/ folder:

cargo test -p ch11-behavior --test exercises ex11_2
Hint 1

Use the formula in the chapter; clamp the result to [0, 1].

Hint 2

With zero trials, the honest answer is the whole interval.

Exercise 11.3

How many trials?

Implement power. Use it to find how many trials per policy you would need to detect a 5-point improvement from 85 to 90 percent, and say what that implies for papers that report a few points of gain from 50 real-world trials.

Check your answer from the rust/ folder:

cargo test -p ch11-behavior --test exercises ex11_3
Hint 1

Sum the probabilities of every pair of outcomes (successes of A, successes of B) for which the test statistic exceeds the critical value.

Hint 2

Each outcome's probability is the product of two binomial probabilities.

What we are not sure about

Where the family ends. The boundary with chapter 5 is porous. Octo takes language through a small encoder; GEN-1 describes its architecture as building on the strengths of vision and language models without starting from one. This book keeps LBMs separate because the absence of a pretrained language backbone changes what data they can use and what instructions they can follow, but some models can be argued either way.

Whether scale replaces the VLM. GEN-0 and GEN-1 are the strongest bet that robot and human-activity data at sufficient scale can do without web pretraining. The evidence is the makers' own, on their own tasks, so the question is open.

How much the primitives matter now. Diffusion heads, flow matching and L1 regression all appear in the same model families, and OpenVLA-OFT reports regression matching diffusion in its setting (chapter 5). Which head is best may depend more on data than on the head.

Further reading

References

10 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.

  1. K. Black, N. Brown, D. Driess and 21 others π0: A Vision-Language-Action Flow Model for General Robot Control. arXiv 2410.24164 v4, 2024-10-31.
  2. C. Chi, Z. Xu, S. Feng and 5 others Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. arXiv 2303.04137 v5, 2023-03-07.
  3. TRI LBM Team, J. Barreiros, A. Beaulieu and 79 others A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation. arXiv 2507.05331 v1, 2025-07-07.
  4. T. Z. Zhao, V. Kumar, S. Levine, C. Finn Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv 2304.13705 v1, 2023-04-23.
  5. K. Black, M. Y. Galliker, S. Levine Real-Time Execution of Action Chunking Flow Policies. arXiv 2506.07339 v2, 2025-06-09.
  6. Octo Model Team, D. Ghosh, H. Walke and 16 others Octo: An Open-Source Generalist Robot Policy. arXiv 2405.12213 v2, 2024-05-20.
  7. Generalist AI GEN-0: Embodied Foundation Models That Scale with Physical Interaction. generalistai.com, 2025-11-04.
  8. Generalist AI GEN-1: Scaling Embodied Foundation Models to Mastery. generalistai.com, 2026-04-02.
  9. Robotics Center (SVRC) Best VLA Models 2026: Complete Vision-Language-Action Guide. roboticscenter.ai, 2026-04.
  10. R. Cadene, S. Aliberts, F. Capuano and 14 others LeRobot: An Open-Source Library for End-to-End Robot Learning. arXiv 2602.22818 v1, 2026-02-26.