Ch.13Part III · Cross-cutting engineering
Data: the pyramid, scaling and provenance
Every family's strengths trace to which data layer it can use. This chapter makes that explicit.
- Math
- L1
- Sources
- 22
- Figures
- 5
- Exercises
- 3
- Last checked
- 9 Oct 2026
One boxed equation per idea
5 papers from 2026
Most are interactive
Rust, with tests
The field moves monthly
Papers cited, by half-year of first arXiv version
Robot trajectories
| Example | Reported size | Source |
|---|---|---|
| Open X-Embodiment | 1M+ trajectories, 22 embodiments | O'Neill et al. 2023, arXiv 2310.08864 |
| DROID | 76k trajectories, 350 hours, 564 scenes | Khazatsky et al. 2024, arXiv 2403.12945 |
| BridgeData V2 | 60,096 trajectories, 24 environments | Walke et al. 2023, arXiv 2308.12952 |
Who can use it: every family for fine-tuning; reactive VLAs (chapter 5) and large behavior models (chapter 11) learn mostly from this layer.
Why this chapter exists
Each family in Part II was defined by what it predicts. Each was limited by what it can learn from. Reactive VLAs need action-labelled robot data, the smallest and most expensive layer of the pyramid. World-action and latent models can learn from human video because they predict futures, or infer actions, instead of needing them. World models used as simulators can turn the largest layer, internet video, into a dynamics prior. Most of the pros and cons in chapters 5 to 11 trace back to this one fact.
This chapter treats data as engineering. How big are the datasets, really, and what does an "hour" of each contain? What is actually established about scaling, and what is a curve drawn through a few points? How do you mix sources that differ in size by a factor of a thousand? And why are reported hours so hard to compare?
After this chapter you will be able to:
- place a dataset on the pyramid of figure 13.1 and say which families can use it directly;
- read a scaling claim critically, using figure 13.3 to see why a few points cannot fix the shape of a curve;
- explain why an hour of teleoperation, an hour of human video and an hour of generated data are not the same unit, and how training mixtures cope.
Unlike AI for the digital world, there is no internet to scrape for large-scale robot interaction data with the physical world.
The three layers, in numbers
The pyramid has three layers. At the top are embodiment-specific trajectories: robot data with the actions that produced it. Open X-Embodiment pooled more than one million of them from 22 robot types, DROID collected 76,000 trajectories (350 hours) in the wild, and BridgeData V2 60,096 on a low-cost arm (Open X-Embodiment Collaboration et al. 2023, arXiv 2310.08864; Khazatsky et al. 2024, arXiv 2403.12945; Walke et al. 2023, arXiv 2308.12952). In the middle is task-relevant human activity: egocentric video of people using their hands, sometimes with tracked hand poses, such as Ego4D, EgoDex, EgoScale and UMI's hand-held gripper data (Grauman et al. 2021, arXiv 2110.07058; Hoque et al. 2025, arXiv 2505.11709; Zheng et al. 2026, arXiv 2602.16710; Chi et al. 2024, arXiv 2402.10329). At the base is task-agnostic internet video, which carries no actions and no task structure but shows how the world moves (NVIDIA et al. 2025, arXiv 2501.03575; Assran et al. 2025, arXiv 2506.09985).
Figure 13.2 puts the reported sizes on one axis. The axis is logarithmic because the sizes span five orders of magnitude, from DROID's 350 hours to the roughly 20 million hours of raw video behind Cosmos.
EgoScale: 20,854 hours
Action-labelled egocentric video with wrist and hand actions.
Source: Zheng et al. 2026, arXiv 2602.16710. Vendor claim.
Reported only in trajectories, so not on this axis: Open X-Embodiment, more than one million trajectories from 22 robot types; BridgeData V2, 60,096 trajectories; Covariant’s fleet data, “tens of millions of trajectories” (vendor claim).
Two things stand out. First, the robot layer has grown, but slowly. TRI's LBM study used about 1,700 hours (TRI LBM Team et al. 2025, arXiv 2507.05331), and LingBot-VLA 2.0's 50,000 hours of robot trajectories across 20 configurations is the largest robot-data figure in this book's sources (Wu et al. 2026, arXiv 2607.06403). Second, the human layer has grown fast, and mostly in company reports. GEN-0 describes 270,000 hours growing by 10,000 a weekvendor claim (Generalist AI 2025), GEN-1 half a million hours from wearable devicesvendor claim (Generalist AI 2026), and Dyna-2 more than a millionvendor claim (Dyna Robotics 2026). NVIDIA attributes GR00T N1.7's improved generalisation and language following to adding 20,000 hours of EgoScale human video to pretrainingvendor claim (NVIDIA 2026).
Scaling: what is established
The strongest public evidence that more human data makes better robots is EgoScale. It trained a VLA on 20,854 hours of action-labelled egocentric video and reports two curves (Zheng et al. 2026, arXiv 2602.16710). The first is offline: validation loss on held-out human data falls along a log-linear law with . The second is what matters, robot performance after post-training. Average task completion rises from 0.30 at 1,000 hours to 0.71 at 20,000 hours, through five measured points, "with no signs of saturation in the explored regime" (Zheng et al. 2026, arXiv 2602.16710).
A log-linear law says each doubling of data adds the same amount:
where and are fitted to the points. It is the natural first model, and EgoScale's loss fits it almost perfectly. The danger is in reading a law into the completion curve beyond its points. Figure 13.3 fits two different curves to the same five completion values.
- Log-linear predicts
- 89%
- Power law predicts
- 107%impossible: above the ceiling
Inside the measured range the two fits are nearly indistinguishable, each missing the points by about three points of completion. At 100,000 hours, the log-linear curve predicts about 89 percent completion and the power law about 108 percent, which is impossible. At GEN-0's 270,000 hours both are above 100 percent. No curve through a handful of points can say how robot performance behaves at ten or fifty times the data, and every real curve must bend before the ceiling. What the scaling evidence establishes is narrower and still important: within the explored range, more human data helped steadily, and the offline validation loss tracked real-robot performance (Zheng et al. 2026, arXiv 2602.16710). GEN-0 adds a different kind of claim, a "phase transition" at 7 billion parameters below which models stop improving with more datavendor claim (Generalist AI 2025). Dyna-2 claims a scaling law that transfers from human to robot datavendor claim (Dyna Robotics 2026). Neither has been independently reproduced in this book's sources.
Synthetic data
Chapter 8 covered the machinery: world models that rewrite appearance, generate new demonstrations and label them with inverse dynamics. For this chapter the question is accounting. NVIDIA reports generating 780,000 synthetic trajectories, which it equates to 6,500 hours of human demonstration, in 11 hours, and improving GR00T N1 by 40 percent when mixing them with real datavendor claim (NVIDIA Newsroom 2025). One industry survey reports that VLAs trained on 40 percent synthetic data matched policies trained on 100 percent real data, citing unnamed CMU and Stanford teamsindustry survey (SVRC (Robotics Center of Silicon Valley) 2026). The book has not found the primary papers, and treats the figure as reported, not established.
Beyond RGB
Most datasets in figure 13.2 are camera video, sometimes with robot joint angles (proprioception) and language. The other modalities buy specific things. Multiple views and depth reduce occlusion and give metric distances: every DROID episode has three cameras, calibration and depth (Khazatsky et al. 2024, arXiv 2403.12945). 3D point flows let one model learn across bodies (Huang et al. 2026, arXiv 2601.03782). Touch tells a model about contact that cameras cannot see, the subject of chapter 9's figure 9.2 (Higuera et al. 2026, arXiv 2602.06001). Sound tells it what is happening inside a bottle being filled (Zhang et al. 2025, arXiv 2512.08405). Human motion capture gives a whole-body controller its prior (Luo et al. 2025, arXiv 2511.07820). Each of these is rare compared with video, which is why the families that rely on them (chapter 9) have the least pretraining data.
An hour is not an hour
Data sizes are reported in hours because hours are easy to count. They are a poor common currency. Figure 13.4 shows what one hour of four kinds of data contains.
Teleoperation
DROID: 350 hours, 76,000 trajectories
- Episodes in one hour
- about 217, of about 17 seconds each
- Actions a robot can execute
- yes, for this robot (a Franka arm)
- Also recorded
- three RGB cameras, depth, calibration, language
- What an hour costs
- about $118 to collect in 2026 (industry survey)
Khazatsky et al. 2024; SVRC, State of Robotics 2026
Egocentric human video with hand tracking
EgoDex: 829 hours, 338,000 demonstrations
- Episodes in one hour
- about 408, of about 9 seconds each
- Actions a robot can execute
- no: 3D hand and finger poses, which must be mapped to a robot
- Also recorded
- 108,000 frames at 30 FPS, language, camera poses
- What an hour costs
- a person wearing a headset; no robot needed
Hoque et al. 2025
Generated trajectories
NVIDIA's blueprint: 780,000 trajectories, called 6,500 hours
- Episodes in one hour
- about 120, of about 30 seconds each
- Actions a robot can execute
- yes, for the robot the model was adapted to, if the generation is faithful (chapter 8)
- Also recorded
- whatever the generator renders
- What an hour costs
- about 6 seconds of generation at NVIDIA's reported rate
NVIDIA press release, 18 March 2025 (vendor claim)
Internet video
Cosmos: about 20 million raw hours, 100 million curated clips
- Episodes in one hour
- about 5 clips survive curation
- Actions a robot can execute
- none
- Also recorded
- whatever was filmed, for whatever reason
- What an hour costs
- storage and filtering compute
NVIDIA 2025, arXiv 2501.03575
Three differences matter. Density: an hour of EgoDex holds about twice as many demonstrations as an hour of DROID, because human demonstrations are shorter. Executability: only teleoperation and faithful generated data carry actions a robot can run directly. Human data needs mapping, through hand poses, latent actions or a learned retargeting. Cost: an hour of teleoperation cost about $118 to collect in 2026 according to one industry surveyindustry survey (SVRC (Robotics Center of Silicon Valley) 2026), while NVIDIA's reported generation rate is about six seconds of compute per "hour". When two models are compared by "hours of pretraining data", none of these differences is visible.
Mixing sources
A training run draws samples from all its sources at once, and has to decide how often to draw from each. The common answer is temperature sampling: give each dataset a share proportional to its size raised to a power ,
With every hour is equally likely to be drawn, so the largest dataset dominates. With every dataset gets the same share, so the smallest is repeated many times. Figure 13.5 shows the trade-off on a mixture built from sizes in this chapter.
- Most repeated dataset
- 5.3×DROID
- Robot-action data in the mix
- 41%teleoperation and generated trajectories
The robot-specific datasets are the smallest and the ones a robot most needs. Proportional sampling gives DROID about 1 percent of the samples. Uniform sampling shows it more than 16 times in one pass, long past the point where EgoScale saw small datasets overfit (Zheng et al. 2026, arXiv 2602.16710). In this mixture, of 0.55 is the smallest temperature that keeps every source below five repetitions. Exercise 13.2 finds it.
In code
The scaling-curve half of the crate is short: the five reported points and the two shapes of curve.
/// EgoScale's average task completion after post-training, against hours of human pretraining
/// data (figure 5, right, of arXiv 2602.16710). A vendor claim, measured on the makers' robots.
pub const EGOSCALE: [(f64, f64); 5] = [(1_000.0, 0.30), (2_000.0, 0.45), (4_000.0, 0.48), (10_000.0, 0.57), (20_000.0, 0.71)];
/// y = a + b ln(x): the shape EgoScale reports for its validation loss.
#[derive(Clone, Copy, Debug)]
pub struct LogLinear {
pub a: f64,
pub b: f64,
}
impl LogLinear {
pub fn fit(points: &[(f64, f64)]) -> LogLinear {
let xs: Vec<f64> = points.iter().map(|p| p.0.ln()).collect();
let n = points.len() as f64;
let mx = xs.iter().sum::<f64>() / n;
let my = points.iter().map(|p| p.1).sum::<f64>() / n;
let sxy: f64 = xs.iter().zip(points).map(|(x, p)| (x - mx) * (p.1 - my)).sum();
let sxx: f64 = xs.iter().map(|x| (x - mx).powi(2)).sum();
let b = sxy / sxx;
LogLinear { a: my - b * mx, b }
}
pub fn predict(&self, x: f64) -> f64 {
self.a + self.b * x.ln()
}
}
/// y = c x^k: the other common shape for a scaling curve.
#[derive(Clone, Copy, Debug)]
pub struct PowerLaw {
pub c: f64,
pub k: f64,
}
impl PowerLaw {
pub fn predict(&self, x: f64) -> f64 {
self.c * x.powf(self.k)
}
}
/// Root-mean-square error of a model on points.
pub fn rmse(points: &[(f64, f64)], f: impl Fn(f64) -> f64) -> f64 {
(points.iter().map(|p| (f(p.0) - p.1).powi(2)).sum::<f64>() / points.len() as f64).sqrt()
}Exercises
The stubs are in rust/ch13-data/src/exercises.rs.
Exercise 13.1
Temperature sampling
Implement mixture_weights. Then say which value of alpha you would start from for a VLA that must work on one specific robot, and why.
Check your answer from the rust/ folder:
cargo test -p ch13-data --test exercises ex13_1Hint 1
Raise each size to the power alpha, then divide by the sum.
Hint 2
Check the two ends: alpha 0 and alpha 1.
Exercise 13.2
How often is the smallest dataset repeated?
Implement max_epochs and alpha_for. Then double the training budget and find the new alpha. What does the answer imply for a lab that can afford to train much longer than its robot data justifies?
Check your answer from the rust/ folder:
cargo test -p ch13-data --test exercises ex13_2Hint 1
Repetitions of a dataset are the budget times its weight divided by its size.
Hint 2
Search alpha from 0 upwards in steps of 0.05 and stop at the first that meets the limit.
Exercise 13.3
The other curve
Implement fit_power_law. Then find the data scale at which each fitted curve first predicts 100 percent completion, and write two sentences a vendor could honestly put under a scaling plot.
Check your answer from the rust/ folder:
cargo test -p ch13-data --test exercises ex13_3Hint 1
Take logarithms of both coordinates and fit a straight line.
Hint 2
The intercept of that line is the log of c.
What we are not sure about
Whether human-data scaling holds across labs. EgoScale, GEN-0 and Dyna-2 all report that more human data keeps helping, each measured by its makers on their own robots. Independent replication at comparable scale is missing from this book's sources.
What an hour should be replaced by. Hours hide density, executability and cost. Episodes, frames or action-labelled transitions each fix one problem and create another. The field has no agreed unit, so comparisons of data scale across papers remain rough.
How much synthetic data is worth. The vendor's 40 percent improvement and the survey's 40 percent equivalence are the most quoted numbers, and neither has a public, controlled study behind it that this book could find.
Further reading
- EgoScale, the clearest public scaling study for human data (Zheng et al. 2026, arXiv 2602.16710).
- Open X-Embodiment, for how pooling robot data became normal (Open X-Embodiment Collaboration et al. 2023, arXiv 2310.08864).
- DROID, for what a well-documented robot dataset looks like (Khazatsky et al. 2024, arXiv 2403.12945).
- EgoDex, for human video collected with robot learning in mind (Hoque et al. 2025, arXiv 2505.11709).
- TRI's careful examination of LBMs, for pretraining mixtures and their effect measured carefully (TRI LBM Team et al. 2025, arXiv 2507.05331).
- The WAM survey's section on data (Lu et al. 2026, arXiv 2609.16074).
References
22 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.
- Open X-Embodiment Collaboration, A. O'Neill, A. Rehman and 291 others Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv 2310.08864 v9, 2023-10-13.
- A. Khazatsky, K. Pertsch, S. Nair and 98 others DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset. arXiv 2403.12945 v2, 2024-03-19.
- H. Walke, K. Black, A. Lee and 11 others BridgeData V2: A Dataset for Robot Learning at Scale. arXiv 2308.12952 v3, 2023-08-24.
- K. Grauman, A. Westbury, E. Byrne and 82 others Ego4D: Around the World in 3,000 Hours of Egocentric Video. arXiv 2110.07058 v3, 2021-10-13.
- R. Hoque, P. Huang, D. J. Yoon and 2 others EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video. arXiv 2505.11709 v3, 2025-05-16.
- R. Zheng, D. Niu, Y. Xie and 12 others EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data. arXiv 2602.16710 v1, 2026-02-18.
- C. Chi, Z. Xu, C. Pan and 5 others Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. arXiv 2402.10329 v3, 2024-02-15.
- NVIDIA, N. Agarwal, A. Ali and 75 others Cosmos World Foundation Model Platform for Physical AI. arXiv 2501.03575 v3, 2025-01-07.
- M. Assran, A. Bardes, D. Fan and 26 others V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. arXiv 2506.09985 v1, 2025-06-11.
- TRI LBM Team, J. Barreiros, A. Beaulieu and 79 others A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation. arXiv 2507.05331 v1, 2025-07-07.
- W. Wu, F. Wang, F. Lu and 21 others From Foundation to Application: Improving VLA Models in Practice. arXiv 2607.06403 v1, 2026-07-07.
- Generalist AI GEN-0: Embodied Foundation Models That Scale with Physical Interaction. generalistai.com, 2025-11-04.
- Generalist AI GEN-1: Scaling Embodied Foundation Models to Mastery. generalistai.com, 2026-04-02.
- Dyna Robotics Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models. dyna.co, 2026-08.
- NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
- NVIDIA Newsroom NVIDIA Announces Isaac GR00T N1, the World's First Open Humanoid Robot Foundation Model, and Simulation Frameworks to Speed Robot Development. nvidianews.nvidia.com, 2025-03-18.
- SVRC (Robotics Center of Silicon Valley) State of Robotics 2026. roboticscenter.ai, 2026-03.
- W. Huang, Y. Chao, A. Mousavian and 4 others PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation. arXiv 2601.03782 v1, 2026-01-07.
- C. Higuera, S. Arnaud, B. Boots and 3 others Visuo-Tactile World Models. arXiv 2602.06001 v1, 2026-02-05.
- F. Zhang, M. Gienger Learning Robot Manipulation from Audio World Models. arXiv 2512.08405 v1, 2025-12-09.
- Z. Luo, Y. Yuan, T. Wang and 25 others SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control. Science Robotics 11 (117), eaed4592 (2026). arXiv 2511.07820 v4, 2025-11-11.
- Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.