Skip to content

Ch.13Part III · Cross-cutting engineering

Data: the pyramid, scaling and provenance

Every family's strengths trace to which data layer it can use. This chapter makes that explicit.

Math
L1

One boxed equation per idea

Sources
22

5 papers from 2026

Figures
5

Most are interactive

Exercises
3

Rust, with tests

Last checked
9 Oct 2026

The field moves monthly

Papers cited, by half-year of first arXiv version

Robot trajectoriesaction-labelled, smallest, most expensiveEgocentric human videotask-relevant, no robot actionsInternet videotask-agnostic, no actions, largestsmallest, most expensivelargest, cheapest

Robot trajectories

ExampleReported sizeSource
Open X-Embodiment1M+ trajectories, 22 embodimentsO'Neill et al. 2023, arXiv 2310.08864
DROID76k trajectories, 350 hours, 564 scenesKhazatsky et al. 2024, arXiv 2403.12945
BridgeData V260,096 trajectories, 24 environmentsWalke et al. 2023, arXiv 2308.12952

Who can use it: every family for fine-tuning; reactive VLAs (chapter 5) and large behavior models (chapter 11) learn mostly from this layer.

Figure 13.1Most of what separates the families is which layer of this pyramid they can learn from. The data pyramid from chapter 2, repeated because this chapter is about it. Select a layer for its datasets, their reported sizes and the families that can use it. Source: sizes from each dataset's paper; EgoScale and Dyna-2 sizes are vendor claims. Same figure as 2.2.

Why this chapter exists

Each family in Part II was defined by what it predicts. Each was limited by what it can learn from. Reactive VLAs need action-labelled robot data, the smallest and most expensive layer of the pyramid. World-action and latent models can learn from human video because they predict futures, or infer actions, instead of needing them. World models used as simulators can turn the largest layer, internet video, into a dynamics prior. Most of the pros and cons in chapters 5 to 11 trace back to this one fact.

This chapter treats data as engineering. How big are the datasets, really, and what does an "hour" of each contain? What is actually established about scaling, and what is a curve drawn through a few points? How do you mix sources that differ in size by a factor of a thousand? And why are reported hours so hard to compare?

After this chapter you will be able to:

  • place a dataset on the pyramid of figure 13.1 and say which families can use it directly;
  • read a scaling claim critically, using figure 13.3 to see why a few points cannot fix the shape of a curve;
  • explain why an hour of teleoperation, an hour of human video and an hour of generated data are not the same unit, and how training mixtures cope.

Unlike AI for the digital world, there is no internet to scrape for large-scale robot interaction data with the physical world.

Peter Chen, chief executive officer and co-founder of Covariant. Covariant introduces RFM-1 to give robots the human-like ability to reason, Covariant announcement, 11 March 2024.

The three layers, in numbers

The pyramid has three layers. At the top are embodiment-specific trajectories: robot data with the actions that produced it. Open X-Embodiment pooled more than one million of them from 22 robot types, DROID collected 76,000 trajectories (350 hours) in the wild, and BridgeData V2 60,096 on a low-cost arm (Open X-Embodiment Collaboration et al. 2023, arXiv 2310.08864; Khazatsky et al. 2024, arXiv 2403.12945; Walke et al. 2023, arXiv 2308.12952). In the middle is task-relevant human activity: egocentric video of people using their hands, sometimes with tracked hand poses, such as Ego4D, EgoDex, EgoScale and UMI's hand-held gripper data (Grauman et al. 2021, arXiv 2110.07058; Hoque et al. 2025, arXiv 2505.11709; Zheng et al. 2026, arXiv 2602.16710; Chi et al. 2024, arXiv 2402.10329). At the base is task-agnostic internet video, which carries no actions and no task structure but shows how the world moves (NVIDIA et al. 2025, arXiv 2501.03575; Assran et al. 2025, arXiv 2506.09985).

Figure 13.2 puts the reported sizes on one axis. The axis is logarithmic because the sizes span five orders of magnitude, from DROID's 350 hours to the roughly 20 million hours of raw video behind Cosmos.

100 h1k h10k h100k h1M h10M hhours, log scale: each gridline is ten times the lastRobot trajectoriesDROID350TRI-Ramen (TRI's own)545MolmoAct2 bimanual720TRI LBM mixture1,695LingBot-VLA 2.0 robot50,000Human activityEgoDex829Ego4D3,670LingBot-VLA 2.0 human10,000EgoScale20,854DreamDojo44,000GEN-0270,000+GEN-1500,000Dyna-21,000,000+Internet videoV-JEPA 21,000,000+Cosmos raw video20,000,000Synthetic and capturePointWorld500SONIC motion capture700GR00T synthetic6,500

EgoScale: 20,854 hours

Action-labelled egocentric video with wrist and hand actions.

Source: Zheng et al. 2026, arXiv 2602.16710. Vendor claim.

Reported only in trajectories, so not on this axis: Open X-Embodiment, more than one million trajectories from 22 robot types; BridgeData V2, 60,096 trajectories; Covariant’s fleet data, “tens of millions of trajectories” (vendor claim).

Figure 13.2Robot data is measured in hundreds or thousands of hours, human data in tens of thousands to a million, internet video in millions; the largest human-data numbers are vendor claims. Reported dataset sizes in hours, grouped by layer, on a log scale. Hollow markers are a company's claims about its own data. Select a dataset for what its hours contain and the source. Source: each dataset's paper or announcement, as listed per dataset; numbers as stated, not normalised.

Two things stand out. First, the robot layer has grown, but slowly. TRI's LBM study used about 1,700 hours (TRI LBM Team et al. 2025, arXiv 2507.05331), and LingBot-VLA 2.0's 50,000 hours of robot trajectories across 20 configurations is the largest robot-data figure in this book's sources (Wu et al. 2026, arXiv 2607.06403). Second, the human layer has grown fast, and mostly in company reports. GEN-0 describes 270,000 hours growing by 10,000 a weekvendor claim (Generalist AI 2025), GEN-1 half a million hours from wearable devicesvendor claim (Generalist AI 2026), and Dyna-2 more than a millionvendor claim (Dyna Robotics 2026). NVIDIA attributes GR00T N1.7's improved generalisation and language following to adding 20,000 hours of EgoScale human video to pretrainingvendor claim (NVIDIA 2026).

Scaling: what is established

The strongest public evidence that more human data makes better robots is EgoScale. It trained a VLA on 20,854 hours of action-labelled egocentric video and reports two curves (Zheng et al. 2026, arXiv 2602.16710). The first is offline: validation loss on held-out human data falls along a log-linear law with R2=0.9983R^2 = 0.9983. The second is what matters, robot performance after post-training. Average task completion rises from 0.30 at 1,000 hours to 0.71 at 20,000 hours, through five measured points, "with no signs of saturation in the explored regime" (Zheng et al. 2026, arXiv 2602.16710).

A log-linear law says each doubling of data adds the same amount:

y=a+bln⁡(hours)\boxed{y = a + b \ln(\text{hours})}

where aa and bb are fitted to the points. It is the natural first model, and EgoScale's loss fits it almost perfectly. The danger is in reading a law into the completion curve beyond its points. Figure 13.3 fits two different curves to the same five completion values.

measured range, 1k to 20k hours0%50%100%150%100%: every task completed, the ceiling1k h10k h100k h1.0M hhours of human pretraining data (log scale)GEN-0 dataDyna-2 data
Log-linear predicts
89%
Power law predicts
107%impossible: above the ceiling
Figure 13.3Two different curves fit EgoScale's five points about equally well and disagree completely beyond them; a scaling curve says nothing outside its measured range. EgoScale's average task completion at five data scales, with a log-linear fit and a power-law fit. Drag the slider to extrapolate, and compare with the data scales other companies now report. Source: points from Zheng et al. 2026 (arXiv 2602.16710), figure 5 (vendor claim); fits in rust/ch13-data; GEN-0 and Dyna-2 scales from their makers' reports.

Inside the measured range the two fits are nearly indistinguishable, each missing the points by about three points of completion. At 100,000 hours, the log-linear curve predicts about 89 percent completion and the power law about 108 percent, which is impossible. At GEN-0's 270,000 hours both are above 100 percent. No curve through a handful of points can say how robot performance behaves at ten or fifty times the data, and every real curve must bend before the ceiling. What the scaling evidence establishes is narrower and still important: within the explored range, more human data helped steadily, and the offline validation loss tracked real-robot performance (Zheng et al. 2026, arXiv 2602.16710). GEN-0 adds a different kind of claim, a "phase transition" at 7 billion parameters below which models stop improving with more datavendor claim (Generalist AI 2025). Dyna-2 claims a scaling law that transfers from human to robot datavendor claim (Dyna Robotics 2026). Neither has been independently reproduced in this book's sources.

Synthetic data

Chapter 8 covered the machinery: world models that rewrite appearance, generate new demonstrations and label them with inverse dynamics. For this chapter the question is accounting. NVIDIA reports generating 780,000 synthetic trajectories, which it equates to 6,500 hours of human demonstration, in 11 hours, and improving GR00T N1 by 40 percent when mixing them with real datavendor claim (NVIDIA Newsroom 2025). One industry survey reports that VLAs trained on 40 percent synthetic data matched policies trained on 100 percent real data, citing unnamed CMU and Stanford teamsindustry survey (SVRC (Robotics Center of Silicon Valley) 2026). The book has not found the primary papers, and treats the figure as reported, not established.

Beyond RGB

Most datasets in figure 13.2 are camera video, sometimes with robot joint angles (proprioception) and language. The other modalities buy specific things. Multiple views and depth reduce occlusion and give metric distances: every DROID episode has three cameras, calibration and depth (Khazatsky et al. 2024, arXiv 2403.12945). 3D point flows let one model learn across bodies (Huang et al. 2026, arXiv 2601.03782). Touch tells a model about contact that cameras cannot see, the subject of chapter 9's figure 9.2 (Higuera et al. 2026, arXiv 2602.06001). Sound tells it what is happening inside a bottle being filled (Zhang et al. 2025, arXiv 2512.08405). Human motion capture gives a whole-body controller its prior (Luo et al. 2025, arXiv 2511.07820). Each of these is rare compared with video, which is why the families that rely on them (chapter 9) have the least pretraining data.

An hour is not an hour

Data sizes are reported in hours because hours are easy to count. They are a poor common currency. Figure 13.4 shows what one hour of four kinds of data contains.

Teleoperation

DROID: 350 hours, 76,000 trajectories

Episodes in one hour
about 217, of about 17 seconds each
Actions a robot can execute
yes, for this robot (a Franka arm)
Also recorded
three RGB cameras, depth, calibration, language
What an hour costs
about $118 to collect in 2026 (industry survey)

Khazatsky et al. 2024; SVRC, State of Robotics 2026

Egocentric human video with hand tracking

EgoDex: 829 hours, 338,000 demonstrations

Episodes in one hour
about 408, of about 9 seconds each
Actions a robot can execute
no: 3D hand and finger poses, which must be mapped to a robot
Also recorded
108,000 frames at 30 FPS, language, camera poses
What an hour costs
a person wearing a headset; no robot needed

Hoque et al. 2025

Generated trajectories

NVIDIA's blueprint: 780,000 trajectories, called 6,500 hours

Episodes in one hour
about 120, of about 30 seconds each
Actions a robot can execute
yes, for the robot the model was adapted to, if the generation is faithful (chapter 8)
Also recorded
whatever the generator renders
What an hour costs
about 6 seconds of generation at NVIDIA's reported rate

NVIDIA press release, 18 March 2025 (vendor claim)

Internet video

Cosmos: about 20 million raw hours, 100 million curated clips

Episodes in one hour
about 5 clips survive curation
Actions a robot can execute
none
Also recorded
whatever was filmed, for whatever reason
What an hour costs
storage and filtering compute

NVIDIA 2025, arXiv 2501.03575

Figure 13.4An hour of teleoperation, an hour of tracked human video, an hour of generated trajectories and an hour of internet video differ in episodes, in whether they carry executable actions, and in cost by orders of magnitude. What one reported hour contains for four kinds of data, computed from each source's own totals. Source: DROID (arXiv 2403.12945); EgoDex (2505.11709); NVIDIA press release, 18 March 2025 (vendor claim); Cosmos (2501.03575); teleoperation cost from SVRC, State of Robotics 2026 (industry survey).

Three differences matter. Density: an hour of EgoDex holds about twice as many demonstrations as an hour of DROID, because human demonstrations are shorter. Executability: only teleoperation and faithful generated data carry actions a robot can run directly. Human data needs mapping, through hand poses, latent actions or a learned retargeting. Cost: an hour of teleoperation cost about $118 to collect in 2026 according to one industry surveyindustry survey (SVRC (Robotics Center of Silicon Valley) 2026), while NVIDIA's reported generation rate is about six seconds of compute per "hour". When two models are compared by "hours of pretraining data", none of these differences is visible.

Mixing sources

A training run draws samples from all its sources at once, and has to decide how often to draw from each. The common answer is temperature sampling: give each dataset a share proportional to its size raised to a power α\alpha,

wi∝hoursi α\boxed{w_i \propto \text{hours}_i^{\,\alpha}}

With α=1\alpha = 1 every hour is equally likely to be drawn, so the largest dataset dominates. With α=0\alpha = 0 every dataset gets the same share, so the smallest is repeated many times. Figure 13.5 shows the trade-off on a mixture built from sizes in this chapter.

Share of training samplesSynthetic 27%EgoScale 49%Times each dataset is seen in one pass of 29,078 hoursDROID350 h, teleoperation5.3×TRI-Ramen545 h, teleoperation4.2×EgoDex829 h, human video3.4×Synthetic6,500 h, generated1.2×EgoScale20,854 h, human video0.7×5×: the limit in exercise 13.2
Most repeated dataset
5.3×DROID
Robot-action data in the mix
41%teleoperation and generated trajectories
Figure 13.5Sampling in proportion to size drowns the small, robot-specific datasets; sampling uniformly repeats them until they overfit. The temperature sets where between the two a mixture sits. A five-source mixture: two teleoperation datasets, tracked human video, generated trajectories and action-labelled human video. Change the sampling temperature and watch each source's share and how often it is repeated.toy simulation Source: sizes from figure 13.2; temperature sampling as implemented in rust/ch13-data.

The robot-specific datasets are the smallest and the ones a robot most needs. Proportional sampling gives DROID about 1 percent of the samples. Uniform sampling shows it more than 16 times in one pass, long past the point where EgoScale saw small datasets overfit (Zheng et al. 2026, arXiv 2602.16710). In this mixture, α\alpha of 0.55 is the smallest temperature that keeps every source below five repetitions. Exercise 13.2 finds it.

In code

The scaling-curve half of the crate is short: the five reported points and the two shapes of curve.

rust/ch13-data/src/lib.rsEgoScale's reported points and two models of a scaling curve
/// EgoScale's average task completion after post-training, against hours of human pretraining
/// data (figure 5, right, of arXiv 2602.16710). A vendor claim, measured on the makers' robots.
pub const EGOSCALE: [(f64, f64); 5] = [(1_000.0, 0.30), (2_000.0, 0.45), (4_000.0, 0.48), (10_000.0, 0.57), (20_000.0, 0.71)];

/// y = a + b ln(x): the shape EgoScale reports for its validation loss.
#[derive(Clone, Copy, Debug)]
pub struct LogLinear {
    pub a: f64,
    pub b: f64,
}

impl LogLinear {
    pub fn fit(points: &[(f64, f64)]) -> LogLinear {
        let xs: Vec<f64> = points.iter().map(|p| p.0.ln()).collect();
        let n = points.len() as f64;
        let mx = xs.iter().sum::<f64>() / n;
        let my = points.iter().map(|p| p.1).sum::<f64>() / n;
        let sxy: f64 = xs.iter().zip(points).map(|(x, p)| (x - mx) * (p.1 - my)).sum();
        let sxx: f64 = xs.iter().map(|x| (x - mx).powi(2)).sum();
        let b = sxy / sxx;
        LogLinear { a: my - b * mx, b }
    }
    pub fn predict(&self, x: f64) -> f64 {
        self.a + self.b * x.ln()
    }
}

/// y = c x^k: the other common shape for a scaling curve.
#[derive(Clone, Copy, Debug)]
pub struct PowerLaw {
    pub c: f64,
    pub k: f64,
}

impl PowerLaw {
    pub fn predict(&self, x: f64) -> f64 {
        self.c * x.powf(self.k)
    }
}

/// Root-mean-square error of a model on points.
pub fn rmse(points: &[(f64, f64)], f: impl Fn(f64) -> f64) -> f64 {
    (points.iter().map(|p| (f(p.0) - p.1).powi(2)).sum::<f64>() / points.len() as f64).sqrt()
}

Exercises

The stubs are in rust/ch13-data/src/exercises.rs.

Exercise 13.1

Temperature sampling

Implement mixture_weights. Then say which value of alpha you would start from for a VLA that must work on one specific robot, and why.

Check your answer from the rust/ folder:

cargo test -p ch13-data --test exercises ex13_1
Hint 1

Raise each size to the power alpha, then divide by the sum.

Hint 2

Check the two ends: alpha 0 and alpha 1.

Exercise 13.2

How often is the smallest dataset repeated?

Implement max_epochs and alpha_for. Then double the training budget and find the new alpha. What does the answer imply for a lab that can afford to train much longer than its robot data justifies?

Check your answer from the rust/ folder:

cargo test -p ch13-data --test exercises ex13_2
Hint 1

Repetitions of a dataset are the budget times its weight divided by its size.

Hint 2

Search alpha from 0 upwards in steps of 0.05 and stop at the first that meets the limit.

Exercise 13.3

The other curve

Implement fit_power_law. Then find the data scale at which each fitted curve first predicts 100 percent completion, and write two sentences a vendor could honestly put under a scaling plot.

Check your answer from the rust/ folder:

cargo test -p ch13-data --test exercises ex13_3
Hint 1

Take logarithms of both coordinates and fit a straight line.

Hint 2

The intercept of that line is the log of c.

What we are not sure about

Whether human-data scaling holds across labs. EgoScale, GEN-0 and Dyna-2 all report that more human data keeps helping, each measured by its makers on their own robots. Independent replication at comparable scale is missing from this book's sources.

What an hour should be replaced by. Hours hide density, executability and cost. Episodes, frames or action-labelled transitions each fix one problem and create another. The field has no agreed unit, so comparisons of data scale across papers remain rough.

How much synthetic data is worth. The vendor's 40 percent improvement and the survey's 40 percent equivalence are the most quoted numbers, and neither has a public, controlled study behind it that this book could find.

Further reading

References

22 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.

  1. Open X-Embodiment Collaboration, A. O'Neill, A. Rehman and 291 others Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv 2310.08864 v9, 2023-10-13.
  2. A. Khazatsky, K. Pertsch, S. Nair and 98 others DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset. arXiv 2403.12945 v2, 2024-03-19.
  3. H. Walke, K. Black, A. Lee and 11 others BridgeData V2: A Dataset for Robot Learning at Scale. arXiv 2308.12952 v3, 2023-08-24.
  4. K. Grauman, A. Westbury, E. Byrne and 82 others Ego4D: Around the World in 3,000 Hours of Egocentric Video. arXiv 2110.07058 v3, 2021-10-13.
  5. R. Hoque, P. Huang, D. J. Yoon and 2 others EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video. arXiv 2505.11709 v3, 2025-05-16.
  6. R. Zheng, D. Niu, Y. Xie and 12 others EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data. arXiv 2602.16710 v1, 2026-02-18.
  7. C. Chi, Z. Xu, C. Pan and 5 others Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. arXiv 2402.10329 v3, 2024-02-15.
  8. NVIDIA, N. Agarwal, A. Ali and 75 others Cosmos World Foundation Model Platform for Physical AI. arXiv 2501.03575 v3, 2025-01-07.
  9. M. Assran, A. Bardes, D. Fan and 26 others V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. arXiv 2506.09985 v1, 2025-06-11.
  10. TRI LBM Team, J. Barreiros, A. Beaulieu and 79 others A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation. arXiv 2507.05331 v1, 2025-07-07.
  11. W. Wu, F. Wang, F. Lu and 21 others From Foundation to Application: Improving VLA Models in Practice. arXiv 2607.06403 v1, 2026-07-07.
  12. Generalist AI GEN-0: Embodied Foundation Models That Scale with Physical Interaction. generalistai.com, 2025-11-04.
  13. Generalist AI GEN-1: Scaling Embodied Foundation Models to Mastery. generalistai.com, 2026-04-02.
  14. Dyna Robotics Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models. dyna.co, 2026-08.
  15. NVIDIA NVIDIA Isaac GR00T N1.7 (Isaac-GR00T repository). github.com, 2026-04-17.
  16. NVIDIA Newsroom NVIDIA Announces Isaac GR00T N1, the World's First Open Humanoid Robot Foundation Model, and Simulation Frameworks to Speed Robot Development. nvidianews.nvidia.com, 2025-03-18.
  17. SVRC (Robotics Center of Silicon Valley) State of Robotics 2026. roboticscenter.ai, 2026-03.
  18. W. Huang, Y. Chao, A. Mousavian and 4 others PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation. arXiv 2601.03782 v1, 2026-01-07.
  19. C. Higuera, S. Arnaud, B. Boots and 3 others Visuo-Tactile World Models. arXiv 2602.06001 v1, 2026-02-05.
  20. F. Zhang, M. Gienger Learning Robot Manipulation from Audio World Models. arXiv 2512.08405 v1, 2025-12-09.
  21. Z. Luo, Y. Yuan, T. Wang and 25 others SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control. Science Robotics 11 (117), eaed4592 (2026). arXiv 2511.07820 v4, 2025-11-11.
  22. Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.