Skip to content

Ch.09Part II · The seven familiesGeometry-first

Geometry-first and structured world models

Predict in an explicit spatial or physical representation (points, Gaussians, particles, occupancy, touch) instead of RGB.

Math
L1

One boxed equation per idea

Sources
14

5 papers from 2026

Figures
5

Most are interactive

Exercises
3

Rust, with tests

Last checked
3 Oct 2026

The field moves monthly

Papers cited, by half-year of first arXiv version

camerahollow: hidden from the cameramugcloth over a box
Explicit
3D position of visible surfaces
Lost
hidden surfaces, unless several views are fused
Used by
PointWorld, ParticleFormer, GeoVLA
Cost
needs depth sensing or reconstruction
Figure 9.1Every representation makes some things explicit and throws others away; geometry-first models choose to make shape, position and contact explicit, at the price of the web-scale pretraining that pixels enjoy. One scene, a mug and a cloth draped over a box, seen from a camera at the top left. Choose a representation to see the scene in it, and what it keeps and loses.schematic, not measured Source: representations and examples from the papers cited in this chapter; scene drawn by this book.

Why this chapter exists

Chapters 6 to 8 predicted the future as video or as features learned from video. That choice inherits something valuable, the billions of hours of video on the web, and something awkward. A video frame is a picture of the world, not the world. It does not say how far away the mug is, whether the cloth is touching the box, or how hard the peg is pressing on the rim of the hole. For many tasks the picture is enough. For contact-rich manipulation it often is not, because the variables control depends on are exactly the ones a camera reports worst.

Geometry-first models predict in an explicit spatial or physical representation instead: point clouds, Gaussian splats, particles, occupancy grids, and in some cases touch or sound. They give up web-scale pretraining in exchange for states a physicist would recognise, predictions that can be checked against a ruler, and simulation that does not need to render a single image.

After this chapter you will be able to:

  • name the main geometric representations, and say with figure 9.1 what each makes explicit and what it loses;
  • explain with figure 9.2 why an RGB rollout can be underspecified for contact, using a concrete pixel-size argument;
  • score a predicted point cloud with chamfer distance and voxel IoU, and say why a point-by-point loss would be wrong.

ETH Zurich's robotics research aims to advance machines that can move, perceive and manipulate reliably in the real world.

Marco Hutter, professor at ETH Zurich's Robotic Systems Lab. NVIDIA Announces NVIDIA Isaac GR00T Reference Humanoid Robot for Academic Research, NVIDIA Newsroom, 31 May 2026.

The mechanism

A geometry-first world model is a forward dynamics model whose state is geometry. Write the scene as a set of points PtP_t, each a 3D position, possibly with a colour or material label. The model predicts how that set moves under an action:

P^t+1=W(Pt, at)\boxed{\hat P_{t+1} = W(P_t,\ a_t)}

The same line covers Gaussians (each element has a position, a shape and a colour), particles (each element is a lump of material, joined to its neighbours) and occupancy grids (each element is a filled cell of space). What they share is that the state is a list of things in space, not a grid of colours. A prediction can then be checked with geometry: is the mug where the model said it would be, to the millimetre?

That shift has a practical consequence for the action too. PointWorld represents actions as "3D point flows": the predicted motion of points on the robot's own body, rather than joint angles. That lets a single model learn from a Franka arm and a bimanual humanoid without an embodiment-specific action space (Huang et al. 2026, arXiv 2601.03782).

What RGB misses

The case for geometry is clearest at contact. Figure 9.2 puts a peg over a hole with a millimetre of clearance and looks at it two ways.

What the camera seeseach square is one camera pixel: 2 mmWhat geometry and touch seecontact:jammedpeg rests on the rim, overlapping it by 0.7 mm
Camera’s estimate of the offset
0.0 to 2.0 mmindistinguishable from a perfect alignment
Outcome
jamsa jam the camera cannot see
Figure 9.2Below the size of a camera pixel, two very different futures look identical; a model predicting in pixels cannot tell them apart, while a model predicting geometry and contact can. A peg above a hole, seen by a camera (left) and by geometry and touch (right). Move the peg off-centre, change the camera's pixel size and the hole's clearance.toy simulation Source: this book's toy, in rust/ch09-geometry; the argument follows the WAM survey (arXiv 2609.16074, section III-E) and VT-WM (2602.06001).

With 2 mm pixels and 1 mm of clearance, every offset between 0.6 and 1.9 mm jams the peg on the rim, and every one of them produces exactly the same camera image as a perfect alignment. Exercise 9.1 counts them. Real cameras are better than 2 mm per pixel at close range, and multiple views help; but contact also happens where the gripper's own fingers hide it, and the quantity that matters, the force on the rim, is not visible at all. A world model trained only on pixels has nothing to learn the difference from.

Visuo-tactile world models make this point with data. VT-WM adds touch to vision and reports 33 percent better object permanence and 29 percent better compliance with the laws of motion in its imagined rollouts, the failure modes being "objects disappearing, teleporting, or moving in ways that violate basic physics" under occlusion. In zero-shot real-robot planning it reports up to 35 percent higher success, with the largest gains on multi-step, contact-rich tasks (Higuera et al. 2026, arXiv 2602.06001). Sound can play the same role where vision is ambiguous: an audio world model predicts the rising pitch of a bottle being filled (Zhang et al. 2025, arXiv 2512.08405).

The representation ladder

Point clouds are the most direct. PointWorld is a large pretrained 3D world model: from one or a few RGB-D images and a sequence of actions, it forecasts per-point 3D displacements. It was trained on about 2 million trajectories and 500 hours of real and simulated manipulation, and runs in real time at 0.1 s per prediction (Huang et al. 2026, arXiv 2601.03782). ParticleFormer is a transformer over point clouds that models rigid, deformable and flexible materials interacting, trained directly on real robot perception data without elaborate scene reconstruction, and used for model predictive control (Huang et al. 2025, arXiv 2506.23126).

Gaussian splats represent a scene as many small coloured blobs that can be rendered back into images, so the model can be trained with photographs and still have an explicit 3D state. GWM predicts how Gaussian primitives propagate under robot actions with a latent diffusion transformer, and serves both as a training signal for imitation and as a simulator for model-based RL (Lu et al. 2025, arXiv 2508.17600). ManiGaussian++ uses a hierarchical Gaussian world model for two-armed manipulation (Yu et al. 2025, arXiv 2506.19842). 4DGS-WAM, from August 2026, separates dynamic objects from the static background so that only the objects need to be predicted. Its first experiments are on a driving dataset, KITTI-MOT, a reminder of how early this line is (Ma et al. 2026, arXiv 2608.25956).

Particles and physics go furthest towards a simulator. PIN-WM identifies a 3D rigid-body system from visual observations through differentiable physics, learns from a few task-agnostic interactions, and bridges the sim-to-real gap with physics-aware "digital cousins" (Li et al. 2025, arXiv 2504.16693). PhysWorld reconstructs a physical world from a generated video and grounds the video's motion in physically accurate actions with residual RL, without collecting real robot data (Mao et al. 2025, arXiv 2511.07416).

Other channels. Cloth-WM feeds surface normals into a DreamerV2-style world model to unfold cloth in the air, deployed zero-shot on a real robot (Rome et al. 2026, arXiv 2602.16675). On the policy side, SpatialVLA injects 3D position information into a VLA pretrained on 1.1 million real robot episodes (Qu et al. 2025, arXiv 2501.15830), and GeoVLA adds a point-cloud encoder next to the VLM (Sun et al. 2025, arXiv 2508.09071). Those two are reactive VLAs with better eyes, not world models, and figure 9.5 places them accordingly.

ModelRGB videoDepth, points, normalsGaussiansParticles or physicsTouchAudio
SpatialVLAJan 2025, 3D-aware VLA
GeoVLAAug 2025, 3D-aware VLA
PIN-WMApr 2025, physics-informed world model
ManiGaussian++Jun 2025, Gaussian world model
ParticleFormerJun 2025, point-cloud world model
GWMAug 2025, Gaussian world model
PhysWorldNov 2025, video to physics
Audio-WMDec 2025, audio world model
PointWorldJan 2026, 3D world model
VT-WMFeb 2026, visuo-tactile world model
Cloth-WMFeb 2026, world model for cloth
4D-WAMAug 2026, world-action model
4DGS-WAMAug 2026, object-centric WAM
  • consumes
  • predicts
  • consumes and predicts
  • training signal only
Figure 9.3Most geometry-first models still read RGB; what changes is what they predict, and several use geometry only to shape training. Thirteen models against six representations: whether each model consumes the representation, predicts it, does both, or uses it only as a training signal. Source: each model's abstract: SpatialVLA (arXiv 2501.15830), GeoVLA (2508.09071), PIN-WM (2504.16693), ManiGaussian++ (2506.19842), ParticleFormer (2506.23126), GWM (2508.17600), PhysWorld (2511.07416), Audio-WM (2512.08405), PointWorld (2601.03782), VT-WM (2602.06001), Cloth-WM (2602.16675), 4D-WAM (2608.08023), 4DGS-WAM (2608.25956).

Scoring a geometric prediction

A predicted video is scored pixel by pixel. A predicted point cloud cannot be, because its points come in no particular order: point 17 of the prediction has no reason to correspond to point 17 of the truth. The standard answer is the chamfer distance. For each predicted point, find the nearest real point; for each real point, the nearest predicted one; average both:

CD(P,Q)=1∣P∣∑p∈Pmin⁡q∈Q∥p−q∥  +  1∣Q∣∑q∈Qmin⁡p∈P∥q−p∥\boxed{\mathrm{CD}(P, Q) = \frac{1}{|P|}\sum_{p \in P} \min_{q \in Q} \lVert p - q \rVert \; + \; \frac{1}{|Q|}\sum_{q \in Q} \min_{p \in P} \lVert q - p \rVert}

The first term punishes predicted points with nothing real nearby, such as a handle where there is none. The second punishes real points the prediction missed, such as a handle the model forgot. Figure 9.4 lets you see both, and shows what goes wrong with a point-by-point loss once the order is shuffled. ParticleFormer's "hybrid point cloud reconstruction loss" is built from terms of this kind (Huang et al. 2025, arXiv 2506.23126).

hollow: the real mugfilled: the model’s predictiongrey lines: each predicted pointto its nearest real point
Chamfer distance
8.6 mmorder-free; unchanged by shuffling
Voxel IoU, 10 mm cells
0.89
Point-by-point error, assuming an order
6.0 mmmeaningless once the order is shuffled
Figure 9.4Point clouds have no order, so a geometric loss must match each point to its nearest neighbour; chamfer distance does, and a point-by-point error does not. A real mug (hollow) and a predicted one (filled), with a line from each predicted point to its nearest real point. Shift and rotate the prediction, drop its handle, or shuffle its point order.toy simulation Source: this book's toy; chamfer distance and voxel IoU as implemented in rust/ch09-geometry.

Occupancy is the coarse alternative: divide space into cells, mark each cell as filled or empty, and compare the sets with intersection over union. It forgives small errors inside a cell and punishes errors that cross cells, so a 10 mm grid scores a 5 mm shift almost as perfect (an IoU of 0.89 in the crate's mug) and a 10 mm shift as poor (0.25). Exercise 9.3 builds both.

The family on the four-component template

Past observationsand instructiono<t, ℓFuture geometrypoints, splats, particlesActionsa t:t+HPolicyp(a | o<t, ℓ)Visual planningp(o | o<t, ℓ)Forward dynamicsp(o | o<t, a)Inverse dynamicsp(a | o<t, o)

Geometric world model, then plan (ParticleFormer, PointWorld, PIN-WM). Learns how points, particles or rigid bodies move under an action, then searches for actions with model predictive control or trains a policy in it.

Figure 9.5Geometry can enter at three places: as input to a policy, as the state of a world model, or only as a training signal; only the second makes the family a world model. Chapter 4's template for four styles of geometry-first model. Select a style, and switch between training and inference. Source: placements from each paper's abstract, as in figure 9.3.

Good at, bad at

Why it exists. RGB rollouts are underspecified for contact-rich manipulation; geometry exposes the variables control depends on; and simulation without rendering is cheaper.

What it is good at. States are physically grounded and interpretable: a predicted point cloud can be measured. Geometry fits naturally with classical physics engines and with structured sensing such as depth cameras, LiDAR and tactile sensors. And it can be fast: PointWorld predicts in 0.1 s (Huang et al. 2026, arXiv 2601.03782).

What it is bad at. It loses most of the web-scale pretraining that video and RGB models inherit, because the internet has far less 3D and tactile data than video. It is sensor-specific: a model trained on one depth camera or one tactile skin does not transfer freely. And the research thread is young, with no shared benchmark suite.

In code

The peg-and-camera model behind figure 9.2 is three short methods.

rust/ch09-geometry/src/lib.rsWhat the camera reports, and what the contact does
/// A peg above a hole, seen by a camera whose pixels each cover `px_mm` of the scene. The peg is
/// `offset_mm` to the side of the hole's centre; the hole is `clearance_mm` wider than the peg.
#[derive(Clone, Copy, Debug, PartialEq)]
pub struct PegInHole {
    pub offset_mm: f64,
    pub clearance_mm: f64,
    pub px_mm: f64,
}

impl PegInHole {
    /// The pixel column the peg's edge falls in: all a camera can report about the offset.
    pub fn edge_pixel(&self) -> i64 {
        (self.offset_mm / self.px_mm).floor() as i64
    }
    /// The peg jams on the rim when it is off-centre by more than half the clearance.
    pub fn jams(&self) -> bool {
        self.offset_mm.abs() > self.clearance_mm / 2.0
    }
    /// How far the peg overlaps the rim, which is what a contact or force sensor would feel.
    pub fn overlap_mm(&self) -> f64 {
        (self.offset_mm.abs() - self.clearance_mm / 2.0).max(0.0)
    }
}

Exercises

The stubs are in rust/ch09-geometry/src/exercises.rs.

Exercise 9.1

Jams the camera cannot see

Implement invisible_jams. Then find, for a 1 mm clearance, the largest pixel size at which no jam is invisible, and explain in two sentences what that says about camera placement for insertion tasks.

Check your answer from the rust/ folder:

cargo test -p ch09-geometry --test exercises ex9_1
Hint 1

Compute the edge pixel of a centred peg once, then compare each offset's edge pixel to it.

Hint 2

Use the PegInHole methods rather than re-deriving the conditions.

Exercise 9.2

Chamfer distance

Implement chamfer. The tests check that it ignores point order and penalises a single stray point. Which term of the formula does a missing handle increase, and which does an extra one?

Check your answer from the rust/ folder:

cargo test -p ch09-geometry --test exercises ex9_2
Hint 1

A brute-force nearest-neighbour search is fine at these sizes.

Hint 2

Compute the mean in each direction separately, then add the two.

Exercise 9.3

Voxels and IoU

Implement voxelize and iou. Then compare the IoU of a 5 mm and a 10 mm shift of the mug at cell sizes of 5, 10 and 20 mm, and say which you would use to train a model and which to check a grasp is collision-free.

Check your answer from the rust/ folder:

cargo test -p ch09-geometry --test exercises ex9_3
Hint 1

Use floor, not truncation, so negative coordinates land in the right cell.

Hint 2

Intersection over union is the size of the intersection divided by the size of the union.

What we are not sure about

Whether geometry needs to be predicted at all. Several of the strongest results use geometry only as an input (SpatialVLA, GeoVLA) or only as a training signal (4D-WAM). It may be that the best design reads geometry and predicts in a learned latent space, which would fold this family into chapters 5 and 7.

Whether pretraining can catch up. The family's biggest handicap is data. PointWorld's 500 hours is large for 3D but tiny next to video. Whether 3D reconstruction of web video can supply the missing scale, as PhysWorld and EnerVerse suggest in different ways, is an open bet (Mao et al. 2025, arXiv 2511.07416; Huang et al. 2025, arXiv 2501.01895).

How to compare models. With no shared benchmark, claims of improvement are relative to each paper's own baselines. A common suite, ideally on real robots with contact-rich tasks, would do more for this family than any single model.

Further reading

References

14 sources, in order of first citation. Local copies are in reference/; arXiv entries link to the abstract page, and the version given is the one read for this chapter.

  1. W. Huang, Y. Chao, A. Mousavian and 4 others PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation. arXiv 2601.03782 v1, 2026-01-07.
  2. C. Higuera, S. Arnaud, B. Boots and 3 others Visuo-Tactile World Models. arXiv 2602.06001 v1, 2026-02-05.
  3. F. Zhang, M. Gienger Learning Robot Manipulation from Audio World Models. arXiv 2512.08405 v1, 2025-12-09.
  4. S. Huang, Q. Chen, X. Zhang and 2 others ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation. arXiv 2506.23126 v4, 2025-06-29.
  5. G. Lu, B. Jia, P. Li and 4 others GWM: Towards Scalable Gaussian World Models for Robotic Manipulation. arXiv 2508.17600 v2, 2025-08-25.
  6. T. Yu, G. Lu, Z. Yang and 7 others ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model. arXiv 2506.19842 v1, 2025-06-24.
  7. Y. Ma, Z. Xu, I. King 4DGS-WAM: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting. arXiv 2608.25956 v1, 2026-08-26.
  8. W. Li, H. Zhao, Z. Yu and 4 others PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation. arXiv 2504.16693 v2, 2025-04-23.
  9. J. Mao, S. He, H. Wu and 9 others Robot Learning from a Physical World Model. arXiv 2511.07416 v1, 2025-11-10.
  10. J. Rome, S. James, S. Ramamoorthy Learning to unfold cloth: Scaling up world models to deformable object manipulation. arXiv 2602.16675 v1, 2026-02-18.
  11. D. Qu, H. Song, Q. Chen and 8 others SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model. Robotics: Science and Systems, 2025. arXiv 2501.15830 v5, 2025-01-27.
  12. L. Sun, B. Xie, Y. Liu and 3 others GeoVLA: Empowering 3D Representations in Vision-Language-Action Models. arXiv 2508.09071 v2, 2025-08-12.
  13. S. Huang, L. Chen, P. Zhou and 8 others EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation. arXiv 2501.01895 v3, 2025-01-03.
  14. Z. Lu, H. Zhai, G. Wang and 13 others World-Action Models for Robot Learning and Control: A Survey. arXiv 2609.16074 v1, 2026-09-13.