Human demonstrations are attractive because people already solve a wide range of physical tasks efficiently. The harder question is how to record those demonstrations in a form that a model can use.

Ego data, or egocentric data, records those demonstrations from the actor's point of view and connects perception to movement over time.

A useful point of view

Egocentric capture follows the demonstrator's working viewpoint. It records what was visible at the moment of action and preserves the changing relationship between the observer, the hands, the objects, and the environment.

That alignment can reduce ambiguity. An external camera may show the whole scene, but it can lose the detail of what the person attended to or what was temporarily occluded from another viewpoint.

Demonstration is more than video

Video remains only one signal. Pose, motion, timing, task boundaries, and object state give the observation a structure that can support imitation learning, world modeling, and evaluation.

The collection system must also be comfortable enough for natural behavior. A technically rich device is less valuable if it changes how people move, shortens sessions, or cannot operate in the target environment.

Choose the right fidelity

Not every task needs every sensor. The right design balances signal quality, collection scale, operator burden, and processing cost. The training objective should decide which measurements are essential and which can be estimated downstream.