RGB and stereo video
First-person observations from the point of action.
Egocentric data infrastructure
Ego data captures first-person observation, motion, and action as synchronized multimodal episodes for world models, physical reinforcement learning, and action fine-tuning.
Definition
Ego data is short for egocentric data: first-person multimodal recordings that preserve what a person sees, how they move, and what they do over time.
Unlike ordinary video, an ego data episode connects visual observations to pose, inertial motion, geometry, task boundaries, and labels. That alignment makes the experience usable as Physical AI training data rather than footage alone.
Inside an episode
A program selects the signals required by its learning objective and task schema.
First-person observations from the point of action.
Motion and interaction expressed on a shared timeline.
High-frequency inertial measurements synchronized with vision.
Scene structure associated with each observation and action.
Explicit starts, outcomes, phases, and failure conditions.
Human-reviewed objects, actions, states, and relationships.
Point of view matters
| Dimension | Ego data | Third-person video |
|---|---|---|
| Viewpoint | Actor's point of view | External observer |
| Attention | What is visible during action | What the camera frames |
| Action context | Perception and motion stay aligned | Actions are observed from a distance |
| Occlusion | Matches the actor's real constraints | Depends on external camera placement |
| Training value | Observation-action episodes | Behavioral and scene context |
Zerolaw ego data pipeline
ZL Capture hardware records task-relevant first-person signals.
ZL Core cleans, slices, synchronizes, calibrates, and fuses sensor streams.
Human review adds task semantics and calibrates quality rules.
ZL Corpus organizes reusable ego data episodes for model development.
Training applications
Learn how environments, objects, and people evolve from continuous first-person observation.
Build task-grounded inputs for simulation, alignment, reward design, and policy learning.
Use human demonstrations as multimodal supervision for embodied and robotic behaviors.
Ego data FAQ
Ego data is short for egocentric data: first-person multimodal recordings that preserve what a person sees, how they move, and what they do over time.
It keeps perception, motion, and action aligned from the actor's point of view, providing supervision for world models, physical reinforcement learning, and action fine-tuning.
It can contain synchronized RGB or stereo video, head and hand pose, object tracks, IMU streams, depth or geometry, task boundaries, and semantic labels.
Third-person video observes a person from outside. Ego data records from the actor's point of view, preserving attention, reach, occlusion, motion, and action context as experienced during the task.
Zerolaw combines ZL Capture hardware with synchronization, calibration, cleaning, sensor fusion, annotation, and corpus construction to deliver task-ready Physical AI episodes.
Build with Zerolaw