Celebrating 25 years of DDD's Excellence and Social Impact.

Human Egocentric Video Dataset for Embodied AI and Robotics

Stereo egocentric video with synchronized head and hand pose, captured on mixed-reality hardware and ready for your training pipeline.

Explore the Dataset

Conference Room

Kitchenette Pantry

Open Office Desk

Reception Desk

Our egocentric dataset is available to license immediately, delivering production-grade stereo video, synchronized head and hand pose, and VLA-ready annotations from real human manipulation tasks captured across real-world environments. 
rectangle 1 1
  • Rectified stereo video above 1080p with 21 to 26-joint hand tracking per hand
  • Delivered in HDF5 and LeRobot v2 formats for immediate pipeline ingestion
  • Commercial usage rights with PII removed at the point of capture
  • Purposeful manipulation episodes, so every clip earns its place in your training set
Environment Task Frames Duration Download
Conference Room Distribute notepads and pens around a conference table 3,674 61.2
Conference Room Erase a whiteboard and rewrite a 3-row agenda 6,296 104.9
Conference Room Set up a water carafe and six glasses on the conference table 4,136 68.9
Conference Room Set up six chairs around a conference table 5563 93
Kitchenette Pantry Heat a lunch container in a kitchenette microwave 6,649 110
Kitchenette Pantry Pour hot water from a dispenser for instant tea 5350 107
Open Office Desk Wipe down a microwave interior with a damp paper towel 4750 79
Open Office Desk Refill a self-inking stamp with new ink drops 3,734 62.2
Open Office Desk Replace a dry whiteboard marker tip cap and store in tray 4401 73
Reception Desk Sign a visitor into a paper logbook and issue a badge 3939 66

Why Choose DDD Egocentric Dataset

ODD Analysis 4

Mixed-Reality Headset Hardware

Captured using Pico 4 Ultra, Apple Vision Pro, Meta Quest 3, and Project Aria, devices with on-board SLAM, inside-out tracking, and hardware-level hand and head pose pipelines built for highest quality and precision manipulation capture.

ODD Analysis 3

Stereo Video with Full Pose Telemetry

Every episode delivers rectified stereo video above 1080p at 60fps, per-frame 6DoF head pose, a 21 to 26-joint hand skeleton per hand in both camera and world frames, and IMU data. Frame-to-pose synchronization is within 10 milliseconds at p95 across every episode.

ODD Analysis 2

Pipeline-Ready Delivery

Every episode ships as rectified stereo MP4 files, a raw HDF5 pose file with the full per-frame schema, and a complete LeRobot v2 dataset with a 371-dimensional float32 observation state.

ODD Analysis 1

Annotated for the Reasoning Layer

Every episode carries a natural-language task description and temporal sub-task segments of one to five seconds, each naming the exact object and action involved.

Human Egocentric Video Dataset, Ready for Your Training Pipeline

Scroll to Top