Data/Physical AI Data
Robot manipulation trajectories, human egocentric video, and action-conditioned footage at scale — for VLA models, world models, and manipulation policies. Built on the same collection, annotation and delivery infrastructure that serves our audio and vision lines.
50,000+
Hours
3M+
Episodes
500+
Unique Tasks
20+
Robot Embodiments
Task, object and scene taxonomies ship with every dataset.
Teleoperated and scripted-playback trajectories on physical robots, across mobile and fixed-base platforms. Proprioceptive state, action and gripper channels included.
| Data Type | Description |
|---|---|
| Bimanual, Mobile Base | Wheeled or tracked base with dual arms; cross-floor navigation combined with manipulation |
| Bimanual, Fixed Base | Tabletop pick-and-place, handover and tool use, with multi-view and wrist cameras |
| Single-Arm Industrial | Machine tending, sorting, palletizing and material handling |
| Humanoid, Whole-Body | Locomotion combined with manipulation, with joint-level state and action |
| Dexterous Hand | Multi-finger hands, in-hand reorientation, deformable objects and granular material |
| Failure & Recovery | Labeled failure, retry and correction trajectories, collected to an agreed ratio — for reward models and data selection |
| Long-Horizon Composite | Multi-step tasks with substep segmentation |
Head-mounted first-person recordings of human task execution, delivered as original camera output so the telemetry embedded in the source container is preserved.
| Data Type | Description |
|---|---|
| Head-Mounted RGB | 1080p+ at 30–60fps, original camera output |
| IMU-Synchronized | Accelerometer and gyroscope at approx. 200Hz, timestamped per sample |
| 3D Hand Pose | Per hand: wrist 6DoF and 21 keypoints, with MANO parameters available |
| Camera Trajectory | Per-frame 6DoF pose with camera intrinsics and distortion coefficients |
| Industrial & Warehouse | Assembly, picking, packing and machine tending, recorded at contracted worksites |
| Retail & Service | Store operations, shelf work, food preparation and counter service |
| Language-Annotated | Task, substep and frame-level alignment, in Japanese and English |
Video paired with a control signal — camera trajectory, human wrist and hand pose, or robot action — for learning environment dynamics rather than appearance. Selected out of a multi-hundred-terabyte HD library against a physical-event taxonomy, so a request returns a curated, event-labeled set.
| Data Type | Description |
|---|---|
| Control Track | Camera 6DoF trajectory, human wrist and hand pose, or robot action array, with sampling rate and time base |
| Contact & Deformation Dynamics | Pouring, cutting, stacking, deformable objects and granular material, collisions and tool contact, with contact-event timestamps |
| Single-Shot Long-Horizon | Continuous single-shot recordings with no scene cuts |
| Multi-View Synchronized | Synchronized multi-camera recordings with extrinsics and measured synchronization error |
Measured human motion and engine-generated trajectories, delivered as a separate layer with a recommended mixing ratio.
| Data Type | Description |
|---|---|
| Motion Capture | Marker-based and markerless, with joint tree, units and coordinate frame |
| Simulation-Generated | Engine, physics solver and domain-randomization ranges included |
Coverage is what decides whether a dataset trains. With production capacity across East Asia, South Asia, Europe and North America, we collect across the scenarios below.
Every collection programme starts from a quota matrix across the dimensions below, and the matrix ships with the data so you can split by domain from day one.
What we capture alongside the video.
| Channel | Specification |
|---|---|
| RGB video | 1080p+ at 30fps, 60fps for high-dynamic tasks; H.264 High or H.265 |
| IMU | Accelerometer and gyroscope at approx. 200Hz, timestamped per sample |
| 3D hand pose | Wrist 6DoF and 21 keypoints per hand; MANO parameters available |
| Camera trajectory | Per-frame 6DoF pose with intrinsics and distortion coefficients |
| Eye gaze | Per-frame gaze direction, and gaze depth on supported devices |
| Depth | Aligned depth maps with format and unit |
| Stereo / RGB-D | Baseline and per-camera calibration |
| Robot state and action | Joint positions and velocities, end-effector pose, gripper width, and the action representation with its control rate |
| Force and tactile | Sensor model, sampling rate and mounting position |
| Audio | Original track with sample rate and channel count |
| Language annotation | Task, substep and frame-level, in Japanese and English |
| Retargeting to your embodiment | Inverse kinematics for the wrist, joint mapping for the hand |
We generate the delivery layer against your stack, and agree the format before collection starts.
Formats
Included with every dataset
You already hold head-mounted footage. We make it trainable.
Four engagement models, and they compose — most programmes run two of them in sequence.
01
02
03
04