Data/Physical AI Data

Data for Models That Act in the Physical World

Robot manipulation trajectories, human egocentric video, and action-conditioned footage at scale — for VLA models, world models, and manipulation policies. Built on the same collection, annotation and delivery infrastructure that serves our audio and vision lines.

50,000+

Hours

3M+

Episodes

500+

Unique Tasks

20+

Robot Embodiments

Task, object and scene taxonomies ship with every dataset.

Data Catalog

Real-Robot Manipulation

Teleoperated and scripted-playback trajectories on physical robots, across mobile and fixed-base platforms. Proprioceptive state, action and gripper channels included.

Data TypeDescription
Bimanual, Mobile BaseWheeled or tracked base with dual arms; cross-floor navigation combined with manipulation
Bimanual, Fixed BaseTabletop pick-and-place, handover and tool use, with multi-view and wrist cameras
Single-Arm IndustrialMachine tending, sorting, palletizing and material handling
Humanoid, Whole-BodyLocomotion combined with manipulation, with joint-level state and action
Dexterous HandMulti-finger hands, in-hand reorientation, deformable objects and granular material
Failure & RecoveryLabeled failure, retry and correction trajectories, collected to an agreed ratio — for reward models and data selection
Long-Horizon CompositeMulti-step tasks with substep segmentation

Human Egocentric Video

Head-mounted first-person recordings of human task execution, delivered as original camera output so the telemetry embedded in the source container is preserved.

Data TypeDescription
Head-Mounted RGB1080p+ at 30–60fps, original camera output
IMU-SynchronizedAccelerometer and gyroscope at approx. 200Hz, timestamped per sample
3D Hand PosePer hand: wrist 6DoF and 21 keypoints, with MANO parameters available
Camera TrajectoryPer-frame 6DoF pose with camera intrinsics and distortion coefficients
Industrial & WarehouseAssembly, picking, packing and machine tending, recorded at contracted worksites
Retail & ServiceStore operations, shelf work, food preparation and counter service
Language-AnnotatedTask, substep and frame-level alignment, in Japanese and English

Action-Conditioned & Interaction Video

Video paired with a control signal — camera trajectory, human wrist and hand pose, or robot action — for learning environment dynamics rather than appearance. Selected out of a multi-hundred-terabyte HD library against a physical-event taxonomy, so a request returns a curated, event-labeled set.

Data TypeDescription
Control TrackCamera 6DoF trajectory, human wrist and hand pose, or robot action array, with sampling rate and time base
Contact & Deformation DynamicsPouring, cutting, stacking, deformable objects and granular material, collisions and tool contact, with contact-event timestamps
Single-Shot Long-HorizonContinuous single-shot recordings with no scene cuts
Multi-View SynchronizedSynchronized multi-camera recordings with extrinsics and measured synchronization error

Motion Capture & Simulation

Measured human motion and engine-generated trajectories, delivered as a separate layer with a recommended mixing ratio.

Data TypeDescription
Motion CaptureMarker-based and markerless, with joint tree, units and coordinate frame
Simulation-GeneratedEngine, physics solver and domain-randomization ranges included

Task & Scene Coverage

Coverage is what decides whether a dataset trains. With production capacity across East Asia, South Asia, Europe and North America, we collect across the scenarios below.

  • Industrial Assembly & Machine Tending
  • Logistics & Warehouse
  • Retail & Store Operations
  • Kitchen & Food Preparation
  • Laboratory & Precision Handling
  • Household Chores
  • Outdoor & Mobile

Coverage Is Designed Before Collection

Every collection programme starts from a quota matrix across the dimensions below, and the matrix ships with the data so you can split by domain from day one.

  • Lighting — daylight, mixed, artificial, low-light
  • Recording site and background variants
  • Object classes and instances, with colour, material and size variants
  • Hand used and posture — standing, seated, crouching
  • Container and tool types
  • Motion speed — nominal, fast, interrupted

Channels

What we capture alongside the video.

ChannelSpecification
RGB video1080p+ at 30fps, 60fps for high-dynamic tasks; H.264 High or H.265
IMUAccelerometer and gyroscope at approx. 200Hz, timestamped per sample
3D hand poseWrist 6DoF and 21 keypoints per hand; MANO parameters available
Camera trajectoryPer-frame 6DoF pose with intrinsics and distortion coefficients
Eye gazePer-frame gaze direction, and gaze depth on supported devices
DepthAligned depth maps with format and unit
Stereo / RGB-DBaseline and per-camera calibration
Robot state and actionJoint positions and velocities, end-effector pose, gripper width, and the action representation with its control rate
Force and tactileSensor model, sampling rate and mounting position
AudioOriginal track with sample rate and channel count
Language annotationTask, substep and frame-level, in Japanese and English
Retargeting to your embodimentInverse kinematics for the wrist, joint mapping for the hand

Delivery

We generate the delivery layer against your stack, and agree the format before collection starts.

Formats

  • LeRobotDataset v3.0
  • RLDS / TFDS
  • HDF5
  • Zarr
  • WebDataset
  • MCAP (rosbag2)

Included with every dataset

  1. 01A datasheet covering coordinate frames, units and timestamps
  2. 02A runnable loader example for PyTorch or TensorFlow
  3. 03A converter to your own conventions
  4. 04Visualization scripts — skeleton overlay and 3D viewer
  5. 05Distribution statistics across duration, task, object and variant

Raw Data Upgrade

You already hold head-mounted footage. We make it trainable.

  • Telemetry extraction
  • 3D hand pose
  • Per-frame camera trajectory
  • Retargeting to your embodiment
  • Format conversion

How We Work With You

Four engagement models, and they compose — most programmes run two of them in sequence.

01

Existing Datasets

  • Your requirements mapped against what we can supply, within one business day
  • Sample package and datasheet within three business days
  • Scope and terms confirmed before you commit

02

Custom Collection

  • Scenario specification and quota matrix agreed first
  • Pilot batch before full-scale collection
  • Recording begins once the spec is fixed

03

Processing & Upgrade

  • Telemetry extraction, hand pose, camera trajectory, retargeting
  • Format conversion and structural restructuring
  • Applied to data you already own
See Raw Data Upgrade above

04

Data Programme Advisory

  • Capability-gap diagnosis on an evaluation set first, then mix design against that gap
  • Evaluation set and baseline delivered with the data
  • Data plus training guidance through the run