Data

AI Data for Every Kind of Model

Two data lines — LLM & multimodal, and physical AI — on one collection, annotation and delivery infrastructure. Choose a line to explore its catalog.

01

LLM & Multimodal Data

What the world looks, sounds and reads like.

Audio, image and video at web scale — for LLMs, multimodal and generative models.

  • Audio & Speech

    Multilingual speech, conversation, TTS, music and sound effects

  • Image & Vision

    Web-scale images, image-text pairs, portraits, advertising and app screenshots

  • Video & Motion

    HD footage, film & TV, digital humans and advertising video

Explore

02

Physical AI Data

How the world responds to action.

Real-world interaction data — for embodied AI and world models.

  • Real-Robot Manipulation

    Bimanual and single-arm platforms, mobile and fixed-base

  • Human Egocentric Video

    Head-mounted recordings with IMU, hand pose and camera motion

  • Action-Conditioned Video

    Video paired with a control signal, for world models

  • Motion Capture & Simulation

    Measured human motion and simulation-generated data

Explore