One egocentric video in. Paired wrist and head views out. Fully annotated.

We capture a person doing a task, and we deliver robot-ready training data from that single capture. No teleoperation rig.

Egocentric (left), wrist camera (right)
Evidence

A wrist camera sees less of the room than a head camera. That is why it works.

Adding a wide third-person view helps a policy learn, and measurably harms it outside its training distribution. (Hsu et al., ICLR 2022)

When eye-in-hand human video was added to robot demonstrations, success rates rose by 58 percentage points on average. Several tasks went from not working at all.

Task Robot data only With eye-in-hand human video
Reaching 10% 90%
Cube grasping 0% 52%
Pick-and-place, unseen task 0% 80%

(Chen et al., Giving Robots a Hand)

Head video teaches a model what the world looks like. A wrist view teaches a policy what it will see when it acts.

We deliver both, from one recording, in the robot's coordinate frame.

Output

What we deliver

01

Hand and finger pose in 3D.

The position of the wrist and each finger joint in space, frame by frame.

02

Object masks with track IDs.

The exact pixels of each object, with the same ID kept on that object across every frame.

03

Contact events.

The frame where the hand grasps the object, and the frame where it lets go.

04

Task and subtask labels.

Named steps with start and end times.

05

Language descriptions.

A text description of each step, aligned to the time boundaries.

06

End-effector trajectories.

The gripper path in a robot-usable frame, retargeted to a generic gripper.

Capture

Capture where you need it

Capture needs a phone, a printed marker, and a trained operator. No teleoperation rig, no robot, no laboratory.

We recruit and train operators from the Southeast Asian remote-work market, where hundreds of thousands of workers are already registered on established platforms.

A trained team of eight can be capturing within about six weeks of a signed scope.

Pipeline

How it works

  1. Step 01

    Capture

    A person wears an egocentric camera and does the task.

  2. Step 02

    Extract

    The pipeline reconstructs the wrist view and the head view, time-synced.

  3. Step 03

    Annotate

    The pipeline adds the 6 annotation types above.

  4. Step 04

    Deliver

    The customer receives the dataset in a standard format.

Format

What lands on your disk

Every episode ships in LeRobot v3 format, with the units, the source of each field, and the quality evidence in the file.

One episode.

tree
episode_0001/
  meta/info.json      schema, units, provenance
  meta/tasks.parquet  task and subtask labels with time bounds
  meta/qc.json        quality evidence for this episode
  data/frames.parquet one row per timestep
  videos/             head view, rendered wrist view
  masks/              per-object mask track, stable ID

Each field says what it is, what units it is in, and whether it was measured or rendered.

meta/info.json
"observation.images.wrist": {
  "dtype": "video", "source": "rendered",
  "method": "gaussian_splat", "standoff_m": 0.15
},
"observation.hand_joints": {
  "shape": [21, 3], "units": "metres",
  "frame": "world", "source": "measured"
},
"action.ee_pose": {
  "shape": [7], "layout": "xyz + quaternion",
  "units": "metres", "frame": "world"
}

And the evidence, per episode.

meta/qc.json
{
  // distance from the solved camera pose to the printed marker
  "camera_pose_vs_marker": { "median_cm": null, "p90_cm": null },
  // share of frames the tracker solved
  "registration_rate":     null,
  // how far reprojected points land from where they were seen
  "reprojection_error_px": null,
  // frames where the wrist frame flipped about its roll axis
  "roll_flips":            null,
  // detected contact events against a human label
  "grasp_state_agreement": null,
  // set when the numbers above are measured, not before
  "verdict":               "pending"
}

Interested? Reach out at contact@humanpriors.com