Hand and finger pose in 3D.
The position of the wrist and each finger joint in space, frame by frame.
We capture a person doing a task, and we deliver robot-ready training data from that single capture. No teleoperation rig.
Interested? Reach out.
Email usA wrist camera sees less of the room than a head camera. That is why it works.
Adding a wide third-person view helps a policy learn, and measurably harms it outside its training distribution. (Hsu et al., ICLR 2022)
When eye-in-hand human video was added to robot demonstrations, success rates rose by 58 percentage points on average. Several tasks went from not working at all.
| Task | Robot data only | With eye-in-hand human video |
|---|---|---|
| Reaching | 10% | 90% |
| Cube grasping | 0% | 52% |
| Pick-and-place, unseen task | 0% | 80% |
(Chen et al., Giving Robots a Hand)
Head video teaches a model what the world looks like. A wrist view teaches a policy what it will see when it acts.
We deliver both, from one recording, in the robot's coordinate frame.
The position of the wrist and each finger joint in space, frame by frame.
The exact pixels of each object, with the same ID kept on that object across every frame.
The frame where the hand grasps the object, and the frame where it lets go.
Named steps with start and end times.
A text description of each step, aligned to the time boundaries.
The gripper path in a robot-usable frame, retargeted to a generic gripper.
Capture needs a phone, a printed marker, and a trained operator. No teleoperation rig, no robot, no laboratory.
We recruit and train operators from the Southeast Asian remote-work market, where hundreds of thousands of workers are already registered on established platforms.
A trained team of eight can be capturing within about six weeks of a signed scope.
A person wears an egocentric camera and does the task.
The pipeline reconstructs the wrist view and the head view, time-synced.
The pipeline adds the 6 annotation types above.
The customer receives the dataset in a standard format.
Every episode ships in LeRobot v3 format, with the units, the source of each field, and the quality evidence in the file.
One episode.
episode_0001/
meta/info.json schema, units, provenance
meta/tasks.parquet task and subtask labels with time bounds
meta/qc.json quality evidence for this episode
data/frames.parquet one row per timestep
videos/ head view, rendered wrist view
masks/ per-object mask track, stable ID
Each field says what it is, what units it is in, and whether it was measured or rendered.
"observation.images.wrist": {
"dtype": "video", "source": "rendered",
"method": "gaussian_splat", "standoff_m": 0.15
},
"observation.hand_joints": {
"shape": [21, 3], "units": "metres",
"frame": "world", "source": "measured"
},
"action.ee_pose": {
"shape": [7], "layout": "xyz + quaternion",
"units": "metres", "frame": "world"
}
And the evidence, per episode.
{
// distance from the solved camera pose to the printed marker
"camera_pose_vs_marker": { "median_cm": null, "p90_cm": null },
// share of frames the tracker solved
"registration_rate": null,
// how far reprojected points land from where they were seen
"reprojection_error_px": null,
// frames where the wrist frame flipped about its roll axis
"roll_flips": null,
// detected contact events against a human label
"grasp_state_agreement": null,
// set when the numbers above are measured, not before
"verdict": "pending"
}
Interested? Reach out at contact@humanpriors.com