truelabelRequest dataEarnRequest

Use case · Egocentric data

Egocentric Data for Imitation Learning

An imitation learning dataset is a set of task demonstrations a policy learns to clone. Egocentric human demonstrations — first-person video of a person doing the task — are a cheaper, faster-to-scale alternative to teleoperation-rig data, because humans generate manipulation behavior far faster than a robot seat can. The catch has two parts. First, cloning needs an action signal, and passive first-person video doesn't contain robot actions — so it has to be sensorized or labeled. Second, a human hand is not your robot's hand, so demonstrations need embodiment alignment (retargeting, or paired robot data) before a policy can execute them. EgoScale is the clean evidence for both halves: it reports a log-linear relationship between egocentric data hours and validation loss (R²=0.9983), and separately a 54% average success-rate improvement over no-pretraining on a 22-DoF dexterous hand — and the recipe still uses aligned human-robot mid-training and robot evaluation (EgoScale, paper). So the real planning question isn't "video or teleop." It's which demonstration source supplies the scale, the action labels, and the embodiment alignment your skill needs — and usually the answer is a blend. This page gives you a rubric to choose the blend and a label checklist to spec it. For the wider architecture, see the integrative guide; for the sources, the evidence matrix.

Updated 2026-07-1912 min read
By Truelabel Team
Reviewed by Truelabel Team ·
imitation learning dataset

Quick facts

Resolution
1080p baseline; stereo 2160p on request for depth and 3D-hand fidelity
Field of view
≥120° horizontal, with hands and manipulated objects kept in frame
Mount
Head-mounted (glasses or head-rig), not chest or handheld — preserves the gaze-aligned manipulation viewpoint a policy learns from
Sensors
RGB, IMU (head motion), optional 25-joint-per-hand 3D pose, optional gaze, optional stereo depth
Labels
frame-aligned action/skill segments; per-take task success and failure flags; object and grasp-state annotations
Volume
80–400 accepted demonstration-hours per task family (pilot batch in days)

Key papers

Hard citations for the claims above. Each entry pairs a specific number with the paper that reports it.

  1. EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video

    Hoque et al., Apple · 2025 · arXiv:2505.11709

    829 hours, 194 tasks. EgoDex pairs 829 hours of egocentric video across 194 tabletop tasks with 3D hand and finger tracking captured on Apple Vision Pro — the largest and most diverse dexterous human-manipulation dataset to date.

  2. Ego4D: Around the World in 3,000 Hours of Egocentric Video

    Grauman et al., Meta AI · 2022 · arXiv:2110.07058

    3,670 hours, 74 locations. Ego4D spans 3,670 hours of daily-life first-person video from 931 camera wearers across 74 locations in 9 countries, collected under consenting-participant privacy and de-identification standards.

  3. Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

    Grauman et al., Meta AI · 2024 · arXiv:2311.18259

    1,286 hours, 740 participants. Ego-Exo4D pairs simultaneously-captured egocentric and exocentric video of skilled activity from 740 participants across 13 cities — 1,286 hours with multichannel audio, eye gaze, 3D point clouds, camera poses, and IMU.

How this connects to VLA fine-tuning and robot foundation models

Imitation learning is the layer where a foundation-model plan turns human behavior into policy behavior. Egocentric demonstrations feed it two things web data can't: task intent and affordances. HRP shows the mechanism concretely — hand, object, and contact affordance labels pulled from human video improved robot pretraining by over 15% across 5 real-world tasks and 3,000+ trials, on 3 robot morphologies (HRP[1]). That's human demonstrations teaching a robot not just "this is a mug" but "grasp it here, pour from this lip."

The connection to VLAs is direct: a VLA fine-tuned for a dexterous, contact-rich skill benefits from demonstration data that carries frame-aligned actions and contact events, and egocentric human demos are the scalable source of that behavior — once they're action-labeled and retargeted. The narrower technique underneath most of this is behavioral cloning: the policy learns to reproduce the demonstrated action given the observation, which is exactly why the action label is non-negotiable and why passive video, however plentiful, isn't a substitute. What egocentric demos don't remove is the need for embodiment-aligned data. UMI's own paper is explicit that unstructured human video carries a large embodiment gap, which is why it captures with a sensorized handheld gripper rather than training on passive clips (UMI[2]). Robot-native demonstrations stay in the plan for alignment and evaluation — DROID's 76,000 trajectories across 564 scenes and 86 tasks, or BridgeData V2's ~60,000 language-conditioned trajectories across 24 environments, are the grounding layer human video doesn't supply (DROID[3], BridgeData V2[4]).

The scale-vs-embodiment alignment rubric

Four demonstration sources, ranked on what actually differentiates them: how far they scale, how well they align to your robot, and how strong an action label they carry. Use it to decide the blend per skill.

The rubric encodes one decision rule worth stating plainly: the more your success depends on precise, target-embodiment control, the further down this table you have to buy. Representation and affordances come from the top rows cheaply; deployable control on a specific robot comes from the bottom rows expensively. Most real programs pretrain with the top rows and align/evaluate with the bottom rows — the same shape EgoScale's recipe shows (EgoScale[5]).

SourceScalabilityEmbodiment alignmentAction-label strengthBest useNot enough when…
Passive egocentric videoHighLowLow (none until labeled)Representation & affordance pretraining; task/intent coverageYou need action supervision or target-robot control
Sensorized human demos (glove/wrist-cam/pose)Medium–highMediumMedium–highDexterity priors; a retargetable action bridgeThe retargeting gap to your hand/arm is large
TeleoperationLow–mediumHighHighEmbodiment-specific fine-tuning; precise labelsYou need broad, cheap task/scene coverage
Robot demonstrationsMediumHighHighPolicy training and evaluationYou need scale beyond your hardware/task set

When human first-person demos are enough — and when they aren't

Enough on their own when the goal is representation learning, affordance discovery, or task/procedural coverage: what parts of objects are actionable, how tasks are sequenced, what human intent looks like across variation. Here scalable egocentric video is genuinely the right tool, and adding sensor channels (hand pose, IMU) upgrades it from appearance data to action-relevant data.

Not enough when the goal is deployable control on a specific robot. Three reasons, each sourced:

• No robot actions in raw video. Cloning needs the action; passive clips don't carry joint commands, gripper states, or forces. Sensorization or labeling is required, and even then it's a human action space, not your robot's. • The embodiment gap is real. Human dexterity, reach, and compliance differ from a robot's; UMI designs its whole interface around this rather than pretending the gap away (UMI[6]). • Fine-grained interaction isn't solved. Current egocentric video-language models bias toward objects over temporal dynamics and degrade on fine-grained hand-object interaction, so "more demonstrations" doesn't automatically buy better contact understanding (EgoNCE++[7]).

The honest framing the field converged on: human demonstrations reduce the amount of robot-in-the-loop collection you need; they don't eliminate it. "Teleop is dead" is not what the evidence says.

A worked decision: sensorized human demos vs teleoperation

The rubric is abstract until you point it at a real skill, so here's the decision most teams actually face. Say you're fine-tuning a policy for a bimanual kitchen task — pouring from a jug into a bowl without overfilling — and you have budget for exactly one demonstration source beyond your existing pretraining data. Sensorized human demos or teleoperation?

Walk the rubric columns. Scale and variation: if you need hundreds of takes across jugs, fill levels, bowl positions, and lighting, sensorized human capture (a UMI-style handheld gripper with a wrist camera) collects that breadth far faster than a teleop seat, because a person pours faster than an operator can drive a robot through the same motion (UMI[2]). Embodiment alignment: if your target end-effector is close to a parallel-jaw gripper, the retargeting from a handheld-gripper demo is tractable and the human-demo layer is usable; if your robot has a high-DoF or non-anthropomorphic hand, the retargeting gap widens and teleop's directly-aligned trajectories start to pay for themselves. Action-label fidelity: teleop gives you measured joint and gripper actions with proprioception at capture time; sensorized human demos give you a retargetable action signal that still has to be mapped to your robot. Contact and force: if the skill's success hinges on force control — a controlled pour, a compliant grasp — the proprioceptive and force channels teleop captures are hard to recover from human demos alone.

The honest answer is usually sequenced, not either/or: sensorized human demos for breadth and dexterity priors, then a smaller teleop or robot-native slice for embodiment alignment and a clean evaluation — the same top-then-bottom shape EgoScale's recipe shows, where large egocentric pretraining sits under a much smaller aligned-robot layer (EgoScale[5], DROID[3]). The three spec choices that decide whether the cheaper human-demo layer is even usable: name the retargeting target (which robot hand and arm), state whether contact forces matter for the skill, and reserve the evaluation split before you capture anything. Get those wrong and the human demos become a re-collection, not a saving.

The demonstration label checklist

This is what turns footage into an imitation-learning dataset. A demonstration that's missing the top rows is video, not a demo. Spec these per skill.

• Instruction / task language — what the demonstrator was asked to do, aligned to the clip. • Task-phase boundaries — start/end of each subgoal (approach, contact, manipulate, release). • Action signal — frame-aligned hand/wrist pose, or a retargeted robot action vector; the thing a policy clones. • Contact events — when hand or tool touches the object; the moments dynamics actually happen. • Object identity and state transitions — what changed, and to what. • Success / failure flags per take — and, ideally, deliberate failure/recovery takes. • Optional gaze / audio — where available, useful for intent and narration. • Rights artifacts — wearer consent, location release, per-clip provenance. • QA gates — hands-in-frame threshold, motion-blur/stability limit, annotation agreement, documented rejection reasons. • Delivery schema — LeRobot-compatible episodes, RLDS, MCAP, or custom per-episode JSON (actions, poses, flags).

The two checklist items people skip and later regret are action signal (without it there's nothing to clone) and rights artifacts (without them the demonstrations aren't shippable, however good they look).

What's worth demonstrating — and how densely

The rubric tells you which source to use; this is what to point it at. Imitation learning rewards demonstrations of the contact-rich, dexterous tasks that policies still fail at, captured with enough repetition to give behavior cloning the sample density it needs.

• Bimanual kitchen manipulation — pouring, chopping, opening jars, loading a dishwasher: two-hand, contact-rich sequences with clear success and failure states. • Tool use and light assembly — driving screws, using a drill, threading cables, seating connectors: dexterous sequences where the moment of contact carries most of the signal. • Deformable and articulated-object handling — folding laundry, opening drawers and doors, manipulating bags, cloth, and cabling under natural clutter, where object state is hard to predict from appearance alone. • Retail and warehouse pick-and-place — grasping varied items, bin-to-bin transfers, packing at realistic pace and lighting.

Two things determine whether these become useful demonstrations. The first is take density: behavior cloning learns a distribution, so a handful of takes per skill is rarely enough — repeated demonstrations of the same target skill, across operators and small variations, give the policy the coverage it needs. The exact number is task-dependent, which is why the spec should name it per skill rather than inherit a fixed academic count. The second is matching the taxonomy to your robot: demonstrations of objects and skills your policy will never face are wasted collection. This is the concrete difference between a roundup dataset and demonstrations that transfer — you spec the objects, skills, and success criteria around the policy you're actually training, and EgoScale's scaling evidence only holds when the added hours are diverse and on-distribution (EgoScale[8], paper[5]).

The rights reality: why "action-labeled AND commercially licensed" is the hard part

Here's the procurement judgment that TrueLabel actually adds, and it's the one the existing page got right: the useful intersection — demonstrations that are both action-labeled and cleanly licensed for commercial training — is thin in the open. Plenty of open egocentric corpora are large but unlabeled (usable for representation, not demonstrations), and plenty of the highest-fidelity annotated corpora ship under non-commercial or no-derivatives terms (great for research, unusable in a product you sell). You can experiment on public data; shipping a policy trained on it is a different question that a license and intended-use review has to answer.

That gap is why teams commission custom demonstration capture: it can be scoped for intended-use review with consent artifacts, provenance, and buyer-reviewed license terms, captured to your demonstration taxonomy — your objects, your skills, your success criteria, your number of takes per task — with per-clip consent and per-trajectory provenance. TrueLabel delivers exactly that: rights-cleared demonstration data with consent artifacts and per-trajectory provenance, in RLDS, LeRobot, MCAP, or custom schemas, with a sample packet and QA evidence before scale. No claimed sample-efficiency multiplier, no delivery-time commitment, and no model-performance or legal-clearance promise; license and intended-use review sit with your counsel, and the parts we can state are enough to plan a defensible run.

Use these to move from category-level context into specific task, dataset, format, and comparison detail.

External references and source context

  1. HRP: Human Affordances for Robotic Pre-Training

    Primary or official source cited by the authored page

    arXiv ↩
  2. Project site

    Primary or official source cited by the authored page

    umi-gripper.github.io ↩
  3. Project site

    Primary or official source cited by the authored page

    droid-dataset.github.io ↩
  4. Project site

    Primary or official source cited by the authored page

    rail-berkeley.github.io ↩
  5. EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data

    Primary or official source cited by the authored page

    arXiv ↩
  6. Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots

    Primary or official source cited by the authored page

    arXiv ↩
  7. Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?

    Primary or official source cited by the authored page

    arXiv ↩
  8. EgoScale: Scaling Human Video to Unlock Dexterous Robot Intelligence

    Primary or official source cited by the authored page

    NVIDIA Research GEAR Lab ↩
  9. VLAs, World Models, and Egocentric Data

    Authored internal route

    truelabel.ai
  10. Robot Foundation Model Data Evidence Matrix

    Authored internal route

    truelabel.ai

FAQ

Why not just use a public egocentric dataset for imitation learning?

Because the demonstrations you can legally ship a policy on are the hard part. Large open egocentric corpora are often unlabeled (representation data, not demonstrations), and the highest-fidelity annotated ones frequently carry non-commercial or no-derivatives licenses. Public data is excellent for scoping and benchmarking; commercial training needs a license and intended-use review, and often custom, action-labeled capture.

Can human egocentric demonstrations replace teleoperation?

They complement it. Human capture collects far more demonstration behavior per unit effort than a teleop seat, and it reduces how much robot-in-the-loop collection you need — but it carries an embodiment gap and no native robot actions, so teleop or robot demonstrations still close the alignment and evaluation layers (UMI, DROID).

How do you get an action signal from monocular first-person video?

By capturing side channels: hand/wrist pose, head IMU, and — where the skill needs it — depth, time-synced to the RGB stream, then retargeted toward the robot's action space. That's what turns a clip into a demonstration a policy can clone; EgoScale's recipe uses exactly this kind of action-labeled, retargeted egocentric data (EgoScale).

Does more demonstration data always help?

More diverse, action-labeled data helps predictably in the settings where scaling laws have been measured (EgoScale reports R²=0.9983 between egocentric hours and validation loss). But raw hour count without action labels, or without target-embodiment alignment, hits diminishing returns fast — and fine-grained interaction understanding remains a known weak spot (EgoNCE++).

What should a buyer specify for an imitation-learning dataset?

Use the label checklist above: instruction language, task-phase boundaries, an action signal (hand/wrist pose or retargeted actions), contact events, object-state transitions, success/failure flags, rights artifacts, QA gates, and a delivery schema — all specified around your robot's skills and objects, not a fixed academic task list.

Planning an imitation-learning run?

Draft the demonstration taxonomy, label checklist, and rights requirements with the data-spec generator, or post a spec to scope a rights-cleared sample packet with QA evidence. TrueLabel delivers action-labeled human demonstrations — with the robot-demo slice you need for alignment — in RLDS, LeRobot, MCAP, and custom schemas, with per-trajectory provenance.

Post a spec