truelabelRequest dataEarnRequest

Plain-language bridge

First-Person Video Data

First-person video data is video captured from the viewpoint of the person, wearable device, or robot experiencing a task. In physical AI it overlaps heavily with what researchers call egocentric video — "first-person" is the plain-language term many buyers use; "egocentric" is the technical one. It matters for robot training because it records what the actor sees: hands approaching objects, tools in use, the moment of contact, and how a scene changes when someone acts. The thing to plan from day one is rights: first-person footage routinely captures faces, voices, screens, homes, and bystanders, so consent, redaction, and licensing are design decisions, not afterthoughts. This page is the terminology-and-rights entry point. It explains how first-person, egocentric, robot POV, and exocentric relate; where first-person video fits in VLA and world-model data plans; and what governance to plan before capture. For the full data architecture, follow the advanced path to the integrative guide.

Updated 2026-07-1912 min read
By Truelabel Team
Reviewed by Truelabel Team ·
first-person video data

Quick facts

Ego4D benchmark scale
Use Ego4D as a public first-person reference for scoping, not supply: about 3,670 hours across 74 locations and 9 countries.
Ego-Exo4D paired-view scale
Use Ego-Exo4D when the task needs paired ego/exo context: 740 participants or camera wearers, 13 cities, 123 sites, and about 1,286 hours.
EPIC-KITCHENS-100 boundary
Use EPIC-KITCHENS-100 for kitchen-action terminology: 100 hours, 45 kitchens, 20M frames, and non-commercial public annotation terms.

Comparison

First-Person Video Data comparison table
Capture methodWhen it helpsConstraint to plan
Wearable cameraHuman task demonstrationsConsent, bystanders, and stable framing
AR glassesHands-free task contextDevice availability and privacy review
Head-mounted action cameraHands and tools in frameComfort, calibration, and motion blur
Robot-mounted cameraRobot embodiment contextSensor sync and task metadata

First-person vs egocentric vs robot POV vs exocentric

These terms get used loosely. For procurement, the distinctions are small but they change what you spec.

• First-person video — the buyer-facing, plain-language term for footage shot from the actor's viewpoint. Broad; includes any actor (human or device). • Egocentric video — the technical term in computer vision and robotics for the same actor-viewpoint footage, usually implying a head- or body-mounted camera on a human and often paired with sensors (IMU, gaze, hand pose). • Robot POV — actor-view footage where the "actor" is the robot, from a head- or wrist-mounted camera. It looks first-person but carries robot embodiment context and, usually, robot action/proprioception streams. • Exocentric video — third-person, external-view footage of the task. It shows global body pose and workspace layout that a first-person view occludes, which is why paired ego/exo datasets exist.

You don't need to relitigate these definitions every time; the glossary owns the formal ones. The practical point is that "first-person" and "egocentric" usually name the same footage, while "robot POV" and "exocentric" name genuinely different signals you might add on purpose.

Why first-person video appears in VLA and world-model data plans

First-person video shows up in serious robot-training plans for one reason: it records the actor's viewpoint, which is where near-field manipulation is legible. Web and third-person video teach what tasks look like from outside; first-person video shows the reach, contact, and object-state change in the frame where a policy has to decide. The reason it feeds VLA and world-model training specifically is a division of labor: vision-language-action policies of the kind RT-2 defined import object semantics from the web, then need the actor-view sequence of reach, contact, and state change that first-person video supplies (RT-2[1]). And that egocentric pretraining has measurable downstream effect — EgoScale reports a log-linear relationship between egocentric data hours and validation loss on a dexterous-manipulation setup (EgoScale[2], paper[3]). Research corpora exist at meaningful scale — Ego4D spans 3,670 hours across 74 locations and 9 countries (Ego4D[4]), and Ego-Exo4D pairs first- and third-person views across 1,286.3 hours of skilled activity (Ego-Exo4D[5]) — which is why teams reach for it.

A concrete example makes the value obvious. Watch someone pour from a jug in first-person and the causal chain is right there in the frame the way a policy needs it: the hand enters, approaches the handle, grips, tilts, liquid moves, the target fills, the hand rights the jug. A third-person clip of the same pour shows that it happened; the first-person clip shows the reach timing, the grip point, and the contact moment — the parts a manipulation policy actually has to reproduce. Multiply that across thousands of everyday tasks and you have why first-person corpora became a training substrate rather than a curiosity.

But "appears in the plan" is not "is the whole plan." First-person video is a scale and affordance layer; it becomes training-ready only when paired with labels (task language, hand/wrist pose, contact and object-state), and for deployment it still gets aligned with robot data. The viewpoint gives you the raw causal signal; the labels and the robot-alignment layer are what turn it into something a policy can be trained and evaluated on. That full stack — and what each layer proves — is the subject of the integrative guide and the evidence matrix. If you're already choosing between demonstration sources, the imitation-learning rubric is the next step; if you're pretraining a dynamics model, the world-models page covers the sequence side.

Buyer-term translation table

Buyers and researchers use different words for the same thing, and the mismatch causes mis-scoped requests. This table translates the common buyer terms into what they mean technically, what they're used for, what they're usually missing, and the rights concern attached.

The row that saves the most rework is the first one: a request for "first-person video" often means the buyer wants demonstration-grade egocentric data with hand pose and action labels — but the phrase alone doesn't say so. Naming the missing labels up front is the difference between a usable dataset and a folder of clips.

Buyer termTechnical termTypical model useMissing by defaultConsent / rights concern
First-person videoEgocentric videoRepresentation & affordance pretrainingAction labels, proprioceptionWearer consent; bystanders in frame
POV footageEgocentric RGBActivity/action recognition, hand-object studyDepth, pose, action vectorsFaces, screens, private spaces
Body-cam / wearable videoEgocentric (body-mounted)Task context, procedural coverageFine hand detail (mount is low/off-axis)Location releases; workplace/home privacy
AR-glasses captureEgocentric (head-mounted + sensors)Hands-free task context, gaze, IMURobot-embodiment fitDevice consent; bystander capture
"Robot's-eye" videoRobot POVEmbodiment context, on-robot policy dataHuman-scale diversityUsually low PII; sensor-sync metadata
Overhead / room videoExocentric videoScene layout, third-person task stateActor-view intent, near-field occlusionBystanders; multi-person scenes

Common first-person capture pitfalls

Most first-person datasets that disappoint fail for the same handful of reasons. Naming them makes them cheap to avoid.

• Wrong mount, lost detail. A comfortable chest or handheld mount is easy to wear but drops the hands out of the near-field, gaze-aligned frame that makes first-person data worth collecting. For manipulation, the mount is a data-quality decision, not a comfort one. • Occlusion at the moment that matters. Hands and objects occlude each other exactly when contact happens — the highest-value instant. Plan camera position and, where needed, a paired exocentric view so the key moment isn't a blur. • Ego-motion treated as noise. The camera moves with the wearer's head, so without a synced IMU or head-pose track, "the object moved" and "I looked away" are indistinguishable. Capture the motion channel; don't try to recover it later. • Incidental PII collected without a plan. Faces, screens, and bystanders enter frame by accident. Deciding redaction and consent rules after collection is expensive and sometimes unrecoverable — plan them before the first clip. • Footage without labels called a "dataset." Raw clips are a starting point; task boundaries, contact events, and an action or pose signal are what make them usable. And even well-labeled first-person video has limits — current egocentric video-language models skew toward recognizing objects over temporal dynamics, so more footage doesn't automatically buy better fine-grained interaction understanding (EgoNCE++[6]). The buyer-term translation table above exists to catch this before you buy.

Capture methods and what each is good for

How the camera is worn changes what the footage can teach. Match the method to the signal you need.

A head-mounted mount that keeps hands and manipulated objects in frame is the workhorse for manipulation data, because it preserves the gaze-aligned viewpoint a policy learns from. A chest- or handheld mount is easier to wear but loses the near-field detail that makes first-person data valuable in the first place.

Capture methodWhen it helpsConstraint to plan
Wearable cameraHuman task demonstrationsConsent, bystanders, stable framing
AR glassesHands-free task context, gaze/IMUDevice availability and privacy review
Head-mounted action cameraHands and tools in frameComfort, calibration, motion blur
Robot-mounted cameraRobot embodiment contextSensor sync and task metadata

Data quality requirements

First-person video that will train a model has to clear quality gates that casual footage never does. At minimum:

• Hands and manipulated objects in frame for the bulk of manipulation moments — the whole point of the viewpoint. • Camera stability within a motion-blur and non-nausea threshold; ego-motion is already a modeling challenge without shaky capture making it worse. • Clear task boundaries — start and end of each task and subgoal — so segments are usable. • Task completion / success labeling per take, including deliberate failure and recovery where it's useful. • Metadata completeness — timestamps synced across video and any sensor streams, capture context, and source documentation. • Rights suitability — whether the source actually has the rights for your intended use.

Treat these as acceptance criteria in the spec, with documented rejection reasons, not aspirations. A QA sample packet up front is how you check them before committing to scale.

Consent and privacy: plan governance before capture

First-person video is privacy-dense by nature. It can record identifiable faces, voices, screens, homes, workplaces, bystanders, and location context — often incidentally. The right posture is to treat viewpoint video as sensitive by default and design governance before collection starts.

This is also where the "public dataset" trap lives. Public first-person corpora such as Ego4D and Ego-Exo4D are excellent for scoping terminology, modalities, and benchmarks — but they are evidence and reference sources, not automatic commercial training supply. A dataset's availability for research is not permission to train and ship a commercial model on it; license and intended-use review is a separate, required step, and it belongs at the start of a data plan rather than at the end when a model is already trained. TrueLabel's role is to deliver first-person data that is rights-cleared from the start: contributor consent artifacts, location releases where applicable, and per-trajectory provenance, in RLDS, LeRobot, MCAP, or custom schemas.

A first-person consent-and-governance workflow

Governance reads like a legal checklist, but for first-person capture it's really a capture-spec input — each item below is a field with an acceptance gate, the same way a blur threshold is. Run these steps before the first clip, in this order, because each one constrains the next.

1. Define intended use first. Name what the data will train and where the model ships. This is the step that sets everything downstream — the consent language, the license terms, and the redaction bar all depend on whether this is internal research or a shipped commercial product. 2. Capture informed participant consent. The wearer consents to the specific intended use, not a generic "you're being recorded." Vague consent collected for one purpose rarely covers a later commercial one. 3. Decide bystander handling up front. Incidental people, faces, and screens enter a first-person frame constantly. Choose the rule before capture: avoid-capture zones, posted on-site notice, or a mandatory redaction pass — and write it into the spec. 4. Secure location releases. For capture in homes, workplaces, or other private or commercial spaces, get the location owner's release. A wearer's consent does not cover the room. 5. Run a de-identification review. Blur or remove faces, screens, and documents in a pass before the data leaves the capture pipeline, not after a model has already trained on it. 6. Log provenance per clip. Record who captured, where, under what consent basis, and with what device, so the chain is auditable later. Provenance you didn't capture at the source is usually unrecoverable. 7. Complete an intended-use and license review with counsel. The final gate before training — the exact step a public-dataset download lets you skip and a commercial run cannot.

The reason to treat this as a workflow rather than a disclaimer is that a downloaded research corpus gives you none of it: no per-clip consent tuned to your use, no bystander policy, no auditable provenance. That absence is precisely why commercial first-person capture is commissioned with the workflow built in, and why the license-and-intended-use question belongs at the start of the plan, not after a model already exists.

Use these to move from category-level context into specific task, dataset, format, and comparison detail.

External references and source context

  1. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

    Primary or official source cited by the authored page

    Proceedings of Machine Learning Research ↩
  2. EgoScale: Scaling Human Video to Unlock Dexterous Robot Intelligence

    Primary or official source cited by the authored page

    NVIDIA Research GEAR Lab ↩
  3. EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data

    Primary or official source cited by the authored page

    arXiv ↩
  4. Ego4D

    Primary or official source cited by the authored page

    Ego4D Consortium ↩
  5. Ego-Exo4D project site

    Primary or official source cited by the authored page

    ego-exo4d-data.org ↩
  6. Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?

    Primary or official source cited by the authored page

    arXiv ↩
  7. truelabel egocentric data licensing hub

    Authored internal route

    truelabel.ai
  8. Egocentric video privacy and consent

    Authored internal route

    truelabel.ai
  9. First-person video data in Europe

    Authored internal route

    truelabel.ai

More glossary terms

FAQ

What is first-person video data?

It's video recorded from the viewpoint of the actor, device, or robot experiencing the task. In physical AI it usually means egocentric human video — head- or body-mounted footage centered on hands and manipulated objects.

Is first-person video the same as egocentric video?

They usually name the same footage. "Egocentric" is the technical term used in computer vision and robotics; "first-person" is the plain-language term many buyers use. Small nuance: "egocentric" often implies a human wearer plus sensors, while "robot POV" is the robot's own actor view.

How is first-person video used in AI training?

As a scalable source of actor-viewpoint behavior — object approach, tool use, task sequencing, hand-object interaction — for representation and affordance pretraining, and, once labeled with pose and actions, for imitation learning and world-model sequences. It's paired with robot data for deployment, not used alone.

How do teams collect first-person video data safely?

They define participant consent, bystander handling, location rules, capture hardware, retention, de-identification review, and QA gates before collection starts — and keep a documented consent chain per clip. Viewpoint video is treated as sensitive by default.

When should I add an exocentric (third-person) view?

Add it when the task needs global body pose or workspace layout that the first-person view occludes — long-horizon or whole-body tasks, or anything where you need to label subgoal boundaries from outside. For close, near-field manipulation where the actor's hands carry the signal, first-person alone is usually the cleaner source. Paired ego/exo datasets exist precisely for the cases where you need both.

Can I train a commercial model on public first-person datasets?

Only after a license and intended-use review. Public corpora like Ego4D and Ego-Exo4D are strong reference and benchmark sources, but availability for research is not commercial-use clearance; commercial training often needs domain-specific, rights-cleared capture.

Scoping a first-person capture plan?

Draft the capture, label, consent, and delivery requirements with the data-spec generator, or post a spec to review a consent-ready, rights-cleared sample packet with QA evidence. TrueLabel is a physical AI data marketplace — post a spec, get samples from vetted suppliers, review before you scale.

Post a spec