Vision-language-action models
VLA training data
VLA training data is the set of synchronized triplets a vision-language-action policy learns from: a visual observation, a language instruction, and the executed action at each timestep. The mistake teams make when they buy it is treating it as one homogeneous asset. It isn't. A workable VLA dataset is a mix — egocentric human video for affordance and dynamics pretraining, sensorized human demonstrations for action structure, robot or teleoperation data for embodiment-aligned actions, and a separate held-out set for evaluation. The reason the mix matters is evidence-backed: OpenVLA, a 7B model trained on 970,000 robot episodes, reports outperforming the 55B RT-2-X by 16.5% absolute in its own evaluation (OpenVLA) — source-specific evidence that a well-chosen data mix can outweigh raw model scale in that evaluation, not a universal law that diversity always beats parameter count. But its model card is equally clear that zero-shot use is bounded by the embodiments and domains in that mix (model card). TrueLabel is a physical AI data marketplace: you post a spec, vetted suppliers return samples. This page explains how to specify the mix for a VLA program, what each stream can and cannot supply, and how to scope a rights-cleared sample packet before you fund scale. For the full conceptual model, see the integrative guide on how VLAs, world models, and egocentric data fit together.
Quick facts
- Observation
- Egocentric, wrist, or external video
- Task context
- Instruction text, object set, environment, success criteria
- Action
- Joint states, end-effector poses, or discrete action labels
- Format
- RLDS, LeRobot, HDF5, MCAP, or buyer-defined schema
- QA
- Observation-action sync and complete task episodes
Comparison
| Model need | Data requirement | Bounty implication |
|---|---|---|
| Visual grounding | Diverse real-world observations | Specify camera and environment |
| Language grounding | Instructions and task labels | Define task taxonomy |
| Action grounding | States, actions, trajectories | Set sync and format rules |
| Evaluation | Held-out accepted episodes | Start with an eval request |
Provider list — VLA training data
10 providers covering VLA training data. Each entry summarizes the provider's strongest fit and a buyer-bottleneck signal so you can shortcut the discovery loop.
#1
OpenVLA
Open-source 7B-parameter VLA model from Stanford / TRI / UC Berkeley — released with weights and training recipe.
Best for: Reference model when designing VLA training data shape (vision token format, action representation, instruction grounding).
#2
RT-2 (Google DeepMind)
Vision-language-action model that co-fine-tunes web-scale VLM with robotics data — defines the modern VLA benchmark.
Best for: Architecture reference; the data recipe (web VLM data + robot trajectories) is the template most production VLA programs follow.
#3
π0 (Physical Intelligence)
Foundation VLA model from Physical Intelligence trained on a large mix of teleop, manipulation, and language data.
Best for: Frontier VLA reference; informs scale and diversity requirements for production training data.
#4
NVIDIA GR00T N1
NVIDIA's open VLA foundation model for humanoids with synthetic-data-heavy training recipe and public weights.
Best for: Sim-first VLA training pattern; useful when synthetic data is part of the production mix.
#5
Open X-Embodiment / RT-X
22-institution cross-embodiment dataset that anchors the dominant VLA pretraining recipe.
Best for: Cross-robot VLA pretraining corpus before deployment-specific fine-tune.
#6
DROID
76k Franka demonstrations with synchronized vision + language annotation in many cases.
Best for: Real-world manipulation slice for VLA fine-tune when single-arm Franka matches deployment.
#7
BridgeData V2
60,096 instruction-conditioned manipulation trajectories with language labels.
Best for: Affordable, well-documented VLA training data with strong instruction grounding.
#8
Hugging Face LeRobot Bridge / DROID variants
Curated LeRobot conversions of canonical VLA datasets (Bridge, DROID, ALOHA) in modern Parquet format.
Best for: Off-the-shelf ingestion path for VLA training when you want modern format conventions.
#9
RoboCat training set
DeepMind self-improving foundation agent — reference for VLA scaling laws.
Best for: Architecture-of-thought reference for self-improvement loops; underlying corpus not redistributable.
#10
SayCan
Affordance-grounded language-to-action work from Google — defines the language → action grounding pattern many VLAs imitate.
Best for: Reference for language grounding shape in VLA training data (especially affordance tags + step verbs).
What VLA training data actually is
A vision-language-action model unifies perception, language grounding, and motor control in one backbone. Unlike a vision-language model that emits text, a VLA emits actions conditioned on what it sees and what it's told to do. RT-2 defined the recipe: co-fine-tune a vision-language model on both web vision-language tasks and robot trajectories, and express actions as tokens. RT-2 reported on the order of 6,000 robot evaluation trials and, notably, emergent semantic behavior — following instructions about objects it never saw in the robot data (RT-2[1]). That's the value a VLA imports from the web: it knows nouns.
What the web can't teach is the sensorimotor mapping. A model that recognizes a mug still can't infer the gripper trajectory to grasp it without action-labeled examples. That's why the open, reproducible VLA — OpenVLA — is trained on 970,000 robot episodes drawn from Open X-Embodiment, and it's worth stating precisely because the web repeats an error: OpenVLA predicts discretized action tokens, not a diffusion action head (OpenVLA[2]). So the shape of a VLA dataset follows directly from the architecture: at each timestep you need the observation, the instruction, and the action vector, time-synchronized, in a format a training pipeline can read (RLDS, LeRobot-compatible structure, or MCAP).
The harder part is not the format. It's that "the action vector" is embodiment-specific — a 7-DoF arm's joint angles are not a parallel-jaw gripper's open/close is not a mobile base's velocity. Cross-embodiment training exists (Open X-Embodiment aggregated data across a wide range of robot platforms), but it requires per-embodiment normalization and, in practice, target-robot fine-tuning. Which is the whole reason the mix has to be designed rather than bought off a shelf.
Where egocentric video and world models fit in a VLA data plan
This is the section most VLA buyers under-plan. Egocentric — first-person — human video has become a serious VLA ingredient because it densely records hands, tools, contact, and object-state change from the actor's viewpoint, at a scale robot teleoperation can't match. EgoScale is the sharpest evidence: pretraining a VLA on 20,854 hours of action-labeled egocentric video (wrist motion plus retargeted hand actions), then doing aligned human-robot mid-training, produced a near-perfect log-linear scaling law between data hours and validation loss (R²=0.9983) and a 54% average success-rate improvement over no-pretraining on a 22-DoF dexterous hand (EgoScale[3], paper[4]).
Read that recipe carefully, because it's the whole argument in one result: the scaling win came from egocentric hours plus aligned robot data plus action labels plus robot evaluation. Not from raw video. HRP makes the same point from a different angle — hand/object/contact affordance labels extracted from human video improved robot pretraining by over 15% across 5 tasks and 3,000+ trials, in that setup (HRP[5]).
World models are the other half of the frontier. Instead of mapping observation-and-instruction to action, they predict what happens next given an action. Action-conditioned video prediction has supported manipulation planning since Deep Visual Foresight (Deep Visual Foresight[6]); the modern joint future-video-and-action version (NVIDIA's DreamZero) reports over 2× improvement versus VLA baselines on new tasks in its lab evaluation (DreamZero[7]). For a buyer, the practical takeaway is that VLA-style policy data and world-model-style sequence data are converging, and NVIDIA's own framing is that the winner may be a hybrid rather than one approach killing the other (NVIDIA WAM[8]). Plan for the mix, not the monoculture.
The VLA data-mix selector
Use this to decide what to collect at each stage of a VLA program. "Missing signal" is the column people skip — it's the thing that stage of data won't give you, and therefore what the next stage has to cover.
The rule the selector encodes: the cheap, scalable stream (egocentric) and the expensive, aligned stream (robot) are complements. Buy only the first and you don't have a cheaper policy — you have an unvalidated one. The evidence matrix puts a primary source behind each row of this logic.
| Stage | Best data stream | What it teaches | Missing signal | Labels to require | TrueLabel scoping question |
|---|---|---|---|---|---|
| Semantic pretraining | Web + egocentric human video | Objects, language, affordances, hand-object timing | Robot actions, proprioception | Task narration, hand/wrist pose, contact/object state | "Which environments and tasks must the first-person distribution cover?" |
| Instruction grounding | Language-annotated demonstrations | Mapping instructions → task phases | Fine-grained low-level control | Dense instructions aligned to temporal segments | "How dense should language be, and in which languages?" |
| Imitation / demonstration | Sensorized human demos or robot demos | Action-relevant structure, task intent | Direct target-embodiment fit (human) or scale (robot) | Frame-aligned actions, success/failure, phase boundaries | "Passive video, sensorized human, or robot demos for this skill?" |
| Embodiment fine-tuning | Teleoperation / robot-native demos | Embodiment-aligned actions + proprioception | Cheap scale, task breadth | Joint/gripper/base actions, proprioception, calibration | "Which robot, action space, and control rate?" |
| Evaluation | Held-out robot demonstrations | Whether the policy actually works | Nothing — it's the truth layer | Task success criteria, failure taxonomy, split rules | "What's the acceptance threshold and eval split?" |
What egocentric video does not replace
For deployment-relevant grounding you still need robot-native data. DROID contributes 76,000 trajectories (350 hours) across 564 scenes and 86 tasks with synchronized RGB, calibration, depth, and language (DROID[9]); BridgeData V2 contributes roughly 60,000 real trajectories across 24 environments with language and goal conditioning (BridgeData V2[10]). Neither is universal coverage — both are starting points you check against your environment and embodiment.
Teleoperation, human demos, and the alignment layer
Teleoperation yields high-quality, precisely-labeled demonstrations, but it scales with operator hours and robot uptime — a real ceiling, not a rhetorical one. The response the field converged on isn't "kill teleop." It's to move as much collection as possible up the scale curve — egocentric and sensorized human data for breadth — and reserve teleop and robot demonstrations for the embodiment-alignment and evaluation layers where nothing else substitutes.
Sensorized human demonstration is the bridge between the two. UMI's handheld gripper with a wrist camera collects human demonstrations that transfer to robot policies, precisely because the capture interface encodes robot-relevant action structure that passive video lacks (UMI[11]). For a VLA buyer, the decision is per-skill: passive egocentric video for representation and affordances, sensorized human demos where you need action structure at scale, and robot/teleop demos where you need embodiment-aligned control and a clean evaluation. The imitation-learning rubric works through that decision in detail.
Three ways a VLA data buy goes wrong
These are the failure modes worth naming before you post a spec, because each is expensive to fix after collection.
• Buying hours instead of labels. Ten thousand hours of unlabeled first-person video is a representation-pretraining asset, not demonstrations. If your policy needs action supervision, raw hour count is the wrong headline — frame-aligned actions, contact events, and phase boundaries are. EgoScale's result is instructive precisely because it's action-labeled ego video, not raw footage, that scales against loss (EgoScale[4]). • Collecting for the wrong embodiment. Data captured with an action space that doesn't map to your robot forces post-hoc retargeting and, often, a re-collection. Specify the target robot, action space, and control rate before capture, not after. This is the concrete meaning of OpenVLA's model-card caveat that zero-shot use is bounded by represented embodiments (OpenVLA card[12]). • Skipping the evaluation split. A training set with no held-out physical evaluation gives you a number you can't trust. Plausible-looking predictions — especially from world-model rollouts — are not physical success, so reserve a robot-native evaluation set with explicit acceptance criteria from the start.
Rights, provenance, and why "public dataset" is not "cleared to ship"
Public datasets are evidence and benchmarks, not automatic commercial supply. Many large egocentric and robot corpora carry research-only or non-commercial terms, and even permissively-licensed video may lack the consent chain and provenance a commercial training run needs. The honest position — the one this cluster holds throughout — is that a dataset's existence is not permission to train a commercial model on it; license and intended-use review is a required step, not a formality.
This is where TrueLabel's role is concrete and deliberately unhyped. TrueLabel delivers rights-cleared data: contributor consent artifacts, location releases where applicable, and per-trajectory provenance and metadata. Delivery is in RLDS, LeRobot, MCAP, or custom schemas, to S3, GCS, or Azure. What you will not find on this page: a legal-clearance promise, a delivery-time commitment, or a promised model gain. Those are either not shipped capabilities or not things any honest data vendor should promise — and the facts we can state are enough to plan a defensible dataset.
How TrueLabel scopes a VLA dataset
The marketplace model is simple: you post a spec describing the task, embodiment, viewpoint, sensors, labels, rights, and QA acceptance; vetted suppliers return sample packets with QA evidence; you review samples before committing to scale. TrueLabel works with around 10,000 collectors across 100 countries and maintains a research catalog of 750+ profiled public and commercial physical-AI datasets, spanning egocentric, exocentric, teleoperation/robot-demonstration, and directed-capture modalities. Sample packets come before scale, with QA evidence included for buyer review.
Concretely, a VLA data request specifies:
• Task and embodiment — the skill, the target robot, the action space, and the control rate. • Viewpoint and sensors — first-person camera position, field of view, frame rate, plus IMU, optional hand/wrist pose, optional depth. • Labels — frame-aligned actions, task-phase boundaries, contact/object-state, success/failure flags, dense instruction language. • Rights — consent artifacts, location releases, per-trajectory provenance, and intended-use terms. • QA and delivery — hands-in-frame and stability thresholds, annotation agreement, rejection reasons, and a delivery schema (RLDS/LeRobot/MCAP/custom).
Draft the whole thing with the data-spec generator, or post a spec to get a rights-cleared sample packet scoped to your task.
Related pages
Use these to move from category-level context into specific task, dataset, format, and comparison detail.
External references and source context
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Primary or official source cited by the authored page
Proceedings of Machine Learning Research ↩ - OpenVLA: An Open-Source Vision-Language-Action Model
Primary or official source cited by the authored page
arXiv ↩ - EgoScale: Scaling Human Video to Unlock Dexterous Robot Intelligence
Primary or official source cited by the authored page
NVIDIA Research GEAR Lab ↩ - EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data
Primary or official source cited by the authored page
arXiv ↩ - HRP: Human Affordances for Robotic Pre-Training
Primary or official source cited by the authored page
arXiv ↩ - Deep Visual Foresight for Planning Robot Motion
Primary or official source cited by the authored page
arXiv ↩ - DreamZero: World Action Models are Zero-shot Policies
Primary or official source cited by the authored page
NVIDIA Research ↩ - Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models
Primary or official source cited by the authored page
NVIDIA Technical Blog ↩ - Project site
Primary or official source cited by the authored page
droid-dataset.github.io ↩ - Project site
Primary or official source cited by the authored page
rail-berkeley.github.io ↩ - Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots
Primary or official source cited by the authored page
arXiv ↩ - OpenVLA 7B model card
Primary or official source cited by the authored page
OpenVLA on Hugging Face ↩ - VLAs, World Models, and Egocentric Data
Authored internal route
truelabel.ai - Project site
Primary or official source cited by the authored page
umi-gripper.github.io - Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
Primary or official source cited by the authored page
arXiv
FAQ
Why do VLA models need action labels if vision-language pretraining is so strong?
Vision-language pretraining teaches semantics — object recognition, spatial relations, goals — but not the sensorimotor mapping needed to act. RT-2 showed web pretraining improves generalization when fine-tuned on robot demonstrations, but the fine-tuning is mandatory (RT-2). OpenVLA showed a 7B model beating a 55B one by leaning on 970,000 action-labeled robot episodes (OpenVLA). Semantics get you a model that knows what a mug is; action labels get you one that can grasp it.
Can egocentric human video replace robot demonstrations for a VLA?
No — it complements them. It's a strong, scalable source for affordance and dynamics pretraining (EgoScale, HRP), but it carries no robot action vectors or proprioception, and unstructured human video has a real embodiment gap (UMI). Deployment-ready control still needs sensorized demos or robot/teleop data aligned to your embodiment, plus held-out evaluation.
What data format should VLA training data be delivered in?
RLDS or a LeRobot-compatible episode structure for observation-action episodes, MCAP for multi-sensor logs, or a custom schema. TrueLabel delivers in RLDS, LeRobot, MCAP, and custom schemas to S3, GCS, or Azure, with per-trajectory provenance and consent artifacts.
Are public VLA datasets commercially usable?
It depends on each dataset's license and intended-use terms — many are research-only or non-commercial, and even permissive video may lack a consent chain. Treat public datasets as evidence and benchmarks, and run a license/intended-use review before any commercial training.
How do I scope a VLA dataset without over-buying?
Use the data-mix selector above: match each stage (pretraining, instruction grounding, imitation, embodiment fine-tuning, evaluation) to the stream that supplies its signal, and specify the labels each stage needs. Then post a spec for a sample packet before committing to scale.
Building or fine-tuning a VLA?
Draft your capture, label, and rights requirements with the physical AI data-spec generator, or post a spec to scope a rights-cleared sample packet with QA evidence before you fund scale. TrueLabel is a physical AI data marketplace — post a spec, get samples from vetted suppliers, review before you commit.
Post a spec